chart-token-latency
A four-state LLM serving-latency scatter: time to first token against time per output token for every request, nearest-rank p50/p95 crosshairs, an inclusive budget crosshair that cuts the plane into four filterable zones, and queue/prefill/decode phase bars for the median and p95 requests.
chart-token-latency
A four-state LLM serving-latency scatter: time to first token against time per output token for every request, nearest-rank p50/p95 crosshairs, an inclusive budget crosshair that cuts the plane into four filterable zones, and queue/prefill/decode phase bars for the median and p95 requests.