Llama-3.1-8B-Instruct · open-loop Poisson arrivals, at least 3 repeats per point, clocks locked, kernels asserted from engine logs.
Every point below is a rate a server could actually be provisioned for. Oversubscribed runs — where offered load exceeded capacity and latency reflects run duration rather than load — are excluded here and published in the repository.
Repository Findings Methodology
Because every configuration ran on the same GPU, the hourly price is a linear scalar and cannot reorder this ranking. The latency budget can, and does — it decides which offered rates are admissible, which caps sustainable throughput, which sets cost per token.
Click a legend entry to hide a configuration. The dashed line is the current budget; points above it do not meet the SLA. Log scale on the y axis, because tail latency spans orders of magnitude between the linear region and saturation.
| Config | Rate | Throughput | TTFT p95 | TPOT p95 | $/1M tokens | Repeats | SLA |
|---|
Throughput is mean across repeats. Cost assumes $0.40/GPU-hour (RunPod A40 community cloud, accessed 2026-08-10) — an assumption, stated as one, because the measurement GPU is a lab machine with no invoice.