Quantized LLM serving on an NVIDIA A40

Llama-3.1-8B-Instruct · open-loop Poisson arrivals, at least 3 repeats per point, clocks locked, kernels asserted from engine logs.

Every point below is a rate a server could actually be provisioned for. Oversubscribed runs — where offered load exceeded capacity and latency reflects run duration rather than load — are excluded here and published in the repository.

Pick a latency budget

500 ms

Because every configuration ran on the same GPU, the hourly price is a linear scalar and cannot reorder this ranking. The latency budget can, and does — it decides which offered rates are admissible, which caps sustainable throughput, which sets cost per token.

Latency vs throughput

Click a legend entry to hide a configuration. The dashed line is the current budget; points above it do not meet the SLA. Log scale on the y axis, because tail latency spans orders of magnitude between the linear region and saturation.

Operating points

ConfigRateThroughputTTFT p95 TPOT p95$/1M tokensRepeatsSLA

Throughput is mean across repeats. Cost assumes $0.40/GPU-hour (RunPod A40 community cloud, accessed 2026-08-10) — an assumption, stated as one, because the measurement GPU is a lab machine with no invoice.