Qwen3-8B on vLLM 0.9.3, single 4090
BetterBench v1 corpus, 20 measured passes after 2 warmup
Self-reported — not verified by launch80Decode holds 142 t/s weighted across eight categories. The concurrency knee sits at 8 requests, past which aggregate throughput flattens while p99 TTFT nearly doubles. Prefill scales cleanly to 32K and falls off at 64K.
- Combined decode142.4t/sweighted across categories
- TTFT p50304ms
- TTFT p99512ms
- ITL 1% low96.1t/sthe stutter floor
Read these first
- ITL 99% high on `code` and `reasoning` rests on 20 runs; BetterBench flagged both as under-sampled. Treat those tails as indicative, not settled.
- Single host, single run. No paired A/B, so a sub-3% difference against another result on different hardware is not meaningful.
| category | passes | TTFT p50 | TTFT p99 | PP t/s med | ITL 1% low | ITL med | ITL 99% high | decode med | ±IQR |
|---|---|---|---|---|---|---|---|---|---|
| code† | 20 | 312 | 498 | 4210 | 96.1 | 151.0 | 203.4 | 148.2 | 6.1 |
| prose | 20 | 298 | 471 | 3980 | 101.3 | 142.8 | 188.0 | 139.7 | 4.4 |
| reasoning† | 20 | 341 | 602 | 4055 | 88.4 | 138.2 | 191.7 | 136.1 | 8.9 |
| summarization | 20 | 489 | 744 | 5120 | 99.0 | 145.6 | 196.2 | 141.9 | 5.2 |
† a percentile on this row rests on too few samples to be reliable.
| level | aggregate t/s | TTFT p50 | TTFT p99 | per-req decode med |
|---|---|---|---|---|
| 1 | 142.4 | 304 | 512 | 142.4 |
| 4 | 421.6 | 486 | 918 | 105.4 |
| 8 | 612.8 | 918 | 1840 | 76.6 |
| 16 | 634.1 | 1974 | 4120 | 39.6 |
| target depth | PP median | TTFT p50 |
|---|---|---|
| 2K | 3840 | 0.52 |
| 8K | 4210 | 1.94 |
| 32K | 4165 | 7.88 |
| 64K | 3102 | 21.14 |
Setup
One RTX 4090 on an otherwise idle host, vLLM 0.9.3 in the official container, mxfp4 weights with an fp8 KV cache.
The BetterBench v1 corpus was used unmodified. Nonce prefixes were left on, so prefill numbers reflect a cold prefix cache rather than accidental cache hits.
Reading the concurrency table
Aggregate throughput gains almost nothing between 8 and 16 concurrent requests, while p99 TTFT more than doubles. For an interactive workload the useful ceiling on this box is 8.
- 1 to 4: near-linear aggregate gain, modest latency cost
- 4 to 8: 45% more aggregate throughput, p99 TTFT doubles
- 8 to 16: 3% more throughput, p99 TTFT doubles again