mlx-community/Qwen3.8-27B-oQ6
Self-reported — not verified by launch80- Combined decode7.3t/sweighted across categories
- Combined ITL 1% low2.5t/sthe trustworthy tail metric
- Combined TTFT p501,641mssingle-stream, batch = 1
Read these first
- Quick smoke run: 5 passes per category after 1 warmup. Too few passes for any percentile to hold; treat every figure as indicative, not as a result.
- Under-sampled: TTFT p99 in chat (n≤5). A p99 needs 500 samples; the daggered cells are roughly the worst observed, not percentiles.
- 3 of 5 runs stopped at max_tokens (60%). On a thinking model a truncated run measures the thinking phase, not a complete answer.
Hover a bar for its IQR and coefficient of variation — a high CV means the category's passes disagree, so read small differences there with care.
Full numbers
| category | passes | TTFT p50 | TTFT p99 | ITL 1% low | ITL med | ITL 99% high | decode med | ±IQR | CV |
|---|---|---|---|---|---|---|---|---|---|
| chat | 5 | 1,641.2 | 1,751.0† | 2.5 | 73.5 | 17,376.7 | 7.3 | 0.7 | 7.3% |
Generated by BetterBench 0.4.0 from a corpus v1.0 run. Corpus hash 706e6cf4b5165492. Combined-score weights — code 0.3, reasoning 0.2, prose 0.15, json 0.15, file_edit 0.1, summarization 0.1. Results are only comparable within a corpus version. Stopped at max_tokens: 3/5 runs (60%) — on a thinking model a truncated run measures the thinking phase, not a complete answer. 1 percentile is marked † — it rests on fewer samples than n · tail ≥ 5 requires (a p99 needs 500 observations), so read it as roughly the worst observed rather than as a percentile. The full list is under sample_gate in the results JSON. See METHODOLOGY.md §sample-size.