We publish our own measurements because vendor tables don't tell you what a $0.53/h endpoint actually does at 1M context. All runs: llama.cpp fork, Qwen3.8-27B, single server, no other tenants. Reproduce anything — configs below.
| SKU | Quant | Context | tok/s (decode) | TTFT | $/hr (us) |
|---|---|---|---|---|---|
| Solo | FP8 + Q5_K KV | 1,000,000 | 47 | — | $1.14 |
| Fast | NVFP4 | 200,000 ×2 · 400K solo | 120 | — | $2.12 |
| Team | NVFP4 | 260,000 ×13 · 500K ×5 · 1M ×2 | 160 | — | $4.27 |
| Lite | Q5_K | 260,000 | 30 | — | $1.48 |
| Micro | mixed 2-bit | 200,000 | 20 | — | $0.60 |
Reference points (community, same model): 218 tok/s single-stream (short ctx) · 140–260 tok/s p50, 0.156s TTFT. TODO at launch: fill TTFT column + link to raw benchmark logs.
| Card | Capacity | Vision (mmproj) |
|---|---|---|
| Blackwell 24 GB | 1 user @ 260K | second model on CPU |
| Blackwell 32 GB | 2 users @ 200K each · or 1 @ 400K | native |
| Blackwell 48 GB | 5 users @ 200K each · or 2 @ 500K · or 1 @ 1M | native |
| Blackwell 96 GB | 13 users @ 260K each · or 5 @ 500K · or 2 @ 1M | native |
| Ada (any size) | max 1 user per card · 1M ctx on 48 GB has slow prefill | — |
| Ampere (any size) | max 1 user per card · 500K ctx max recommended (very slow prefill) | — |
These are the numbers behind the SKU capacity column on the pricing page. Concurrent users = concurrent in-flight generations (max-num-seqs).
| Benchmark | Qwen3.6-27B | Qwen3.8-27B |
|---|---|---|
| Terminal-Bench 2.1 | 63.4 | 73.0 |
| DeepSWE 1.1 | 13.3 | 42.2 |
| OSWorld-Verified | 63.9 | 84.3 |
| SWE-bench Pro | — | 61.7 |
| LiveCodeBench v6 | — | 90.3 |
Artificial Analysis Intelligence Index: 52 (GLM-5.2: 53, Kimi K3: higher but cluster-scale). #9 on Code Arena WebDev — the only small model in the top 10. Scores are Alibaba's; we flag that instead of hiding it.
10 turns/hr × 200K context + 2K output:
| Option | $/hr |
|---|---|
| noTOKEN.cloud Solo (dedicated) | $1.14 |
| GLM-5.2 ($1.40/$4.40 per M) | $3.02 |
| Kimi K3 ($3/$15 per M) | $6.75 |
And this is a light workload — 10 turns an hour. At 100 turns/hr × 100K context (10M input tokens/hr), GLM-5.2 reads $14.88/hr and Kimi K3 reads $33/hr, while the dedicated endpoint stays at $1.14. Per-token pricing bills every context re-read; dedicated pricing doesn't care how hard your agent works.