The numbers, with methodology

We publish our own measurements because vendor tables don't tell you what a $0.53/h endpoint actually does at 1M context. All runs: llama.cpp fork, Qwen3.8-27B, single server, no other tenants. Reproduce anything — configs below.

SKUQuantContexttok/s (decode)TTFT$/hr (us)
SoloFP8 + Q5_K KV1,000,00047$1.14
FastNVFP4200,000 ×2 · 400K solo120$2.12
TeamNVFP4260,000 ×13 · 500K ×5 · 1M ×2160$4.27
LiteQ5_K260,00030$1.48
Micromixed 2-bit200,00020$0.60

Reference points (community, same model): 218 tok/s single-stream (short ctx) · 140–260 tok/s p50, 0.156s TTFT. TODO at launch: fill TTFT column + link to raw benchmark logs.

Measured capacity per card (Qwen3.8-27B, llama.cpp)

CardCapacityVision (mmproj)
Blackwell 24 GB1 user @ 260Ksecond model on CPU
Blackwell 32 GB2 users @ 200K each · or 1 @ 400Knative
Blackwell 48 GB5 users @ 200K each · or 2 @ 500K · or 1 @ 1Mnative
Blackwell 96 GB13 users @ 260K each · or 5 @ 500K · or 2 @ 1Mnative
Ada (any size)max 1 user per card · 1M ctx on 48 GB has slow prefill
Ampere (any size)max 1 user per card · 500K ctx max recommended (very slow prefill)

These are the numbers behind the SKU capacity column on the pricing page. Concurrent users = concurrent in-flight generations (max-num-seqs).

Quality (vendor-reported, Qwen model card)

BenchmarkQwen3.6-27BQwen3.8-27B
Terminal-Bench 2.163.473.0
DeepSWE 1.113.342.2
OSWorld-Verified63.984.3
SWE-bench Pro61.7
LiveCodeBench v690.3

Artificial Analysis Intelligence Index: 52 (GLM-5.2: 53, Kimi K3: higher but cluster-scale). #9 on Code Arena WebDev — the only small model in the top 10. Scores are Alibaba's; we flag that instead of hiding it.

Cost vs per-token APIs (agentic workload)

10 turns/hr × 200K context + 2K output:

Option$/hr
noTOKEN.cloud Solo (dedicated)$1.14
GLM-5.2 ($1.40/$4.40 per M)$3.02
Kimi K3 ($3/$15 per M)$6.75

And this is a light workload — 10 turns an hour. At 100 turns/hr × 100K context (10M input tokens/hr), GLM-5.2 reads $14.88/hr and Kimi K3 reads $33/hr, while the dedicated endpoint stays at $1.14. Per-token pricing bills every context re-read; dedicated pricing doesn't care how hard your agent works.