Chinese open-weight labs have captured the majority of token volume on OpenRouter — but a single Western lab still captures the majority of the dollars. The gap between those two facts is the whole investment thesis.
| vendor | vol | Δ% | share | Δpp |
| deepseek | 110.9T | +44% | 24.0% | -4.3pp |
| z-ai | 77.6T | +353% | 16.8% | +10.5pp |
| openai | 74.4T | +129% | 16.1% | +4.1pp |
| tencent | 73.0T | +120% | 15.8% | +3.6pp |
| xiaomi | 29.2T | -5% | 6.3% | -5.0pp |
| 28.8T | +13% | 6.2% | -3.1pp | |
| anthropic | 21.1T | -4% | 4.6% | -3.5pp |
| moonshotai | 8.1T | +7% | 1.8% | -1.0pp |
| vendor | $ | Δ% | share | Δpp |
| anthropic | $12.6M | -1% | 54.8% | -2.8pp |
| openai | $2.4M | -1% | 10.5% | -0.5pp |
| deepseek | $2.2M | +36% | 9.4% | +2.2pp |
| $1.6M | +1% | 6.9% | -0.2pp | |
| minimax | $1.0M | +322% | 4.6% | +3.5pp |
| tencent | $616.0K | +11% | 2.7% | +0.2pp |
| qwen | $579.3K | -1% | 2.5% | -0.1pp |
| xiaomi | $576.0K | -33% | 2.5% | -1.4pp |
Δpp = percentage-point shift in share of the total — the zero-sum view of who's taking ground.
The clearest evidence that volume migration threatens revenue: the highest-volume apps run both Anthropic and DeepSeek for the same job. 6 of the top 10 apps mix them — and DeepSeek V4 Flash undercuts Claude Sonnet by ~77× on output price. When a workflow already calls both, switching share is a config change, not a migration cost.
| app | Anthropic | DeepSeek | mix |
| Hermes Agent | 2% | 45% | both |
| Claude Code | 14% | 9% | both |
| Kilo Code | 0% | 7% | |
| Cline | 2% | 31% | both |
| pi | 3% | 44% | both |
| Codex | 0% | 7% | |
| OpenClaw | 8% | 42% | both |
| OpenHands | 5% | 71% | both |
| ISEKAI ZERO | 0% | 22% | |
| Portkey AI | 0% | 6% |
Share of the app's tokens by vendor. Even Anthropic's own Claude Code routes ~13% to DeepSeek; roleplay/agent apps lean majority-DeepSeek. The cheaper model is already inside the funnel.
Individual versions churn constantly, which makes per-model churn misleading. The real unit is the family: usage migrates within a family (e.g. Claude Opus 4.6 → 4.7 → 4.8) while the family's total tells you whether the franchise is winning. "Lead version" share is a stickiness signal — how fast users consolidate onto the newest release.
| family | recent | Δ% | lead version (stickiness) | migration |
| DeepSeek DeepSeek Flash | 66.0T | +8% | DeepSeek V4 Flash 07 71% | vV4 0423→vV4 0731 |
| OpenAI GPT Luna (batch) | 53.5T | +198% | GPT-5.6 Luna (batch) 98% | — |
| Xiaomi MiMo | 25.6T | -10% | MiMo-V2.5 100% | — |
| Tencent Hy3 | 19.0T | -43% | Hy3 100% | — |
| nvidia/nemotron ultra 550b a55 | 17.9T | +34% | nvidia/nemotron-3-ul 100% | — |
| Google Gemini Flash (batch) | 16.1T | +30% | Gemini 3.8 Flash (ba 45% | v3.6→v3.8 |
| Z.ai GLM (batch) | 11.5T | new | GLM 5.3 (batch) 100% | — |
| DeepSeek DeepSeek Pro | 9.9T | -24% | DeepSeek V4 Pro 0423 53% | vV4 0423→vV4 0813 |
| Z.ai GLM | 9.4T | -41% | GLM 5.2 (free) 93% | v5.2→v4.6 |
Every tracked model placed on its curve — launch → ramp → peak → decline → death:
A long "declining" tail is normal — it's last-generation versions bleeding into their successors. "Dead" = had real volume, now ~zero with no in-family heir absorbing it.
Models don't age gracefully; they get cannibalised by their own successors. Measuring each generation's useful life (launch → usage falling below half its peak) exposes how fast the treadmill runs — and it runs much faster in the East. Average useful life: China ~63 days vs West ~92 days.
| generation | peak share | useful life | origin |
| Deepseek 3.1 | 5.9% | 10d | China |
| Qwen 2.5 | 13.0% | 11d | China |
| Deepseek 2.5 | 5.8% | 14d | China |
| Deepseek 3 | 9.4% | 18d | China |
| Google 3.7 | 3.5% | 22d | West |
| Minimax 2.5 | 26.2% | 26d | China |
| Google 3.5 | 1.8% | 28d | West |
| Google 3.6 | 3.6% | 30d | West |
| Minimax 2 | 4.5% | 38d | China |
| Anthropic 4.7 | 9.2% | 46d | West |
| Minimax 3 | 10.7% | 49d | China |
| Minimax 2.1 | 5.9% | 53d | China |
Shortest-lived first. The fastest Chinese generations turn over in under two weeks, while Western flagships hold for two to four months — the table shows the full spread. On average a Chinese generation's useful life is ~63d vs ~92d in the West. Faster churn = faster iteration, but a brutal amortisation window on training spend.