CodeSOTAIntelligence
Research Memo · ort.fabryka.ai

The OpenRouter LLM Market:
Who Owns Volume vs. Who Owns Revenue

Chinese open-weight labs have captured the majority of token volume on OpenRouter — but a single Western lab still captures the majority of the dollars. The gap between those two facts is the whole investment thesis.

635 days of data (2024-11-09 → 2026-08-05) · measured throughput ~36T tokens/wk · generated 2026-08-05
Thesis — two clocks. Volume and revenue are not rival metrics; they are the leading and lagging indicators of the same migration. Today Anthropic holds ~13% of token volume but ~65% of spend while DeepSeek holds ~22% of volume but only ~5% of spend. The bear case isn't "ignore volume" — it's the opposite: a good-enough model at ~83× lower list price per output token is taking volume inside the same apps that pay for Claude. That is classic low-end disruption, where volume moves first and revenue follows with a lag. The revenue lead is real but is the clock that moves last — so the volume curve is the one to underwrite.
~13% / 65%
Anthropic — volume / spend share
~22% / 5%
DeepSeek — volume / spend share
~36T tokens/wk
measured throughput
635d
daily history captured

1 · Vendor share — the two leaderboards

By token volume (30d)

vendorvolΔ%shareΔpp
deepseek44.7T+30%22.0%+2.7pp
xiaomi36.3T+88%17.9%+7.0pp
anthropic25.6T-10%12.6%-3.2pp
google18.1T+7%8.9%-0.5pp
z-ai16.1T+70%7.9%+2.6pp
openai15.8T+36%7.8%+1.3pp
minimax13.4T-26%6.6%-3.5pp
tencent11.2T-26%5.5%-3.0pp

By estimated spend (7d)

vendor$Δ%shareΔpp
anthropic$24.4M-2%65.2%-1.9pp
openai$4.5M-1%12.0%-0.3pp
google$3.0M+1%8.1%-0.1pp
deepseek$1.8M+33%4.8%+1.1pp
minimax$1.0M+388%2.7%+2.1pp
qwen$581.4K-1%1.6%-0.0pp
xiaomi$576.0K-33%1.5%-0.8pp
z-ai$563.3K-11%1.5%-0.2pp

Δpp = percentage-point shift in share of the total — the zero-sum view of who's taking ground.

2 · Substitution in the wild — the same apps, both vendors

The clearest evidence that volume migration threatens revenue: the highest-volume apps run both Anthropic and DeepSeek for the same job. 10 of the top 10 apps mix them — and DeepSeek V4 Flash undercuts Claude Sonnet by ~83× on output price. When a workflow already calls both, switching share is a config change, not a migration cost.

appAnthropicDeepSeekmix
Hermes Agent3%31%both
Kilo Code2%4%both
Claude Code36%7%both
OpenClaw12%22%both
Cline8%23%both
pi18%34%both
Descript87%0%both
ISEKAI ZERO1%59%both
Janitor AI1%65%both
Lemonade2%0%both

Share of the app's tokens by vendor. Even Anthropic's own Claude Code routes ~13% to DeepSeek; roleplay/agent apps lean majority-DeepSeek. The cheaper model is already inside the funnel.

3 · The substitution engine — families, not versions

Individual versions churn constantly, which makes per-model churn misleading. The real unit is the family: usage migrates within a family (e.g. Claude Opus 4.6 → 4.7 → 4.8) while the family's total tells you whether the franchise is winning. "Lead version" share is a stickiness signal — how fast users consolidate onto the newest release.

familyrecentΔ%lead version (stickiness)migration
Xiaomi MiMo33.8T+98%MiMo-V2.5 100%
DeepSeek DeepSeek Flash29.6T+43%DeepSeek V4 Flash 04 88%
Z.ai GLM15.8T+75%GLM 5.2 89%v5.1→v5.2
DeepSeek DeepSeek Pro12.8T+36%DeepSeek V4 Pro 100%
Anthropic Claude Opus12.7T-32%Claude Opus 4.8 49%
MiniMax MiniMax M312.6T-25%MiniMax M3 100%
Tencent Hy311.2T-26%Hy3 95%v—→v—
nvidia/nemotron ultra 550b a5510.6T+219%nvidia/nemotron-3-ul 100%
Google Gemini Flash8.8T+5%Gemini 3 Flash Previ 47%v3.5→v3.6
Anthropic Claude Sonnet8.8T+9%Claude Sonnet 5 50%v4.6→v5
StepFun Step Flash5.9T+18%Step 3.7 Flash 100%v3.5→v3.7
Worked example — Claude Opus: the family grew while usage rotated off 4.6 (declining) onto 4.7 (surging) with 4.8 ramping behind it. Healthy in-family substitution = retention, not churn.

4 · Lifecycle of the ecosystem

Every tracked model placed on its curve — launch → ramp → peak → decline → death:

18
LAUNCHING
young & climbing
6
GROWING
rising toward peak
15
MATURE
at peak, plateaued
108
DECLINING
past peak, falling
32
DEAD
retired / substituted

A long "declining" tail is normal — it's last-generation versions bleeding into their successors. "Dead" = had real volume, now ~zero with no in-family heir absorbing it.

5 · Winners & losers (family churn, 30d)

Gaining

  • nvidia/nemotron ultra 550b +219%
  • Xiaomi MiMo +98%
  • Z.ai GLM +75%
  • DeepSeek DeepSeek Flash +43%
  • DeepSeek DeepSeek Pro +36%

Losing

  • Anthropic Claude Opus -32%
  • Tencent Hy3 -26%
  • MiniMax MiniMax M3 -25%
  • Google Gemini Flash +5%
  • Anthropic Claude Sonnet +9%

6 · The model treadmill — and China churns fastest

Models don't age gracefully; they get cannibalised by their own successors. Measuring each generation's useful life (launch → usage falling below half its peak) exposes how fast the treadmill runs — and it runs much faster in the East. Average useful life: China ~54 days vs West ~100 days.

generationpeak shareuseful lifeorigin
Deepseek 3.15.9%10dChina
Qwen 2.513.0%11dChina
Deepseek 2.55.8%14dChina
Deepseek 39.4%18dChina
Google 1.546.7%22dWest
Deepseek 14.2%25dChina
Minimax 2.526.2%26dChina
Google 3.51.8%28dWest
Minimax 24.5%38dChina
Anthropic 4.79.2%46dWest
Minimax 310.7%49dChina
Minimax 2.15.9%53dChina

Shortest-lived first. The fastest Chinese generations turn over in under two weeks, while Western flagships hold for two to four months — the table shows the full spread. On average a Chinese generation's useful life is ~54d vs ~100d in the West. Faster churn = faster iteration, but a brutal amortisation window on training spend.

7 · Implications