CodeSOTAIntelligence
Research Memo · ort.fabryka.ai

The OpenRouter LLM Market:
Who Owns Volume vs. Who Owns Revenue

Chinese open-weight labs have captured the majority of token volume on OpenRouter — but a single Western lab still captures the majority of the dollars. The gap between those two facts is the whole investment thesis.

686 days of data (2024-11-09 → 2026-09-25) · measured throughput ~36T tokens/wk · generated 2026-09-25
Thesis — two clocks. Volume and revenue are not rival metrics; they are the leading and lagging indicators of the same migration. Today Anthropic holds ~5% of token volume but ~55% of spend while DeepSeek holds ~24% of volume but only ~9% of spend. The bear case isn't "ignore volume" — it's the opposite: a good-enough model at ~77× lower list price per output token is taking volume inside the same apps that pay for Claude. That is classic low-end disruption, where volume moves first and revenue follows with a lag. The revenue lead is real but is the clock that moves last — so the volume curve is the one to underwrite.
~5% / 55%
Anthropic — volume / spend share
~24% / 9%
DeepSeek — volume / spend share
~36T tokens/wk
measured throughput
686d
daily history captured

1 · Vendor share — the two leaderboards

By token volume (30d)

vendorvolΔ%shareΔpp
deepseek110.9T+44%24.0%-4.3pp
z-ai77.6T+353%16.8%+10.5pp
openai74.4T+129%16.1%+4.1pp
tencent73.0T+120%15.8%+3.6pp
xiaomi29.2T-5%6.3%-5.0pp
google28.8T+13%6.2%-3.1pp
anthropic21.1T-4%4.6%-3.5pp
moonshotai8.1T+7%1.8%-1.0pp

By estimated spend (7d)

vendor$Δ%shareΔpp
anthropic$12.6M-1%54.8%-2.8pp
openai$2.4M-1%10.5%-0.5pp
deepseek$2.2M+36%9.4%+2.2pp
google$1.6M+1%6.9%-0.2pp
minimax$1.0M+322%4.6%+3.5pp
tencent$616.0K+11%2.7%+0.2pp
qwen$579.3K-1%2.5%-0.1pp
xiaomi$576.0K-33%2.5%-1.4pp

Δpp = percentage-point shift in share of the total — the zero-sum view of who's taking ground.

2 · Substitution in the wild — the same apps, both vendors

The clearest evidence that volume migration threatens revenue: the highest-volume apps run both Anthropic and DeepSeek for the same job. 6 of the top 10 apps mix them — and DeepSeek V4 Flash undercuts Claude Sonnet by ~77× on output price. When a workflow already calls both, switching share is a config change, not a migration cost.

appAnthropicDeepSeekmix
Hermes Agent2%45%both
Claude Code14%9%both
Kilo Code0%7%
Cline2%31%both
pi3%44%both
Codex0%7%
OpenClaw8%42%both
OpenHands5%71%both
ISEKAI ZERO0%22%
Portkey AI0%6%

Share of the app's tokens by vendor. Even Anthropic's own Claude Code routes ~13% to DeepSeek; roleplay/agent apps lean majority-DeepSeek. The cheaper model is already inside the funnel.

3 · The substitution engine — families, not versions

Individual versions churn constantly, which makes per-model churn misleading. The real unit is the family: usage migrates within a family (e.g. Claude Opus 4.6 → 4.7 → 4.8) while the family's total tells you whether the franchise is winning. "Lead version" share is a stickiness signal — how fast users consolidate onto the newest release.

familyrecentΔ%lead version (stickiness)migration
DeepSeek DeepSeek Flash66.0T+8%DeepSeek V4 Flash 07 71%vV4 0423→vV4 0731
OpenAI GPT Luna (batch)53.5T+198%GPT-5.6 Luna (batch) 98%—
Xiaomi MiMo25.6T-10%MiMo-V2.5 100%—
Tencent Hy319.0T-43%Hy3 100%—
nvidia/nemotron ultra 550b a5517.9T+34%nvidia/nemotron-3-ul 100%—
Google Gemini Flash (batch)16.1T+30%Gemini 3.8 Flash (ba 45%v3.6→v3.8
Z.ai GLM (batch)11.5TnewGLM 5.3 (batch) 100%—
DeepSeek DeepSeek Pro9.9T-24%DeepSeek V4 Pro 0423 53%vV4 0423→vV4 0813
Z.ai GLM9.4T-41%GLM 5.2 (free) 93%v5.2→v4.6
Worked example — Claude Opus: the family grew while usage rotated off 4.6 (declining) onto 4.7 (surging) with 4.8 ramping behind it. Healthy in-family substitution = retention, not churn.

4 · Lifecycle of the ecosystem

Every tracked model placed on its curve — launch → ramp → peak → decline → death:

25
LAUNCHING
young & climbing
7
GROWING
rising toward peak
15
MATURE
at peak, plateaued
129
DECLINING
past peak, falling
39
DEAD
retired / substituted

A long "declining" tail is normal — it's last-generation versions bleeding into their successors. "Dead" = had real volume, now ~zero with no in-family heir absorbing it.

5 · Winners & losers (family churn, 30d)

Gaining

  • Z.ai GLM (batch) new
  • OpenAI GPT Luna (batch) +198%
  • nvidia/nemotron ultra 550b +34%
  • Google Gemini Flash (batch +30%
  • DeepSeek DeepSeek Flash +8%

Losing

  • Tencent Hy3 -43%
  • Z.ai GLM -41%
  • DeepSeek DeepSeek Pro -24%
  • Xiaomi MiMo -10%
  • DeepSeek DeepSeek Flash +8%

6 · The model treadmill — and China churns fastest

Models don't age gracefully; they get cannibalised by their own successors. Measuring each generation's useful life (launch → usage falling below half its peak) exposes how fast the treadmill runs — and it runs much faster in the East. Average useful life: China ~63 days vs West ~92 days.

generationpeak shareuseful lifeorigin
Deepseek 3.15.9%10dChina
Qwen 2.513.0%11dChina
Deepseek 2.55.8%14dChina
Deepseek 39.4%18dChina
Google 3.73.5%22dWest
Minimax 2.526.2%26dChina
Google 3.51.8%28dWest
Google 3.63.6%30dWest
Minimax 24.5%38dChina
Anthropic 4.79.2%46dWest
Minimax 310.7%49dChina
Minimax 2.15.9%53dChina

Shortest-lived first. The fastest Chinese generations turn over in under two weeks, while Western flagships hold for two to four months — the table shows the full spread. On average a Chinese generation's useful life is ~63d vs ~92d in the West. Faster churn = faster iteration, but a brutal amortisation window on training spend.

7 · Implications