OpenRouter is a marketplace: it takes a cut of the inference dollars flowing through it. So follow the dollars — not the tokens. Here's the GMV, OpenRouter's likely cut, and the profit pool per model maker.
OpenRouter doesn't sell tokens — it sells convenience: one API, one bill, automatic fallback and routing across 400+ models. It monetises with a take-rate on the dollars routed (credit fees + BYOK fee + provider spread). Applied to ~$1.1B of GMV:
This is why OpenRouter's revenue tracks dollar spend, not token volume. The Chinese open-weight wave drives most tokens but little spend — so OpenRouter's P&L rides the premium, Western, high-$/token traffic. A pure race-to-cheap would compress its GMV-per-token and squeeze the take.
| vendor | revenue / yr | gross profit @ 90% | spend share |
| anthropic | $653M | $588M | 59% |
| openai | $126M | $113M | 11% |
| $82M | $74M | 7% | |
| deepseek | $82M | $74M | 7% |
| minimax | $30M | $27M | 3% |
| qwen | $30M | $27M | 3% |
| xiaomi | $30M | $27M | 3% |
| z-ai | $29M | $26M | 3% |
| moonshotai | $17M | $15M | 2% |
| tencent | $11M | $10M | 1% |
Revenue = tokens × list price (exact). Gross profit assumes a 90% token margin (inference cost ~10% of price). This is only the OpenRouter slice of each vendor's business — direct + cloud revenue is far larger.
Take one release — say Opus 4.5. It cost (assume) $100M to train; here's what each Opus has generated. "Via OpenRouter" is measured; "est. total" scales that up ×10 (OpenRouter is only ~10% of Anthropic's API).
| model | rev via OR (measured) | est. total rev (×10) | gross profit @ 90% | train cost | return |
Return = est. gross profit ÷ training cost. Even on the OpenRouter slice alone, the bigger Opus releases out-earn a $$100M-class training run within a few months; scaled to full API, each Opus returns several × its compute cost. The model isn't the cost centre — it's the asset. Training spend is dwarfed by the revenue an in-demand model throws off across its ~2-month prime.
Caveats: training cost and the ×10 scale-up are assumptions (Anthropic doesn't disclose either); "cost" here is the compute run only, excluding R&D, staff, and serving infra. The newest release is still ramping — its lifetime return looks low only because it has had weeks, not months, to accumulate. Treat as an order-of-magnitude frame, not Anthropic's books.
tokens × price/tokenrevenue × 90% (margin = 1 − inference_cost/price)GMV × take-rate
Worked example — Anthropic: at the current run-rate it routes $653M/yr of revenue through OpenRouter, ~$588M/yr gross profit at a 90% margin. Because Anthropic is ~59% of all OpenRouter spend, it alone underwrites the bulk of OpenRouter's take.
Sensitivity: margin scales profit linearly (80% → multiply by 0.8/0.9); take-rate scales OpenRouter's cut linearly. The fragile variable is blended $/M — if it falls from $0.60 toward the open-weight floor (~$0.20), GMV and every downstream number compress with it.