Same model, same token: one euro buys one block of tokens at API list price and more than a hundred blocks on a Claude Max 20x seat.

AI Analysis — the Smartness × Cost Map and the Model Dossier

Posted by:

|

On:

|

Same model, same token: one euro buys one block of tokens at API list price and more than a hundred blocks on a Claude Max 20x seat.

Every model argument you have ever read is about one axis: which one is smartest. Your invoice is about the other axis. And the units people quote in the first argument are almost never the units they actually pay in the second.

So we drew both axes. This page is the live output of two panels inside TinkerClaw: a scatter plot placing every model we can reach by intelligence against what a token actually costs us, and a dossier grading each on what it is good at — and what it will refuse to answer. Not screenshots: the real markup, zoomable and hoverable. Our agent republishes this page itself, every night. There is no second copy anywhere.

As of 21 September 2026: 121 models across 13 vendors on the chart, 45 graded in the dossier across 10 capability columns. 58 of 121 models are discounted by a plan we hold, widest gap 510×.

THALAMUS frontier 16 rungs (€/task) of 169 across 48/121 models · bias 3 (balanced) floor idx 48.4 → pick claude-code/claude-fable-5-1@medium idx 48.9 €0.129/task

1. SMARTNESS × COST — the map

Every mark is an SVG node, not a pixel. Click the map to open it full screen. There you can scroll or pinch to zoom, drag to pan, hover any dot for its full costing, click a vendor chip to isolate that vendor, and flip the two switches at top right: per token or per task, and a log or a linear cost axis. The middle of the plot is deliberately crowded; zooming is how you read it.

click the map to open it full screen · zoom, pan, hover every model · switch per token / per task and log / linear
SMARTNESS × COST
AA intelligence index · effective cost · a stop is plotted when the vendor documents it · solid = AA measured it · dotted = estimated from other public benchmarks (±1σ on hover) · dashed = neither, hung at the headline

€0.05€0.10€0.20€0.50€1.0€2.0€5.0€10€20€5010152025303540455055EFFECTIVE COST · €/Mtok OUTPUT · logEFFECTIVE COST PER AVERAGE TASK · €/task · logAA INTELLIGENCE INDEXSIZE ∝ CONTEXT128k256k1M75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%ling-3.0-flash · Effort ladder unknown — €0.06/Mtok · ~9.6k tok/task · 262k ctx · idx 24.9 (headline — AA published no per-effort split)inkling · Effort ladder unknown — €4.0/Mtok · ~9.6k tok/task · 200k ctx · idx 25.0 (headline — AA published no per-effort split)inkling-small · Effort ladder unknown — €1.2/Mtok · ~9.6k tok/task · 200k ctx · idx 27.8 (headline — AA published no per-effort split)nex-n2-pro · Effort ladder unknown — €1.0/Mtok · ~9.6k tok/task · 200k ctx · idx 28.2 (headline — AA published no per-effort split)75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%gpt-4o (openai)gpt-4o (copilot)gemini-2.0-flashgpt-4.1 (openai)gpt-4.1 (copilot)nemotron-3.5-lightninggemini-2.5-flashgrok-code-fast-1gemini-2.5-pro (google)gemini-2.5-pro (copilot)claude-haiku-4.5 (copilot)Claude Haiku 4.5 (claude-code)longcat-2.0o3gpt-5.1-codex-minigpt-5-minigpt-5.4-nanogemini-3.5-flash-litegpt-5gpt-5.1-codexgpt-5.4-mini (openai)gpt-5.4-mini (copilot)gpt-5.1 (openai)gpt-5.1 (copilot)ling-3.0-flashinklingHunyuan 3 (openrouter·tencent)Hunyuan 3 (tencent)Kimi K2.7 Code (openrouter·moonshotai)Kimi K2.7 Code (moonshot-ai)mimo-v2.5-pro (openrouter·xiaomi)mimo-v2.5-pro (xiaomi)GLM 5.1 (openrouter·z-ai)GLM 5.1 (z-ai)gemini-3-flash-preview (google)gemini-3-flash-preview (copilot)Kimi K2.6 (openrouter·moonshotai)Kimi K2.6 (moonshot-ai)qwen3.6-plusinkling-smallglm-5 (openrouter·z-ai)glm-5 (z-ai)gemini-3-pro-preview (google)gemini-3-pro-preview (copilot)solar-pro4nex-n2-proQwen 3.6 Max Previewgpt-5.2-codexclaude-opus-4.5MiniMax M3 (openrouter·minimax)MiniMax M3 (minimax)claude-sonnet-4.5Qwen 3.7 Maxgemini-3.1-pro-preview (google)gemini-3.1-pro-preview (copilot)Claude Sonnet 4.6 (claude-code)claude-sonnet-4.6 (copilot)gpt-5.2 (openai)gpt-5.2 (copilot)claude-opus-4.6 (copilot)Claude Opus 4.6 (claude-code)gpt-5.3-codex (openai)gpt-5.3-codex (copilot)gemini-3.5-flashQwen3.8 27B (openrouter·qwen)Qwen3.8 27B (darkbloom)Qwen3.8 27B (alibaba)GLM 5.2 (openrouter·z-ai)GLM 5.2 (deepinfra)GLM 5.2 (z-ai)Muse Spark 1.1gemini-3.6-flashdeepseek-v4-flashDeepSeek V4 Flashgpt-5.4-progpt-5.2-progpt-5.1-codex-maxDeepSeek V4 Flash Vision (openrouter·deepseek)DeepSeek V4 Flash Vision (deepinfra)claude-sonnet-4DeepSeek V4 Pro (openrouter·deepseek)DeepSeek V4 Pro (deepseek)deepseek-v4-progpt-5.6-luna (codex · openai · openai-codex)claude-sonnet-5gpt-5.5 (openai)gpt-5.5 (copilot)gpt-5.5 (openai-codex)Grok 4.5gpt-5.4 (openai)gpt-5.4 (copilot)DeepSeek V4.1 Flash (openrouter·deepseek)DeepSeek V4.1 Flash (deepseek)Muse Spark 1.2gemini-3.7-flashQwen3.8 2.4T A95BClaude Opus 4.7 (claude-code)claude-opus-4.7 (copilot)gemini-3.8-flashClaude Opus 4.8GLM 5.3 Flash (openrouter·z-ai)GLM 5.3 Flash (z-ai)gpt-5.6-terra (codex · openai · openai-codex)Kimi K3 (openrouter·moonshotai)Kimi K3 (moonshot-ai)Grok 4.6GLM 5.3 (openrouter·z-ai)GLM 5.3 (z-ai)Qwen 3.8 Max (0902)gpt-5.6-sol (codex · openai · openai-codex)Muse Spark 1.3Claude Opus 5GPT-6 Astra (openai-codex)gpt-6-astra (openai)Claude Fable 5.1THALAMUS frontier 16 rungs (€/task) of 169 across 48/121 models · bias 3 (balanced) floor idx 48.4 → pick claude-code/claude-fable-5-1@medium idx 48.9 €0.129/taskThe envelope is the Pareto frontier, on the €/task axis, of the effort rungs of the models src/shared/thalamus-candidates.ts can reach, computed by src/shared/thalamus-frontier.ts — the same module the reply-path router calls, never a list in this file. A rung is on the frontier when no other rung is both cheaper-or-equal per task AND smarter-or-equal; the BIAS dial picks the cheapest frontier rung within its gap of the best. A model with NO ring is either unreachable (no AA index, dearer than 1.5x the anchor’s effective cost, or its provider’s token window is spent) or reachable but dominated on every rung. ‘cost UNVERIFIED’ means at least one reachable model carried no published price, so the veto could prove it neither cheap nor expensive.THALAMUS frontier 16 rungs (€/task) of 169 across 48/121 models · bias 3 (balanced) floor idx 48.4 → pick claude-code/claude-fable-5-1@medium idx 48.9 €0.129/task · no domain switches at this bias · cost UNVERIFIED · out 73 cost-veto · upper bound: no quota snapshot, no auth filterno domain switches at this bias · cost UNVERIFIED · out 73 cost-veto · upper bound: no quota snapshot, no auth filterThe envelope is the Pareto frontier, on the €/task axis, of the effort rungs of the models src/shared/thalamus-candidates.ts can reach, computed by src/shared/thalamus-frontier.ts — the same module the reply-path router calls, never a list in this file. A rung is on the frontier when no other rung is both cheaper-or-equal per task AND smarter-or-equal; the BIAS dial picks the cheapest frontier rung within its gap of the best. A model with NO ring is either unreachable (no AA index, dearer than 1.5x the anchor’s effective cost, or its provider’s token window is spent) or reachable but dominated on every rung. ‘cost UNVERIFIED’ means at least one reachable model carried no published price, so the veto could prove it neither cheap nor expensive.THALAMUS frontier 16 rungs (€/task) of 169 across 48/121 models · bias 3 (balanced) floor idx 48.4 → pick claude-code/claude-fable-5-1@medium idx 48.9 €0.129/task · no domain switches at this bias · cost UNVERIFIED · out 73 cost-veto · upper bound: no quota snapshot, no auth filter€0.00€20€40€60€80€100€120€140€16010152025303540455055EFFECTIVE COST · €/Mtok OUTPUT · linearEFFECTIVE COST PER AVERAGE TASK · €/task · linearAA INTELLIGENCE INDEXSIZE ∝ CONTEXT128k256k1M75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%ling-3.0-flash · Effort ladder unknown — €0.06/Mtok · ~9.6k tok/task · 262k ctx · idx 24.9 (headline — AA published no per-effort split)inkling · Effort ladder unknown — €4.0/Mtok · ~9.6k tok/task · 200k ctx · idx 25.0 (headline — AA published no per-effort split)inkling-small · Effort ladder unknown — €1.2/Mtok · ~9.6k tok/task · 200k ctx · idx 27.8 (headline — AA published no per-effort split)nex-n2-pro · Effort ladder unknown — €1.0/Mtok · ~9.6k tok/task · 200k ctx · idx 28.2 (headline — AA published no per-effort split)75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%75%gpt-4o (openai)gpt-4o (copilot)gemini-2.0-flashgpt-4.1 (openai)gpt-4.1 (copilot)nemotron-3.5-lightninggemini-2.5-flashgrok-code-fast-1gemini-2.5-pro (google)gemini-2.5-pro (copilot)claude-haiku-4.5 (copilot)Claude Haiku 4.5 (claude-code)longcat-2.0o3gpt-5.1-codex-minigpt-5-minigpt-5.4-nanogemini-3.5-flash-litegpt-5gpt-5.1-codexgpt-5.4-mini (openai)gpt-5.4-mini (copilot)gpt-5.1 (openai)gpt-5.1 (copilot)ling-3.0-flashinklingHunyuan 3 (openrouter·tencent)Hunyuan 3 (tencent)Kimi K2.7 Code (openrouter·moonshotai)Kimi K2.7 Code (moonshot-ai)mimo-v2.5-pro (openrouter·xiaomi)mimo-v2.5-pro (xiaomi)GLM 5.1 (openrouter·z-ai)GLM 5.1 (z-ai)gemini-3-flash-preview (google)gemini-3-flash-preview (copilot)Kimi K2.6 (openrouter·moonshotai)Kimi K2.6 (moonshot-ai)qwen3.6-plusinkling-smallglm-5 (openrouter·z-ai)glm-5 (z-ai)gemini-3-pro-preview (google)gemini-3-pro-preview (copilot)solar-pro4nex-n2-proQwen 3.6 Max Previewgpt-5.2-codexclaude-opus-4.5MiniMax M3 (openrouter·minimax)MiniMax M3 (minimax)claude-sonnet-4.5Qwen 3.7 Maxgemini-3.1-pro-preview (google)gemini-3.1-pro-preview (copilot)Claude Sonnet 4.6 (claude-code)claude-sonnet-4.6 (copilot)gpt-5.2 (openai)gpt-5.2 (copilot)claude-opus-4.6 (copilot)Claude Opus 4.6 (claude-code)gpt-5.3-codex (openai)gpt-5.3-codex (copilot)gemini-3.5-flashQwen3.8 27B (openrouter·qwen)Qwen3.8 27B (darkbloom)Qwen3.8 27B (alibaba)GLM 5.2 (openrouter·z-ai)GLM 5.2 (deepinfra)GLM 5.2 (z-ai)Muse Spark 1.1gemini-3.6-flashdeepseek-v4-flashDeepSeek V4 Flashgpt-5.4-progpt-5.2-progpt-5.1-codex-maxDeepSeek V4 Flash Vision (openrouter·deepseek)DeepSeek V4 Flash Vision (deepinfra)claude-sonnet-4DeepSeek V4 Pro (openrouter·deepseek)DeepSeek V4 Pro (deepseek)deepseek-v4-progpt-5.6-luna (codex · openai · openai-codex)claude-sonnet-5gpt-5.5 (openai)gpt-5.5 (copilot)gpt-5.5 (openai-codex)Grok 4.5gpt-5.4 (openai)gpt-5.4 (copilot)DeepSeek V4.1 Flash (openrouter·deepseek)DeepSeek V4.1 Flash (deepseek)Muse Spark 1.2gemini-3.7-flashQwen3.8 2.4T A95BClaude Opus 4.7 (claude-code)claude-opus-4.7 (copilot)gemini-3.8-flashClaude Opus 4.8GLM 5.3 Flash (openrouter·z-ai)GLM 5.3 Flash (z-ai)gpt-5.6-terra (codex · openai · openai-codex)Kimi K3 (openrouter·moonshotai)Kimi K3 (moonshot-ai)Grok 4.6GLM 5.3 (openrouter·z-ai)GLM 5.3 (z-ai)Qwen 3.8 Max (0902)gpt-5.6-sol (codex · openai · openai-codex)Muse Spark 1.3Claude Opus 5GPT-6 Astra (openai-codex)gpt-6-astra (openai)Claude Fable 5.1THALAMUS frontier 16 rungs (€/task) of 169 across 48/121 models · bias 3 (balanced) floor idx 48.4 → pick claude-code/claude-fable-5-1@medium idx 48.9 €0.129/taskThe envelope is the Pareto frontier, on the €/task axis, of the effort rungs of the models src/shared/thalamus-candidates.ts can reach, computed by src/shared/thalamus-frontier.ts — the same module the reply-path router calls, never a list in this file. A rung is on the frontier when no other rung is both cheaper-or-equal per task AND smarter-or-equal; the BIAS dial picks the cheapest frontier rung within its gap of the best. A model with NO ring is either unreachable (no AA index, dearer than 1.5x the anchor’s effective cost, or its provider’s token window is spent) or reachable but dominated on every rung. ‘cost UNVERIFIED’ means at least one reachable model carried no published price, so the veto could prove it neither cheap nor expensive.THALAMUS frontier 16 rungs (€/task) of 169 across 48/121 models · bias 3 (balanced) floor idx 48.4 → pick claude-code/claude-fable-5-1@medium idx 48.9 €0.129/task · no domain switches at this bias · cost UNVERIFIED · out 73 cost-veto · upper bound: no quota snapshot, no auth filterno domain switches at this bias · cost UNVERIFIED · out 73 cost-veto · upper bound: no quota snapshot, no auth filterThe envelope is the Pareto frontier, on the €/task axis, of the effort rungs of the models src/shared/thalamus-candidates.ts can reach, computed by src/shared/thalamus-frontier.ts — the same module the reply-path router calls, never a list in this file. A rung is on the frontier when no other rung is both cheaper-or-equal per task AND smarter-or-equal; the BIAS dial picks the cheapest frontier rung within its gap of the best. A model with NO ring is either unreachable (no AA index, dearer than 1.5x the anchor’s effective cost, or its provider’s token window is spent) or reachable but dominated on every rung. ‘cost UNVERIFIED’ means at least one reachable model carried no published price, so the veto could prove it neither cheap nor expensive.THALAMUS frontier 16 rungs (€/task) of 169 across 48/121 models · bias 3 (balanced) floor idx 48.4 → pick claude-code/claude-fable-5-1@medium idx 48.9 €0.129/task · no domain switches at this bias · cost UNVERIFIED · out 73 cost-veto · upper bound: no quota snapshot, no auth filter
◎ outline + logo = model · size ∝ context window (see the three sample rings, bottom-left of the plot)— a circle per vendor-documented effort · solid ring = AA measured that effort · dotted ring + dotted line = ESTIMATED from other public per-effort benchmark runs (Epoch AI hub, LMArena) fitted to the AA scale, ±1σ on hover, never above a measured max · dashed ring on a dashed rail = neither, hung at the headline index · a model with no ladder stays one dot△ shaded triangle = the official API list price of the same model · – – – dashed bridge = the gap between what this plan pays and sticker · hover either to light both · 58 of 121 models are discounted by a plan we hold · widest gap 510×prepaid circles (Claude, Grok, ChatGPT) sit at 75% of the quota ceiling — drag one toward its triangle to read other utilisation; it snaps back on dropRead the two marks as different questions. A circle answers “what does this token cost ME” (Anthropic Max 20x amortises Opus 5’s $25 sticker to ~€0.073 — 340×, re-derived 2026-09-02 from live quota utilisation and measured burn, not from dividing the sticker); a triangle answers “what does this model cost”. Metered routes — every Chinese lab, Google, the OpenRouter tail — have no triangle because their circle already IS list price. Comparing a circle to a circle across those two bases is the one thing this chart must never be used for.€/task: tokens-per-task OckBench-anchored (Opus 5, Kimi K3), rest estimated · reference = Opus 5 @ high, Anthropic’s documented default (fixed)– – – dashed = one model, several vendors · Copilot is not a markup: GitHub charges each model’s own list price (13 of 14 triples verified 2026-08-15) — its dots sit right because our Anthropic plan is the deeper discount, and Copilot rows are prospective (we hold no Pro+ seat)Auto not plotted (uncapped) · not plotted: Gemma 4 26B (local, no token price), M365 Copilot (Think deeper) (no intelligence index) · 32 models sold by several vendors at different prices — dashed line joins the routes · 17 Chinese models priced across 30 suppliers (cheapest + lab plotted)

How to read it in sixty seconds

  • Up is the Artificial Analysis Intelligence Index — a composite of public benchmarks, not our opinion.
  • Right is effective cost in € per million output tokens, on a log scale by default. Flip the LOG / LINEAR switch to see why: on a linear axis everything interesting is a smear against the left edge.
  • Bubble size is the context window. The three unlabelled rings at bottom-left are the ruler: 128k, 256k, 1M.
  • Colour and logo are the vendor. The chips along the top switch each one on and off.

The two marks answer different questions — this is the part people get wrong

A circle answers “what does this token cost me?” A triangle answers “what does this model cost?” — the vendor’s published API list price. The dashed bridge between them is the gap a subscription buys. Hover either one and the tooltip spells out which is which.

Worked example. Anthropic lists Opus 5 output at $25 per million tokens. Amortised across a Max 20× seat it costs us roughly €0.07 — a spread of two orders of magnitude.

The number is derived, not divided: fee against live quota utilisation and measured burn. Dividing the sticker by a headline token allowance gives you a much prettier number and a wrong one. The prepaid circles sit at 75% of the quota ceiling — drag one toward its triangle to read other utilisation levels, and it snaps back when you let go.

Two consequences worth internalising:

  • Metered routes have no triangle. Every Chinese lab, Google, the whole OpenRouter tail — their circle already is list price. Nothing to bridge.
  • Never compare a circle to a circle across two different bases. A prepaid seat and a pay-per-use endpoint are not the same unit. That comparison is the one thing this chart must never be used for, and it is exactly the mistake that makes people say “the Chinese models are cheaper” when, in cash actually leaving the account, several of them are dearer than a Claude seat.

Effort is a ladder, not a point

A model does not have one intelligence score. It has one per reasoning-effort setting, and the curve is not always monotonic — Grok 4.6 measurably scores lower at xhigh than at high, because extra serial reasoning buys hard static reasoning at the cost of tool use and answer stability.

  • Solid ring — the vendor documents this effort level and Artificial Analysis measured it. A real number.
  • Dotted ring, dotted line — the level exists but AA never scored it, so it is estimated from other public per-effort benchmark runs (Epoch AI’s hub, LMArena) fitted onto the AA scale. Hover gives ±1σ. An estimate is never drawn above a measured maximum.
  • Dashed ring on a dashed rail — neither. Hung at the headline index, honestly flagged.

A stop the vendor does not document is never invented. An earlier version painted the same six-rung ladder on every model, which gave Grok six dots for two real settings. A fabricated dot at an invented cost is worse than a missing dot.

Why it plots models we cannot even use

Deliberately. The chart’s whole job is to catch a model cheaper or smarter than anything we currently route to — and a new competitor is, by definition, on a provider we do not have an account with yet. An earlier build filtered the catalogue down to configured providers, which quietly deleted the only points worth looking at.

The last line is a decision, not a caption

Along the bottom the router reports live. It is not a ranking — it is a Pareto frontier computed on the cost-per-task axis: a rung sits on the frontier when no other rung is both cheaper-or-equal per task and smarter-or-equal. A bias dial then walks that frontier, picking the cheapest frontier rung within its tolerance of the best available. Turn it toward quality and the pick climbs; turn it toward thrift and it slides down. A model with no ring is either unreachable (no published index, too dear against the anchor, or its token window is spent) or reachable but beaten on every rung.

The chart is not an illustration of that policy — it is drawn from the same module the reply-path router calls, so the two cannot quietly disagree.

2. The DOSSIER — because “smart” is not a scalar

An index number tells you nothing about whether a model is good at your subject, and nothing at all about whether it will simply decline your task. Hover any cell for the evidence behind that grade, and click a capability heading to re-rank the whole table by it. Ranking by COST is the fastest way to see how differently the two halves of the market are priced.

hover any cell for the evidence · click a capability heading to rank the table by it
SMART MODELS · DOSSIER
what each is best at · a column per censored topic · hover a cell for the evidence
CAPABILITYbest in classstrongadequate·not its jobCENSORSHIPrefusesvetted users onlypartial / holds weaklyanswersplain header = ANCHORED to a named benchmark · ? = JUDGED, no public anchor · dashed underline = graded, transport unproven · 1 on a glyph = best configured model in that column (measured Epoch AI percentile where one exists, else the grade)hover any cell · click a subject to rank by it
THALAMUS ROUTES · bias 3 (balanced)GENERAL → Claude Opus 5 @xhigh · idx 49.7 · €0.60/taskCODE → Claude Opus 5 @xhigh · idx 49.7 · €0.60/task · measured p93/2AGENTIC → Claude Opus 5 @xhigh · idx 49.7 · €0.60/task · measured p88/3REASON → Claude Opus 5 @xhigh · idx 49.7 · €0.60/task · measured p91/11WRITE → Claude Opus 5 @xhigh · idx 49.7 · €0.60/task · unmeasuredPSYCH → Claude Opus 5 @xhigh · idx 49.7 · €0.60/task · unmeasuredLONG CTX → Claude Opus 5 @xhigh · idx 49.7 · €0.60/task · unmeasuredVISION → Claude Opus 5 @xhigh · idx 49.7 · €0.60/task · unmeasuredWORLD → GPT-6 Astra @high · idx 50.9 · €1.19/task · measured p99/1
CAPABILITY — what to route to it CENSORSHIP — what it will not answer
MODEL · AA idx BLOC CODE AGENTIC? REASON? WRITE? PSYCH? LONG CTX? VISION? SPEED COST WORLD? CSAM BIO·CHEM EXPLOSIVES DRUGS MALWARE SEC RESEARCH ELECTIONS CN POLITICS ADULT
Claude Fable 5.1 53.4 US 1 ·
GPT-6 Astra 52.7 US 1 · · 1
Claude Opus 5 50.8 US 1 1 · ·
Muse Spark 1.3 48.1 ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ?
GPT-5.6 Sol 47.0 US ·
Qwen 3.8 Max (0902) 45.4 CN
GLM 5.3 44.8 CN 1
Grok 4.6 44.3 US
Kimi K3 43.6 CN ·1 1
GPT-5.6 Terra 42.1 US
GLM 5.3 Flash 41.8 CN
Claude Opus 4.8 41.8 US · ·
gemini-3.8-flash 40.9 US 1
Claude Opus 4.7 40.7 US · ·
Qwen3.8 2.4T A95B 39.9 CN
gemini-3.7-flash 39.6 US
Muse Spark 1.2 39.6 ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ?
DeepSeek V4.1 Flash 39.5 CN ·
Grok 4.5 38.8 US
gpt-5.5 38.4 US · 1 ·
claude-sonnet-5 38.2 US
GPT-5.6 Luna 37.3 US
DeepSeek V4 Pro 36.0 CN ·
DeepSeek V4 Flash Vision 34.8 CN ·
DeepSeek V4 Flash 34.3 CN ·
gemini-3.6-flash 34.0 US
Muse Spark 1.1 33.7 ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ?
GLM 5.2 33.7 CN
Qwen3.8 27B 33.7 CN
gemini-3.5-flash 33.6 US
Claude Opus 4.6 31.9 US · ·
Claude Sonnet 4.6 30.1 US
gemini-3.1-pro-preview 29.7 US
Qwen 3.7 Max 29.5 CN
MiniMax M3 29.2 ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ?
Qwen 3.6 Max Preview 28.4 CN
Kimi K2.6 27.0 CN ·
GLM 5.1 26.1 CN
Kimi K2.7 Code 25.8 CN ·
Hunyuan 3 25.3 ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ?
Claude Haiku 4.5 16.9 US · ·
Hermes 4 405B 9not wired OSS · · ·
Hermes 4 70B not wired OSS · · ·
Dolphin Mistral 24B Venice not wired OSS · · ·
Huihui Qwen3.5 27B abliterated not wired OSS · · · ·
REFUSED BY BOTH CAMPS
  • Sexual content involving minors — absolute in every lab tested, US and Chinese alike
  • Biological, chemical, radiological and nuclear weapons uplift
  • Bomb-making and device construction
  • Fraud, scam and phishing kits
  • Drug synthesis routes and trafficking logistics
WHERE THEY SPLIT
Security research — the big one — US models refuse to READ an exploit, not just to write one. During the July 2026 Hugging Face breach, Claude and GPT declined to process the attacker’s payloads and logs; Hugging Face ran GLM 5.2 in-house over 17,000+ telemetry events instead. Chinese models draw the line at BUILDING malware, not at analysing it. Anthropic’s Cyber Verification Program and OpenAI’s Trusted Access for Cyber exist to walk this back for vetted defenders.
Beijing politics is not a bloc property — On 168 cases covering topics the Chinese state suppresses, Kimi K2.5 scored 98.8% — identical to Claude Opus 4.5 — while DeepSeek V3.2 scored 19%. Same country, opposite behaviour. GLM sits in between and moves with the endpoint: 95.2% on local weights, 79.8% through a hosted API.
The mirror — US models are trained off INFLUENCING politics (campaigning, election persuasion); Chinese models off CRITICIZING power (CCP legitimacy, sovereignty). Both camps politicize refusals — in opposite directions, and each camp’s blind spot is the other’s specialty.
The OSS rows lift a lot — and nothing at the top — Hermes 4 405B, the most permissive model on a mainstream pay-per-use endpoint, answers 57.1% of RefusalBench’s 166 refusal-prone prompts against roughly 17% for GPT-4o and Claude Sonnet. It still declines ~43%. What lifts: over-refusal, dual-use security work, frank medical and legal talk, adult and dark fiction, taking a side. What does not: CSAM (enforced at the provider and legal layer, not the weights), real CBRN capability (the knowledge was never in the base, so past the refusal you get confident regurgitation), and — on abliterated Chinese bases — the political filtering, which is a different circuit from the English safety direction the technique removes. Provider terms still govern the account regardless of what the weights will say.
Refusing on paper ≠ holding under pressure — The CASI jailbreak-resilience index (Jul 2026) splits models more sharply than policy does: Claude Sonnet 5 93.08, Qwen3.5-397B 81.13, MiMo-V2.5 73.80, GLM-5.2 46.58. A ■ in this table means trained refusal, not a guarantee it survives contact with a determined prompt.

CAPABILITY grades are a synthesis across SWE-bench Verified, OckBench, published evaluations and vendor launch claims (marked as claims in the cell) — a routing prior, not a single measured benchmark. Two ★ in a column mean “either is a defensible first choice”, not a tie score — unless the ★ carries a dashed underline, which marks it GRADED, NOT WIRED: right about the model, unproven about our route to it, and therefore not a routable first choice until the transport is probed. Every column now declares its PROVENANCE. ANCHORED means the grades are read off a named public measurement of that column’s own question — only CODE, SPEED and COST clear that bar. JUDGED (a ? on the header) means there is no public anchor and the grade is an opinion with its reasoning in the tooltip. LONG CTX is JUDGED on purpose: its MEASURED line is the ADVERTISED WINDOW while the grade is useful RECALL, and a related figure is not a basis. PSYCH was added 2026-08-30 because the architect routes by it; EQ-Bench 3 is the anchor it should have and does not yet, for the two re-checkable reasons in that column’s tooltip. CENSORSHIP is training-time disposition from usage policies, published testing and the July 2026 Hugging Face incident reporting — not guarantees: jailbreaks, endpoint-side filtering and version drift all exist. Corrected 2026-08-06: an earlier version of this table listed cyberattack tooling as refused by both camps, which averaged away the one asymmetry that mattered. Since 2026-09-02 the 1 on a glyph marks the single best CONFIGURED model in that column — the highest Epoch AI percentile where any configured model has a public per-domain run, the judged grade where none does (WRITE and PSYCH at the time of writing) — and the header tooltip says which. The THALAMUS ROUTES strip above the table is the router’s own answer per task domain at the current BIAS, from the same functions the reply path calls.

CN PROVIDER × MODEL — best price per model

model lab direct DeepInfra Novita GMICloud Venice SiliconFlow AtlasCloud Phala StreamLake cheapest
qwen/qwen3.8-max-0902 sub? 6.00 Alibaba 6.00
qwen/qwen3.8-2.4t-a95b sub? 6.00 6.00 6.00 6.00 6.00 Novita 6.00
qwen/qwen3.8-27b sub? 2.55 1.88 3.00 3.20 2.08 Darkbloom 1.80
qwen/qwen3.7-max sub? 4.42 Alibaba 4.42
moonshotai/kimi-k3 sub 15.00 14.25 12.00 Relace 8.50
moonshotai/kimi-k2.7-code sub 4.00 3.40 3.84 4.00 3.50 3.80 3.00 StreamLake 3.00
moonshotai/kimi-k2.6 sub 4.00 3.50 3.40 3.60 3.50 3.40 4.00 4.60 2.52 DigitalOcean 2.40
z-ai/glm-5.3 sub 4.40 3.00 2.72 3.30 4.40 3.52 4.40 2.86 Morph 2.43
z-ai/glm-5.3-flash sub 0.5000 0.2500 0.4400 0.2500 0.5000 0.5000 0.5000 0.4250 0.4700 GMICloud 0.2500
z-ai/glm-5.2 sub 4.40 1.80 2.04 4.40 4.40 3.74 2.95 3.00 2.02 DeepInfra 1.80
z-ai/glm-5.1 sub 4.40 3.50 4.40 4.40 4.40 3.74 3.96 4.20 3.04 StreamLake 3.04
z-ai/glm-5 sub 3.20 3.20 1.92 3.20 2.55 1.92 StreamLake 1.92
deepseek/deepseek-v4-pro-0813 pay/use 3.96 2.60 2.97 3.17 4.95 3.96 3.96 4.36 2.95 DeepInfra 2.60
deepseek/deepseek-v4.1-flash pay/use 1.20 0.4200 1.14 1.14 1.50 1.20 1.20 1.10 0.8788 DeepInfra 0.4200
deepseek/deepseek-v4-flash-0731 pay/use 0.1800 1.23 0.8580 0.3500 0.6600 1.32 1.32 0.1716 Relace 0.1600
deepseek/deepseek-v4-flash-vision-exp pay/use 0.6468 1.32 1.32 1.32 DeepInfra 0.6468
minimax/minimax-m3 sub? 1.20 1.10 1.20 0.9600 1.20 1.20 1.20 CoreWeave 0.9600
tencent/hy3 pay/use 0.5280 0.5800 0.5800 0.5800 0.8000 0.6400 Tencent 0.5280
xiaomi/mimo-v2.5-pro pay/use 0.8700 1.17 0.9605 0.6090 0.8700 1.04 GMICloud 0.6090

$ per 1M OUTPUT tokens, cheapest endpoint per provider. Bold = best price for that model. Caveat: the cheapest route is often fp4 quantisation — hover a cell for its precision; cheap can buy lower quality. Subscriptions: Z.AI: GLM Coding Plan from $18/mo · Moonshot AI: Kimi Code from $19/mo · Minimax: MiniMax Coding Plan from $49/mo (unconfirmed) · Alibaba: Qwen Token Plan from $6/mo (unconfirmed) · DeepSeek: pay-per-use only. Everything else is pay-per-use credits. Source: openrouter.ai/api/v1/models/{id}/endpoints, fetched 2026-09-21 07:21 — regenerated daily by the model-rank-refresh cron, because these prices move within days.

Left half — what to route to it

Ten columns: CODE, AGENTIC, REASON, WRITE, PSYCH, LONG CTX, VISION, SPEED, COST, WORLD. Four grades: best in class, strong, adequate, not its job.

The provenance matters more than the grade. A plain header means the column is anchored to a named public benchmark — SWE-bench Verified, OckBench, published evaluations — and hovering a cell tells you which. A daggered header means there is no public anchor and the grade is a judgement. A dashed underline means the grade is right about the model but the transport is unproven: correct about the weights, untested through a given route.

Every grade is a routing prior, not a measurement. It is there to pick a first candidate, not to settle an argument.

Right half — what it will not answer

Nine columns from CSAM through BIO-CHEM, EXPLOSIVES, MALWARE, SEC RESEARCH, ELECTIONS, CN POLITICS to ADULT, graded refuses / vetted users only / partial / answers. This is training-time disposition read off usage policies, published testing and incident evidence — not a guarantee. Jailbreaks, endpoint-side filtering and version drift all exist.

What the table actually shows

  • The floor is universal. Sexual content involving minors, CBRN uplift, bomb-making, fraud and trafficking logistics: refused by every lab tested, US and Chinese alike. There is no permissive frontier model on those.
  • Security research is where the camps split. US models refuse to read an exploit, not just write one — during the July 2026 Hugging Face breach, Claude and GPT declined to process the attacker’s own payloads and logs, and Hugging Face ran GLM 5.2 in-house over 17,000+ telemetry events instead. Chinese models draw the line at building malware, not at analysing it. Anthropic’s Cyber Verification Program and OpenAI’s Trusted Access for Cyber exist to walk this back for vetted defenders.
  • Beijing politics is not a bloc property. Across 168 cases on topics the Chinese state suppresses, Kimi K2.5 scored 98.8% — identical to Claude Opus 4.5 — while DeepSeek V3.2 scored 19%. Same country, opposite behaviour. GLM sits in between and moves with the endpoint: 95.2% on local weights, 79.8% through the hosted API.
  • The mirror. US models are trained off influencing politics (campaigning, election persuasion); Chinese models off criticising power. Both camps politicise refusals, in opposite directions, and each camp’s blind spot is the other’s specialty.
  • Open weights lift a lot and nothing at the top. Hermes 4 405B, the most permissive model on a mainstream pay-per-use endpoint, answers 57.1% of RefusalBench’s refusal-prone prompts against roughly 17% for GPT-4o and Claude Sonnet — and still declines ~43%. What lifts is over-refusal, dual-use security work, frank medical and legal talk, adult and dark fiction. What does not lift is CSAM (enforced at the provider and legal layer, not in the weights) or real CBRN capability (the knowledge was never in the base, so past the refusal you get confident regurgitation).
  • Refusing on paper is not holding under pressure. The CASI jailbreak-resilience index splits models far more sharply than policy does: Claude Sonnet 5 at 93.08, Qwen3.5-397B at 81.13, MiMo-V2.5 at 73.00, GLM-5.2 at 46.58.

And the money, at the foot of the table

Scroll the dossier to its end and you hit a provider × model price matrix for the Chinese models, because they are the ones sold by nine different hosts at nine different prices. With the caveat that makes the cheap column dangerous: the cheapest route is very often FP4 quantisation. Hover a cell for its precision. Cheap can buy you a worse model wearing the same name.

Why bother drawing any of this

Because an agent that runs continuously makes thousands of model choices a day, and every one of them is this two-axis decision made badly or well. The chart stops us paying list price for something a plan we already hold covers. The dossier stops us routing a security-analysis task to a model that will refuse to read the log file.

If you want the rest of the machinery: the effort × model slider is how a choice gets made per turn, the command center is where the spend lands, and this is what the bill looks like at the end of a month.

This page is machine-maintained: the embeds and the figures above are regenerated and republished by our own nightly job, straight from the running agent. The one control that cannot cross over is the €/Mtok ↔ €/task switch — the per-task layout is computed live and does not survive an export. Everything else — zoom, pan, vendor filtering, tooltips, column ranking — is the real thing.

Posted by

in

Leave a Reply

Your email address will not be published. Required fields are marked *