tokenstat
tokenstat

Model catalog

Benchmarks

Arena Elo from free open daily snapshots on GitHub (text, code, vision, agent), plus a sparse Artificial Analysis subset OpenRouter embeds. One score set is cross-linked onto provider clones of the same model. Arena may score the same base model at different reasoning efforts, so those rows stay separate and link to the base model only when it exists in the catalog. The open dumps publish 20 models per Arena board (10 on agent), so a full board is short by design. Empty on a single board is a gap, not a ranking of zero.

168 scored rows · 417 linked into the catalog · as of · Models · AI plans

Arena code

Top 10 of 44Full board
#ModelScore
1
OpenAI: GPT-6 AstraOpenAI · gpt-6-astra · max effort
1800
2
Anthropic: Claude Fable 5.1Anthropic · claude-fable-5.1 · max effort
1758
3
Anthropic: Claude Opus 5Anthropic · claude-opus-5 · max effort
1687
4
Qwen3.8 Max 0902Alibaba · qwen3.8-max-0902
1681
5
Moonshot: Kimi K3Moonshot · kimi-k3 · max effort
1674
6
Qwen3.8 MaxAlibaba · qwen3.8 · max effort
1671
7
Anthropic: Claude Opus 5Anthropic · claude-opus-5 · high effort
1660
8
Meta: Muse Spark 1.3Meta · muse-spark-1.3 · max effort
1652
9
Qwen3.8 Flash NextAlibaba · qwen3.8-flash-next
1635
10
Anthropic: Claude Fable 5Anthropic · claude-fable-5 · high effort
1628

Arena agent

10 modelsFull board
#ModelScore
1
Anthropic: Claude Fable 5.1Anthropic · claude-fable-5.1 · max effort
13.71
2
OpenAI: GPT-6 AstraOpenAI · gpt-6-astra · max effort
11.54
3
Anthropic: Claude Opus 5Anthropic · claude-opus-5 · high effort
10.25
4
Anthropic: Claude Opus 5Anthropic · claude-opus-5 · max effort
10.16
5
Anthropic: Claude Fable 5Anthropic · claude-fable-5 · high effort
8.81
6
Anthropic: Claude Opus 4.8 ThinkingAnthropic · claude-opus-4-8 · high effort
8.19
7
OpenAI: GPT-5.6 SolOpenAI · gpt-5.6-sol · xhigh effort
7.1
8
Moonshot: Kimi K3Moonshot · kimi-k3 · max effort
6.22
9
Anthropic: Claude Sonnet 5 ThinkingAnthropic · claude-sonnet-5 · high effort
5.97
10
OpenAI: GPT-5.5OpenAI · GPT 5.5 · xhigh effort
5.03

Arena text

Top 10 of 19Full board
#ModelScore
1
Anthropic: Claude Fable 5Anthropic · claude-fable-5 · high effort
1506
2
Anthropic: Claude 4.6 Opus ThinkingAnthropic · claude-opus-4-6 · high effort
1505
3
Anthropic: Claude 4.7 Opus ThinkingAnthropic · claude-opus-4-7 · high effort
1502
4
Meta: Muse Spark 1.2Meta · muse-spark-1.2 · xhigh effort
1500
5
Anthropic: Claude Fable 5.1Anthropic · claude-fable-5.1 · max effort
1498
6
Anthropic: Claude 4.6 Opus ThinkingAnthropic · claude-opus-4.6
1497
7
Anthropic: Claude 4.7 Opus ThinkingAnthropic · claude-opus-4-7
1494
8
Anthropic: Claude Opus 5Anthropic · claude-opus-5 · high effort
1493
9
Google: Gemini 3.8 FlashGoogle · gemini-3.8-flash · high effort
1493
10
Meta: Muse Spark 1.1Meta · muse-spark-1.1
1493

Arena vision

Top 10 of 19Full board
#ModelScore
1
Anthropic: Claude Fable 5Anthropic · claude-fable-5 · high effort
1310
2
Qwen3.8 MaxAlibaba · qwen3.8 · max effort
1302
3
Anthropic: Claude 4.7 Opus ThinkingAnthropic · claude-opus-4-7 · high effort
1301
4
Anthropic: Claude 4.7 Opus ThinkingAnthropic · claude-opus-4-7
1300
5
Anthropic: Claude 4.6 Opus ThinkingAnthropic · claude-opus-4-6 · high effort
1299
6
Meta: Muse Spark 1.3Meta · muse-spark-1.3 · max effort
1294
7
Anthropic: Claude 4.6 Opus ThinkingAnthropic · claude-opus-4.6
1293
8
Meta: Muse Spark 1.2Meta · muse-spark-1.2 · xhigh effort
1292
9
Anthropic: Claude Fable 5.1Anthropic · claude-fable-5.1 · max effort
1289
10
Anthropic: Claude Opus 5Anthropic · claude-opus-5 · high effort
1289

Artificial Analysis

Top 10 of 130Full board
#ModelScore
1
OpenAI: GPT-6 AstraOpenAI · gpt-6-astra
61AA intelligence
2
xAI: Grok 4.6xAI · grok-4.6 · high effort
61AA intelligence
3
Meta: Muse Spark 1.3Meta · muse-spark-1.3 · xhigh effort
61AA intelligence
4
Anthropic: Claude Fable 5.1Anthropic · claude-fable-5.1
53.4AA intelligence
5
Anthropic: Claude Opus 5Anthropic · claude-opus-5
50.8AA intelligence
6
Anthropic: Claude Fable 5Anthropic · claude-fable-5
49.6AA intelligence
7
OpenAI: GPT-5.6 SolOpenAI · gpt-5.6-sol
47AA intelligence
8
Qwen3.8 Max 0902Alibaba · qwen3.8-max-0902
45.4AA intelligence
9
Mistral: GLM-5.3z-ai · glm-5.3
44.8AA intelligence
10
xAI: Grok 4.6x-ai · grok-4.6
44.3AA intelligence

Sources

Arena Elo from Arena AI, via open dumps by oolong-tea-2026. Artificial Analysis indices via OpenRouter (subset). Artificial Analysis for the indices themselves. Plus vendor-reported launch evals collected from announcements, analyst writeups, launch blogs, and community benchmarks. As of .

Turn benchmark scores into working sessions

Run agents in a local Git-enabled workspace, connect to another machine when needed, and track the sessions and tokens behind your work. See what your AI subscriptions covered.

CLI for macOS, Linux, and Windows

Data and list-price estimates may contain mistakes. Not investment advice.