LLM Stats
NEWRun every top model with one operatorOpen Playground ›

AI Leaderboard — Compare 300+ Top AI Models by Intelligence, Speed & Price

Independent rankings of GPT, Claude, Gemini, Llama, DeepSeek and 300+ AI models — composite LLM Stats Score, updated continuously from public benchmarks and live API metrics. See the full LLM Leaderboard for complete LLM rankings with advanced filters.

leads on reasoning
GPT-5.6 Sol
94.6% gpqa
wins at coding
Claude Opus 5
27 arena
cheapest in the top 10
Grok 4.5
$2.00 /M tok
fastest output
Step-3.5-Flash
813 tok/s
longest context window
Grok-4 Fast Reasoning
2.0M tokens
best open-weights
Kimi K3
93.5% gpqa
358| Open
RANKMODELLLM STATSREASONINGCODINGAGENTCODE ARENAPARAMSCONTEXTSPEEDPRICING $/MLICENSE
1
GPT-5.6 Sol
OpenAI
57.456.950.644.42,1341.1M88 c/s$7.78
2
A\
Claude Opus 5
Anthropic
56.355.442.741.52,6681.0M30 c/s$7.22
3
A\
Claude Fable 5
Anthropic
56.153.748.842.82,0031.0M111 c/s$14.44
4
A\
Claude Mythos PreviewUNRELEASED
Anthropic
56.056.846.738.2
5
Kimi K31
Moonshot AI
54.953.945.941.61,8162.8T1.0M57 c/s$4.33
6
Z
GLM-5.3NEW
Zhipu AI
54.754.945.441.8753B
7
🐋
DeepSeek-V4-Pro-0813NEW
DeepSeek
54.252.243.340.31.6T1.0M178 c/s$0.48
8
Qwen3.8 MaxNEW
Alibaba Cloud / Qwen Team
53.352.042.340.02.4T
9
GPT-5.6 Terra4
OpenAI
52.951.246.340.81,1851.1M113 c/s$3.11
10
A\
Claude Opus 4.84
Anthropic
51.951.443.937.01,6891.0M46 c/s$7.22
11
Muse Spark 1.14
Meta
51.852.437.938.01,0701.0M118 c/s$1.58
12
G
Gemini 3.7 FlashNEW
Google
51.250.138.636.01.0M360 c/s$1.08
13
A\
Claude Sonnet 55
Anthropic
49.649.140.134.01,5881.0M51 c/s$2.89
14
GPT-5.54
OpenAI
49.348.641.234.72,1241.1M$7.78
15
X
Grok 4.56
xAI
47.545.839.434.11,498500K77 c/s$2.44
1-15 of 358
Recent

New Models

Announced in the last 15 days.

All updates
Index

Performance Index

Composite TrueSkill ratings across published benchmarks, paired with each model's blended price (8 input : 1 output) per 1M tokens.

ListPlot
OrgOpennessCountry
Full leaderboard
REASONING INDEXSCORE
01
56.9
02
56.8
03
55.4
04
54.9
05
53.9
06
53.7
07
52.4
08
52.2
09
52.0
10
51.4

TrueSkill conservative rating (μ − 3σ) across all benchmarks tagged in this category. Higher is better. Methodology & full leaderboard →

LLM-STATS.COM · V1
Browse

Explore leaderboards by category

Find the best model for the job — by capability, modality, or industry.

Movement

Leaderboard Changes

Rank movements across arenas — top 10 models by conservative rating.

Open arena
Jul 20Jul 27Aug 3Aug 10#1#2#3#4#5#6#7#8#9#10gemini-3.1-pro-previewclaude-opus-4-6gpt-5.6-solclaude-sonnet-4-6gpt-5.1-medium-2025-11-12claude-opus-4-7qwen3.7-maxclaude-sonnet-4-5-20250929grok-4.20-beta-0309-reasoningkimi-k3

Position by conservative rating (μ − 3σ). Top 10 models over the last 30 days.

Trends

AI Trends

How models, prices, and capabilities are evolving over time.

All trends

Benchmark Saturation

SOTA score progression over time — each step is a new best model

Saturation (85%)LLMSTATS0%20%40%60%80%100%Apr '24Jul '24Oct '24Jan '25Apr '25Jul '25Oct '25Jan '26Apr '26Jul '26
AIME 2025SimpleQAMMMLUBrowseCompMCP AtlasScreenSpot ProMMMUMMMU ProOSWorldSciCodeApex AgentsTerminal Bench
Saturated (85%+)
AIME 2025100%+6.3/mo
SimpleQA97%+4.2/mo
MMMLU93%+0.6/mo
BrowseComp91%+3.3/mo
MCP Atlas88%+3.4/mo
ScreenSpot Pro88%+2.7/mo
MMMU86%+1.5/mo
Fastest improving
Apex Agents57%+4.1/mo
OSWorld79%+4.1/mo
Terminal Bench50%+2.1/mo
SciCode60%+1.6/mo
Hardest to crack
MMMU Pro84%+1.1/mo
SciCode60%+1.6/mo
Terminal Bench50%+2.1/mo
OSWorld79%+4.1/mo
7 saturated5 still challenging12 benchmarks tracked

The Closing Gap

LLM Stats Score gap between proprietary and open-weight models

LLMSTATS0102030405060202420252026
Closed SourceOpen Source

The Price of Intelligence

Cost per LLM Stats Score point over time — each line tracks the cheapest model

$1.0$0.10$0.0130+ Score40+ Score50+ ScoreLLMSTATSJul '24Oct '24Jan '25Apr '25Jul '25Oct '25Jan '26Apr '26Jul '26
30+ Score40+ Score50+ Score

40+ tier: 98% cheaper — $0.09 → $0.002 per LLM Stats Score point

The Context Race

Maximum context window over time — each step is a new record

GPT-3.5 TurboGPT-4 TurboGPT-4.1Grok-4 Fast Reason…10K100K1.0MContext Window (tokens)LLMSTATS202420252026

Current record: Grok-4 Fast Reasoning at 2.0M tokens

Research

Latest Research

Analysis and comparisons from the LLM Stats team.

All posts
Compare

Hot Comparisons

Head-to-head matchups between frontier models.

Compare any models
Side by side

Compare AI models

Free head-to-head playgrounds across image, video, website, game and chat modalities.

All arenas
Arenas

Arena Rankings

Aggregated head-to-head performance across every sub-arena.

View all

Chat Arena

Vote now
#ModelChat Arena
1
G
Gemini 3.1 Pro
Google
2559
2
A\
Claude Opus 4.6
Anthropic
2541
3
GPT-5.6 Sol
OpenAI
2520
4
A\
Claude Sonnet 4.6
Anthropic
2515
5
GPT-5.1 Medium
OpenAI
2490
6
A\
Claude Opus 4.7
Anthropic
2488
7
Qwen3.7 Max
Qwen
2485
8
A\
Claude Sonnet 4.5
Anthropic
2481
9
Kimi K3
MoonshotAI
2473
10
A\
Claude Opus 4.8
Anthropic
2460
1-10 of 87
Previous12345Next

Coding Arena

Vote now
#ModelAVGWebsite3DGameAnimationSVGDataVizMIDI
1
A\
Claude Opus 5
Anthropic
2668165933602984
2
🐋
DeepSeek-V4-Flash-0731
DeepSeek
24072407
3
G
Gemini 3.1 Pro
Google
2164721222312722263295629402775
4
GPT-5.6 Sol
OpenAI
21341415295419072394269218111769
5
A\
Claude Opus 4.6
Anthropic
21271123217815632348220228722604
6
GPT-5.5
OpenAI
21241056244520552699224822692095
7
A\
Claude Fable 5
Anthropic
2003117527241801273727976102176
8
A\
Claude Opus 4.7
Anthropic
1926960211613952463218123632003
9
Kimi K3
MoonshotAI
1816969143228152047
10
GPT-5.4
OpenAI
1783727161913521948204221672626
1-10 of 88
Previous12345Next
Benchmarks

Featured Benchmarks

Performance evaluations across reasoning, coding, math and multimodal domains.

All benchmarks

GPQA

A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Questions are Google-proof and extremely difficult, with PhD ex...

PhysicsReasoningGeneralBiologyChemistry239 models
1A\Claude Mythos Preview0.9
1GPT-5.6 Sol0.9
3GGemini 3.1 Pro0.9
4A\Claude Opus 4.70.9
5GPT-5.50.9
5A\Claude Opus 4.80.9
7Kimi K30.9
8GPT-5.2 Pro0.9
9XGrok 4.50.9
10GPT-5.6 Terra0.9
View full leaderboard

MMLU

Massive Multitask Language Understanding benchmark testing knowledge across 57 diverse subjects including STEM, humanities, social sciences, and professional domains

LanguageLegalMathReasoningFinanceGeneralHealthcare101 models
1GPT-50.9
2o10.9
3o1-preview0.9
3GPT-4.50.9
5Qwen3 VL 235B A22B Thinking0.9
5Sarvam-105B0.9
7A\Claude 3.5 Sonnet0.9
7A\Claude 3.5 Sonnet0.9
9GPT-4.10.9
9Kimi K2 09050.9
View full leaderboard

AIME 2025

All 30 problems from the 2025 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with integer answers from 000-999. U...

MathReasoning115 models
1XGrok-4 Heavy1.0
1GGemini 3 Pro1.0
1GPT-5.2 Pro1.0
1GPT-5.21.0
1Kimi K2-Thinking-09051.0
6A\Claude Opus 4.61.0
7GGemini 3 Flash1.0
8LongCat-Flash-Thinking-26011.0
8GPT-5.1 High1.0
10Nemotron 3 Nano (30B A3B)1.0
View full leaderboard

LiveCodeBench

LiveCodeBench is a holistic and contamination-free evaluation benchmark for large language models for code. It continuously collects new problems from programming contests (LeetCod...

ReasoningGeneralCode75 models
1🐋DeepSeek-V4-Pro-Max0.9
2🐋DeepSeek-V4-Flash-Max0.9
3🐋DeepSeek-V4-Flash-04230.9
4Solar Pro 40.9
5🐋DeepSeek-V3.20.8
5🐋DeepSeek-V3.2 (Thinking)0.8
7MiniMax M20.8
8LongCat-Flash-Thinking-26010.8
9Nemotron 3 Super (120B A12B)0.8
10XGrok-3 Mini0.8
View full leaderboard

MMMU

MMMU (Massive Multi-discipline Multimodal Understanding) is a benchmark designed to evaluate multimodal models on college-level subject knowledge and deliberate reasoning. Contains...

MultimodalReasoningGeneralHealthcareVision63 models
1Qwen3.6 Plus0.9
2GPT-5.10.9
2GPT-5.1 Instant0.9
2GPT-5.1 Thinking0.9
5GPT-50.8
6Qwen3.5-122B-A10B0.8
7Qwen3.6-27B0.8
7o30.8
9Qwen3.5-27B0.8
10GGemini 2.5 Pro Preview 06-050.8
View full leaderboard

TAU-bench Retail

A benchmark for evaluating tool-agent-user interaction in retail environments. Tests language agents' ability to handle dynamic conversations with users while using domain-specific...

ReasoningCommunicationTool Calling25 models
1A\Claude Sonnet 4.50.9
2A\Claude Opus 4.10.8
3A\Claude Opus 40.8
4A\Claude 3.7 Sonnet0.8
5A\Claude Sonnet 40.8
6ZGLM-4.50.8
7ZGLM-4.5-Air0.8
8Qwen3-Coder 480B A35B Instruct0.8
9o4-mini0.7
10o10.7
View full leaderboard
View all benchmarksExplore all 687 benchmark leaderboards.
Notice missing or incorrect data? Let us know
Leaderboard guide

Compare the best AI models with one independent score.

The LLM Stats leaderboard ranks GPT, Claude, Gemini, Llama, DeepSeek, Qwen, Mistral, GLM and more by intelligence, speed and price. Every score is sourced from public benchmarks and live API metrics.

FAQ

Quick answers for choosing, comparing and interpreting today's leading AI models.

On the LLM Stats Leaderboard, GPT-5.6 Sol currently leads on GPQA Diamond (94.6% gpqa), the most discriminating reasoning benchmark at the frontier. This AI leaderboard ranks models by the LLM Stats Score, which aggregates GPQA, SWE-Bench Verified, coding-arena performance and pricing into one comparable AI ranking. Rankings refresh continuously as new benchmark results land.

Join our newsletter and stay up to date with everything AI

There's too much noise in AI, let's filter it for you. Get a curated digest of models, benchmarks, and the analysis that matters, right in your inbox once a week.

No spam, unsubscribe anytime