AI Leaderboard — Compare 300+ Top AI Models by Intelligence, Speed & Price
Independent rankings of GPT, Claude, Gemini, Llama, DeepSeek and 300+ AI models — composite LLM Stats Score, updated continuously from public benchmarks and live API metrics. See the full LLM Leaderboard for complete LLM rankings with advanced filters.
| RANK | MODEL | LLM STATS | REASONING | CODING | AGENT | CODE ARENA | PARAMS | CONTEXT | SPEED | PRICING $/M | LICENSE |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | ◎ GPT-5.6 Sol– OpenAI | 57.4 | 56.9 | 50.6 | 44.4 | 2,134 | — | 1.1M | 88 c/s | $7.78 | |
| 2 | A\ Claude Opus 5– Anthropic | 56.3 | 55.4 | 42.7 | 41.5 | 2,668 | — | 1.0M | 30 c/s | $7.22 | |
| 3 | A\ Claude Fable 5– Anthropic | 56.1 | 53.7 | 48.8 | 42.8 | 2,003 | — | 1.0M | 111 c/s | $14.44 | |
| 4 | A\ Claude Mythos Preview–UNRELEASED Anthropic | 56.0 | 56.8 | 46.7 | 38.2 | — | — | — | — | — | |
| 5 | ≋ Kimi K3↓ 1 Moonshot AI | 54.9 | 53.9 | 45.9 | 41.6 | 1,816 | 2.8T | 1.0M | 57 c/s | $4.33 | |
| 6 | Z GLM-5.3–NEW Zhipu AI | 54.7 | 54.9 | 45.4 | 41.8 | — | 753B | — | — | — | |
| 7 | 🐋 DeepSeek-V4-Pro-0813–NEW DeepSeek | 54.2 | 52.2 | 43.3 | 40.3 | — | 1.6T | 1.0M | 178 c/s | $0.48 | |
| 8 | ✦ Qwen3.8 Max–NEW Alibaba Cloud / Qwen Team | 53.3 | 52.0 | 42.3 | 40.0 | — | 2.4T | — | — | — | |
| 9 | ◎ GPT-5.6 Terra↓ 4 OpenAI | 52.9 | 51.2 | 46.3 | 40.8 | 1,185 | — | 1.1M | 113 c/s | $3.11 | |
| 10 | A\ Claude Opus 4.8↓ 4 Anthropic | 51.9 | 51.4 | 43.9 | 37.0 | 1,689 | — | 1.0M | 46 c/s | $7.22 | |
| 11 | ∞ Muse Spark 1.1↓ 4 Meta | 51.8 | 52.4 | 37.9 | 38.0 | 1,070 | — | 1.0M | 118 c/s | $1.58 | |
| 12 | G Gemini 3.7 Flash–NEW Google | 51.2 | 50.1 | 38.6 | 36.0 | — | — | 1.0M | 360 c/s | $1.08 | |
| 13 | A\ Claude Sonnet 5↓ 5 Anthropic | 49.6 | 49.1 | 40.1 | 34.0 | 1,588 | — | 1.0M | 51 c/s | $2.89 | |
| 14 | ◎ GPT-5.5↓ 4 OpenAI | 49.3 | 48.6 | 41.2 | 34.7 | 2,124 | — | 1.1M | — | $7.78 | |
| 15 | X Grok 4.5↓ 6 xAI | 47.5 | 45.8 | 39.4 | 34.1 | 1,498 | — | 500K | 77 c/s | $2.44 |
New Models
Announced in the last 15 days.
Performance Index
Composite TrueSkill ratings across published benchmarks, paired with each model's blended price (8 input : 1 output) per 1M tokens.
TrueSkill conservative rating (μ − 3σ) across all benchmarks tagged in this category. Higher is better. Methodology & full leaderboard →
Explore leaderboards by category
Find the best model for the job — by capability, modality, or industry.
Leaderboard Changes
Rank movements across arenas — top 10 models by conservative rating.
Position by conservative rating (μ − 3σ). Top 10 models over the last 30 days.
AI Trends
How models, prices, and capabilities are evolving over time.
Benchmark Saturation
SOTA score progression over time — each step is a new best model
The Closing Gap
LLM Stats Score gap between proprietary and open-weight models
The Price of Intelligence
Cost per LLM Stats Score point over time — each line tracks the cheapest model
40+ tier: 98% cheaper — $0.09 → $0.002 per LLM Stats Score point
The Context Race
Maximum context window over time — each step is a new record
Current record: Grok-4 Fast Reasoning at 2.0M tokens
Latest Research
Analysis and comparisons from the LLM Stats team.
2d agoGemini 3.7 Flash Release, Benchmarks And Price
Gemini 3.7 Flash scores 65.3% on DeepSWE v1.1 and 43.6% on FrontierCode 1.1 Main, at $0.75/$3.75 through 2026. Same 1M context as 3.6 Flash. List price doubles on…
3d agoDeepSeek-V4-Pro-0813: Pro Leaves PreviewDeepSeek-V4-Pro-0813 is the GA version behind API id deepseek-v4-pro. $0.435 / $0.003625 cached / $0.87 per 1…
3d agoGrok 4.6 Release, Benchmarks And Agent LoopsGrok 4.6 scores 1753 on GDPVal-AA v2 and 65.9% on DeepSWE v1.1, with 500K context and $2/$6 list pricing. The…
3d agoLFM2.5-VL-3B: Liquid's Faster Edge Vision-Language ModelLiquid AI LFM2.5-VL-3B (~Aug 12, 2026): ~3.1B open-weight edge vision-language model, non-reasoning VL, LFM …
3d agoQwen3.8-Max Open Weights: First Max-Class Qwen You Can DownloadAug 12, 2026 open-weights drop of Qwen3.8-2.4T-A95B — the first Max-class downloadable Qwen. Custom license;…
4d agoMAI-Code-1.1-Flash: Microsoft's Cheaper Vision Coding ModelMAI-Code-1.1-Flash is Microsoft's vision-capable Copilot coding model: 256K context, 138B/5B MoE, ~73% lower l…Hot Comparisons
Head-to-head matchups between frontier models.
Compare AI models
Free head-to-head playgrounds across image, video, website, game and chat modalities.
Arena Rankings
Aggregated head-to-head performance across every sub-arena.
Chat Arena
| # | Model | Chat Arena |
|---|---|---|
| 1 | G Gemini 3.1 Pro Google | 2559 |
| 2 | A\ Claude Opus 4.6 Anthropic | 2541 |
| 3 | ◎ GPT-5.6 Sol OpenAI | 2520 |
| 4 | A\ Claude Sonnet 4.6 Anthropic | 2515 |
| 5 | ◎ GPT-5.1 Medium OpenAI | 2490 |
| 6 | A\ Claude Opus 4.7 Anthropic | 2488 |
| 7 | ✦ Qwen3.7 Max Qwen | 2485 |
| 8 | A\ Claude Sonnet 4.5 Anthropic | 2481 |
| 9 | ≋ Kimi K3 MoonshotAI | 2473 |
| 10 | A\ Claude Opus 4.8 Anthropic | 2460 |
Coding Arena
| # | Model | AVG ↓ | Website | 3D | Game | Animation | SVG | DataViz | MIDI |
|---|---|---|---|---|---|---|---|---|---|
| 1 | A\ Claude Opus 5 Anthropic | 2668 | – | – | 1659 | 3360 | 2984 | – | – |
| 2 | 🐋 DeepSeek-V4-Flash-0731 DeepSeek | 2407 | – | – | – | – | 2407 | – | – |
| 3 | G Gemini 3.1 Pro Google | 2164 | 721 | 2223 | 1272 | 2263 | 2956 | 2940 | 2775 |
| 4 | ◎ GPT-5.6 Sol OpenAI | 2134 | 1415 | 2954 | 1907 | 2394 | 2692 | 1811 | 1769 |
| 5 | A\ Claude Opus 4.6 Anthropic | 2127 | 1123 | 2178 | 1563 | 2348 | 2202 | 2872 | 2604 |
| 6 | ◎ GPT-5.5 OpenAI | 2124 | 1056 | 2445 | 2055 | 2699 | 2248 | 2269 | 2095 |
| 7 | A\ Claude Fable 5 Anthropic | 2003 | 1175 | 2724 | 1801 | 2737 | 2797 | 610 | 2176 |
| 8 | A\ Claude Opus 4.7 Anthropic | 1926 | 960 | 2116 | 1395 | 2463 | 2181 | 2363 | 2003 |
| 9 | ≋ Kimi K3 MoonshotAI | 1816 | 969 | – | 1432 | 2815 | 2047 | – | – |
| 10 | ◎ GPT-5.4 OpenAI | 1783 | 727 | 1619 | 1352 | 1948 | 2042 | 2167 | 2626 |
Featured Benchmarks
Performance evaluations across reasoning, coding, math and multimodal domains.
GPQA
A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Questions are Google-proof and extremely difficult, with PhD ex...
MMLU
Massive Multitask Language Understanding benchmark testing knowledge across 57 diverse subjects including STEM, humanities, social sciences, and professional domains
AIME 2025
All 30 problems from the 2025 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with integer answers from 000-999. U...
LiveCodeBench
LiveCodeBench is a holistic and contamination-free evaluation benchmark for large language models for code. It continuously collects new problems from programming contests (LeetCod...
MMMU
MMMU (Massive Multi-discipline Multimodal Understanding) is a benchmark designed to evaluate multimodal models on college-level subject knowledge and deliberate reasoning. Contains...
TAU-bench Retail
A benchmark for evaluating tool-agent-user interaction in retail environments. Tests language agents' ability to handle dynamic conversations with users while using domain-specific...
Compare the best AI models with one independent score.
The LLM Stats leaderboard ranks GPT, Claude, Gemini, Llama, DeepSeek, Qwen, Mistral, GLM and more by intelligence, speed and price. Every score is sourced from public benchmarks and live API metrics.
FAQ
Quick answers for choosing, comparing and interpreting today's leading AI models.
On the LLM Stats Leaderboard, GPT-5.6 Sol currently leads on GPQA Diamond (94.6% gpqa), the most discriminating reasoning benchmark at the frontier. This AI leaderboard ranks models by the LLM Stats Score, which aggregates GPQA, SWE-Bench Verified, coding-arena performance and pricing into one comparable AI ranking. Rankings refresh continuously as new benchmark results land.





