Lens · L1 · Benchmarks
Coding benchmarks tripled in 18 months. Compute constraints did not.
The capability curve, measured. 86 models, 647 sourced scores across 7 labs — self-reported and independently verified side by side. The data that answers: how fast is the frontier actually moving?
SWE-bench VerifiedvVerifiedMaintained by Princeton / OpenAIHuman-validated subset of real GitHub issues the model must resolve with a working patch.Benchmark docs →
| # | Model ↕ | Source | Score ↓ | As of ↕ |
|---|---|---|---|---|
| 1 | claude Opus 4.8Anthropic | Self | 88.6% | 2026-05-28 |
| 2 | claude Opus 4.7Anthropic | Self | 87.6% | 2026-04-16 |
| 3 | claude Opus 4.5Anthropic | Self | 80.9% | 2025-11-24 |
| 4 | claude Opus 4.6Anthropic | Self | 80.8% | 2026-02-05 |
| 5 | deepseek DeepSeek-V4 Preview (V4-Pro)DeepSeek | Self | 80.6% | 2026-04-24 |
| 6 | gpt GPT-5.2OpenAI | Self | 80% | 2025-12-11 |
| 7 | claude Sonnet 4.6Anthropic | Self | 79.6% | 2026-02-17 |
| 8 | claude Sonnet 4.5Anthropic | Self | 77.2% | 2025-09-29 |
| 9 | gpt GPT-5.1OpenAI | Self | 76.3% | 2025-11-13 |
| 10 | gemini Gemini 3 ProGoogle DeepMind (Alphabet) | Self | 76.2% | 2025-11-18 |
| 11 | claude Haiku 4.5Anthropic | Self | 73.3% | 2025-10-15 |
| 12 | claude Sonnet 4Anthropic | Self | 72.7% | 2025-05-22 |
| 13 | claude Opus 4Anthropic | Self | 72.5% | 2025-05-22 |
| 14 | claude Sonnet 3.7Anthropic | Self | 70.3% | 2025-02-24 |
| 15 | o-series o3OpenAI | Self | 69.1% | 2025-04-16 |
| 16 | o-series o4-miniOpenAI | Self | 68.1% | 2025-04-16 |
| 17 | deepseek DeepSeek-V3.1DeepSeek | Self | 66% | 2025-08-21 |
| 18 | gpt-oss gpt-oss-120bOpenAI | Self | 62.4% | 2025-08-05 |
| 19 | gpt-oss gpt-oss-20bOpenAI | Self | 60.7% | 2025-08-05 |
| 20 | gemini 2.5 FlashGoogle DeepMind (Alphabet) | Self | 60.4% | 2025-06-26 |
| 21 | deepseek DeepSeek-R1-0528DeepSeek | Self | 57.6% | 2025-05-28 |
| 22 | gpt GPT-4.1OpenAI | Self | 54.6% | 2025-04-14 |
| 23 | claude Sonnet 3.5 (v2)Anthropic | Self | 49% | 2024-10-22 |
| 24 | o-series o3-miniOpenAI | Self | 48.9% | 2025-01-31 |
| 25 | gemini Gemini 2.5 Flash-LiteGoogle DeepMind (Alphabet) | Self | 44.9% | 2025-06-17 |
| 26 | o-series o1OpenAI | Self | 40.9% | 2024-12-05 |
| 27 | claude Haiku 3.5Anthropic | Self | 40.6% | 2024-10-22 |
| 28 | gpt GPT-4.5OpenAI | Self | 38% | 2025-02-27 |
| 29 | deepseek DeepSeek-V2.5DeepSeek | Self | 16.8% | 2024-09-05 |
Sources: self-reported scores from lab announcements and system cards; independent scores from Artificial Analysis, LMArena, and academic evaluators. Every row carries a source URL and confidence level. Self-reported and independent numbers are never averaged — divergence is intentionally visible.