Model API pricing
/api/model-pricingList prices and published task-bench scores for major model APIs. Each bench is one board, and an empty cell means that board published no score.
List prices and published task-bench scores for major model APIs. Each bench is one board, and an empty cell means that board published no score.
One chart per board, grouped by the kind of work. Click a name to keep that model, and Command-click to add another. On a point chart, hovering a point draws a line at that score and marks every point above it.
The vertical axis is the score and the horizontal axis is the thinking level.
A desktop task scores only when it finishes completely. Each cell is that model's latest run on XLANG's full set at 500 steps, and partial credit is named in the note.
Desktop tasks with no internet, scored with partial credit. This is OpenAI's offline set, so it is a different task list from the strict row.
Desktop tasks scored with partial credit on OSWorld 2.1. This is Anthropic's published partial score, not the binary strict board above.
The vertical axis is the score and the horizontal axis is the cost of the full run.
Terminal tasks in software, configuration and data analysis. Each point is one model, and the dollar amount is the cost of the full 66-task run.
Anthropic's own Terminal-Bench 4.0 run, published with Sonnet 5.5. It is not the Snorkel board in the row above.
The vertical axis is the score and the horizontal axis is the mean cost per task, on a log scale.
Long coding tasks in real repositories. A line joins one model's thinking levels, and each point is that level's mean cost per task.
Agentic coding tasks run in Cursor. The bar is the score published for that model.
Whether a code change would be merged without extra edits. Sonnet 5.5's marked score is xhigh, and max is named in the note.
Coding tasks from real Cursor sessions. This is CursorBench 4.0, not the 3.2 comparison above.
Scientific workflows that analyse data, run simulations and fit models. A note names a score from another lab's run.
Multi-step business workflows. A note names a score from another lab's run.
Bars show the spread between these scores.
Professional knowledge-work tasks, scored as Elo. A longer bar is a higher score, and the axis runs from the lowest score here to the highest.
Bars show the spread between these scores.
Professional knowledge-work tasks, scored as Elo on GDPval-AA v2.1. A note names a pre-release run or a score that may predate a fix.
Bars show the spread between these scores.
Long-horizon knowledge work, scored as Elo. A note names a pre-release run or a score that may predate a fix.
Multidisciplinary questions, answered with tools. Every cell on this board is a with-tools score.
Reading charts, without tools. A note names a score checked after an image-understanding fix.
Each row is one board, and the shaded cell is the highest score on that row. A cell shows the best score with its thinking level in brackets, and hovering it lists the other figures for that score.
| Bench | Claude Fable 5.1Anthropic | Claude Fable 5Anthropic | Claude Opus 5.5Anthropic | Claude Opus 5Anthropic | Claude Opus 4.8Anthropic | Claude Opus 4.7Anthropic | Claude Sonnet 5.5Anthropic | Claude Sonnet 5Anthropic | Claude Sonnet 4.6Anthropic | GPT-6 AstraOpenAI | GPT-6 SolOpenAI | GPT-5.6 SolOpenAI | GPT-5.6 TerraOpenAI | GPT-5.6 LunaOpenAI | GPT-5.5OpenAI | Gemini 3.8 FlashGoogle | Grok 4.7xAI | Grok 4.6xAI | Grok 4.5xAI | Kimi K2.6Moonshot | Kimi K2.7 CodeMoonshot | DeepSeek V4 ProDeepSeek | Qwen 3.7-PlusAlibaba | MiniMax M3MiniMax |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Computer use OSWorld 2.0 strict | — | — | — | max 44.3%. xhigh 33.3%. high 36.9%. medium 33.0%. low 18.8%. v2.1. Partial 77.7% at max. | Partial 54.8%. Batched tools. | Partial 48.9%. Batched tools. | — | — | max 8.3%. medium 9.3%. Partial 33.9% at medium. | — | — | v2026.08.08. Partial 62.7%. | — | — | Partial 49.5%. | — | — | — | — | Partial 22.1%. | — | — | Partial 21.5%. | Partial 22.3%. |
Computer use OSWorld 2.0 offline | — | — | — | — | — | — | — | — | — | 72.6% | — | 65.7% | — | — | — | — | — | — | — | — | — | — | — | — |
Computer use OSWorld 2.1 partial | — | — | 81.8%Partial. | — | — | — | 80.1%Partial. | 57.0%Partial. | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
Terminal agent Terminal-Bench 4.0 | max, Claude Code. Run $6.2k. | max, Claude Code. Run $7.3k. | — | xhigh, Claude Code. Run $6.1k. | max, Claude Code. Run $6.5k. | — | — | max, Claude Code. Run $9.6k. | — | max, Codex. Run $3.3k. | — | max, Codex. Run $2.5k. | max, Codex. Run $1.7k. | max, Codex. Run $0.3k. | — | high, mini-SWE-agent. Run $1.8k. | xhigh, Grok Build. Run $3.7k. | high, Grok Build. Run $3.6k. | high, Grok Build. Run $2.1k. | — | — | — | — | — |
Terminal agent Terminal-Bench 4.0, Anthropic | — | — | 66.4% (xhigh) | — | — | — | 70.6% | 10.3% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
Coding agent DeepSWE v1.1 | — | max 69.7% · $21.63 a task. xhigh 69.9% · $13.41 a task. high 68.6% · $9.18 a task. medium 65.4% · $6.09 a task. low 59.6% · $3.76 a task | — | max 73.6% · $11.84 a task. xhigh 73.2% · $9.07 a task. high 72.8% · $6.08 a task. medium 68.9% · $3.29 a task. low 58.1% · $1.66 a task | max 59.0% · $13.22 a task. xhigh 54.4% · $8.01 a task. high 51.8% · $4.28 a task. medium 48.7% · $3.44 a task. low 40.8% · $2.29 a task | — | — | max 53.8% · $26.40 a task. xhigh 49.7% · $11.89 a task. high 48.2% · $7.43 a task. medium 39.8% · $4.08 a task. low 30.5% · $2.19 a task | — | max 73.2% · $7.50 a task. xhigh 74.1% · $4.43 a task. high 73.2% · $3.92 a task. medium 72.8% · $3.08 a task. low 67.0% · $1.60 a task | — | max 72.7% · $6.46 a task. xhigh 70.7% · $3.60 a task. high 69.4% · $2.66 a task. medium 61.1% · $1.42 a task. low 45.4% · $0.82 a task | max 69.6% · $3.96 a task. xhigh 60.2% · $1.70 a task. high 53.8% · $0.91 a task. medium 35.1% · $0.47 a task. low 24.1% · $0.34 a task | max 67.2% · $0.61 a task. xhigh 56.9% · $0.31 a task. high 44.2% · $0.16 a task. medium 11.3% · $0.04 a task. low 1.5% · $0.01 a task | — | high 73.8% · $2.36 a task. medium 71.0% · $1.97 a task | — | xhigh 66.7% · $5.50 a task. high 65.2% · $4.38 a task. medium 67.5% · $3.45 a task. low 41.6% · $1.04 a task | $2.42 a task | — | No effort setting.. $2.82 a task | $1.67 a task | — | — |
Coding agent CursorBench 3.2 | 73.4% | 70.5% | — | 70.0% | — | — | — | — | — | — | — | 67.2% | — | — | — | — | — | — | — | — | — | — | — | — |
Coding agent FrontierCode 1.1 | — | — | 54.4% | — | — | — | Xhigh. Max scored 46.2%. | 42.4% | — | — | 49.3% | — | — | — | — | — | — | — | — | — | — | — | — | — |
Coding agent CursorBench 4.0 | — | — | 57.8% | — | — | — | 55.5% | 34.1% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
Research agent Terminal-Bench Science 0.1 | 52.6% | 24.7% | — | 29.0% | — | — | — | — | — | 64.6%OpenAI's run. | — | 22.4% | — | — | — | — | — | — | — | — | — | — | — | — |
Workflow agent AutomationBench | 31.4% | 17.1% | — | 26.9% | — | — | — | — | — | 41.4%OpenAI's run. | — | 19.6% | — | — | — | — | — | — | — | — | — | — | — | — |
Knowledge work GDPval-AA v2 | 1,853 | 1,723 | — | 1,824 | — | — | — | — | — | — | — | 1,711 | — | — | — | — | — | — | — | — | — | — | — | — |
Knowledge work GDPval-AA v2.1 | — | — | 1,846 | — | — | — | 1,844Pre-release run. A structured-output bug may understate this score. | 1,449 | — | — | 1,487May predate an image-understanding fix. | — | — | — | — | — | — | — | — | — | — | — | — | — |
Knowledge work AA-Briefcase v1.1 | — | — | 1,822 | — | — | — | 1,811Pre-release run. A structured-output bug may understate this score. | 1,359 | — | — | 1,483May predate an image-understanding fix. | — | — | — | — | — | — | — | — | — | — | — | — | — |
Reasoning Humanity's Last Exam | — | — | 67.7%With tools. | — | — | — | 64.5%With tools. | 54.9%With tools. | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
Vision Chartography | — | — | 64.4%No tools. | — | — | — | 61.6%No tools. | 15.6%No tools. | — | — | 53.6%No tools. Image-understanding fix. Anthropic's check says this score did not change. | — | — | — | — | — | — | — | — | — | — | — | — | — |
Prices checked 27 September 2026. Bench scores were read from the linked boards on 28 September 2026. A dash means that board published no score. The OSWorld rows use different task sets.