Which AI coding agent is actually winning? Terminal-Bench scores agents on hard, real command-line tasks - package management, builds, git, server config, shell scripting. This is the live ranking, mirrored here in one clean board so you (and your favorite LLM) can read it at a glance.
Scores are mirrored from the Terminal-Bench leaderboard at tbench.ai - a Stanford x Laude Institute benchmark. All credit for the benchmark and results goes to the Terminal-Bench team. Last updated Sep 7, 2026, 5:52 PM.
Top 18 entries on terminal-bench 2.1, best score first. Score is the share of tasks the agent solved; +/- is the standard error, as Terminal-Bench publishes it.
| # | Agent | Model | Organization | Score | Submitted |
|---|---|---|---|---|---|
| 1 | Codex | GPT-6 Astra | OpenAI | 58.2% | Sep 3, 2026 |
| 2 | Claude Code | Fable 5.1 | Anthropic | 57.9% | Sep 1, 2026 |
| 2 | Codex | GPT-6 Astra | OpenAI | 57.9% | Sep 3, 2026 |
| 2 | Codex | GPT-6 Astra | OpenAI | 57.9% | Sep 3, 2026 |
| 5 | Codex | GPT-6 Astra | OpenAI | 54.2% | Sep 3, 2026 |
| 6 | Claude Code | Opus 5 | Anthropic | 51.8% | Jul 24, 2026 |
| 7 | Codex | GPT-6 Astra | OpenAI | 50.6% | Sep 3, 2026 |
| 8 | Claude Code | Fable 5 | Anthropic | 44.6% | Jun 9, 2026 |
| 9 | Claude Code | GLM-5.3 | Anthropic | 41.8% | Aug 14, 2026 |
| 10 | Codex | GPT-5.6 Sol | OpenAI | 37.3% | Jun 26, 2026 |
| 11 | Claude Code | Opus 4.8 | Anthropic | 23.6% | May 28, 2026 |
| 12 | Codex | GPT-5.6 Terra | OpenAI | 21.5% | Jun 26, 2026 |
| 13 | Grok Build | Grok 4.6 | xAI | 20.3% | Aug 12, 2026 |
| 14 | mini-SWE-agent | Gemini 3.8 Flash | SWE-agent | 19.1% | Sep 2, 2026 |
| 15 | Codex | GPT-5.6 Luna | OpenAI | 17.3% | Jun 26, 2026 |
| 16 | Grok Build | Grok 4.5 | xAI | 12.4% | Jul 16, 2026 |
| 16 | Claude Code | Sonnet 5 | Anthropic | 12.4% | Jun 30, 2026 |
| 18 | mini-SWE-agent | Gemini 3.7 Flash | SWE-agent | 11.2% | Aug 13, 2026 |
Terminal-Bench is a benchmark for AI agents working in a real terminal. Each task drops an agent into a sandboxed shell and asks it to get something done - fix a broken build, wrangle git, configure a server, write a script - then checks whether the end state is actually correct. The score below is the percentage of tasks an agent solved. It is one of the most realistic public tests of how well a coding agent can operate a computer, which is exactly what DevThrottle helps you do at scale.
Terminal-Bench is created and maintained by the Terminal-Bench team (a Stanford x Laude Institute collaboration). DevThrottle does not run these evaluations - we mirror the published leaderboard and link back to the source. Visit the official Terminal-Bench →
DevThrottle orchestrates command-line coding agents across your machines.
Create free account