AI Agent Evaluations
Performance results of AI coding agents on Next.js code generation and migration tasks, scored on success rate, average duration, and average list cost.
Agent Performance Results
Agent | |||||
|---|---|---|---|---|---|
Kimi K3 | OpenCode | 199.89s | $0.141 | 92% | 96% |
Claude Fable 5 (high) | Claude Code | 233.93s | $1.13 | 92% | 96% |
Cursor Composer 2.5 | Cursor | 149.82s | $0.046 | 92% | 96% |
GPT 5.6 Sol (ultra) | Codex | 231.83s | $0.741 | 92% | 92% |
GLM 5.2 | OpenCode | 285.17s | $0.253 | 88% | 96% |
Claude Opus 4.8 | Claude Code | 165.91s | $0.348 | 88% | 96% |
Grok 4.5 | OpenCode | 142.52s | $0.073 | 83% | 96% |
GPT 5.3 Codex (xhigh) | Codex | 178.57s | $0.233 | 83% | 92% |
GPT 5.4 (xhigh) | Codex | 227.98s | $0.238 | 83% | 88% |
GPT 5.5 Pro | Codex | 771.63s | $18.21 | 83% | 83% |
Kimi K2.7 Code | OpenCode | 207.72s | $0.060 | 75% | 96% |
MiniMax M3 | OpenCode | 181.88s | $0.032 | 75% | 96% |
GLM 5.1 | OpenCode | 253.24s | $0.077 | 75% | 96% |
Claude Opus 4.7 (max) | Claude Code | 142.48s | $0.445 | 75% | 96% |
Claude Opus 4.6 | Claude Code | 189.27s | $0.204 | 75% | 96% |
Gemini 3.1 Pro Preview | Gemini CLI | 247.45s | $0.110 | 75% | 92% |
Cursor Composer 2.0 | Cursor | 115.52s | $0.038 | 75% | 92% |
Kimi K2.6 | OpenCode | 197.69s | $0.057 | 67% | 96% |
Cursor Composer 1.5 | Cursor | 119.20s | N/A | 67% | 88% |
Gemini 3.0 Pro Preview | Gemini CLI | 260.37s | $0.165 | 67% | 83% |
Claude Sonnet 4.6 | Claude Code | 157.50s | $0.140 | 58% | 96% |
GPT 5.2 Codex (xhigh) | Codex | 149.60s | $0.150 | 58% | 79% |
Claude Sonnet 4.5 | Claude Code | 150.53s | $0.130 | 50% | 83% |
MiniMax M2.7 | OpenCode | 293.14s | $0.026 | 50% | 63% |
Kimi K2.5 | OpenCode | 134.99s | $0.018 | 21% | 58% |
* AGENTS.md provides bundled Next.js documentation for AI coding agents. The column shows additional evals that passed when agents had access to this documentation.
N/A marks evals a model has never run (newly added evals). They count against the success rate so all models are scored over the same eval set. Reworked evals keep a model's last result until it is rerun.
Avg List Cost is the mean cost per eval, estimated from the tokens each run used at the provider's public list price. It is a relative guide, not a bill: your rate depends on caching, discounts, and subscription plans. Avg List Cost shows N/A when a model has no published token price or its runs did not record token usage.