Three sources in one place, and one answer to the only question that matters at the start of a task: which agent do I run. Cost per task from CursorBench, quality drift from AI Stupid Level, availability and price from GitHub Copilot. The effort level is always named — without it a benchmark number means nothing.
Cyan — the numbers line up. Take it.
Violet — a gap. You are paying a lot for a little.
Effort — always named on the model, never implied.
Cyan means the price gap is small next to the quality gap — take the dearer model, even for simple work. Violet means you are paying a lot for very little — stay on the cheaper one.
The jobs that actually land here. On the right, the model and the effort level you start with — so you do not have to compare benchmarks yourself.
| Task | Role | Start with | Score | $ / task |
|---|
Every dot is a model at one effort level: average cost per task across (log scale), CursorBench score up. The line joins the models nothing else beats on price and score at the same time.
The same model under the same name can simply get worse. The sparkline shows individual benchmark runs from the last 7 days — thin line raw, thick line smoothed. Single runs swing hard, so Δ7d compares the average of the first fifth of the window with the average of the last fifth. Effort is not published here — these runs use each provider's default setting.
| Model | Score | 7 days | Δ7d | min–max | Trend |
|---|
Prices per million tokens from GitHub's documentation — this is the part that comes out of the budget. Models without a CursorBench score are not benchmarked end to end: their price is known, their quality is not.
| Model | Category | Input / 1M | Cached / 1M | Output / 1M | CursorBench |
|---|