blinkneuron

Kvasir? Kvasir — in Norse myth, a being made out of the mixed knowledge of two warring tribes of gods: the wisest thing alive, and the one you went to for counsel. Later killed for it, and brewed into the mead of poetry. This page mixes three sources and answers one question. Same idea, fewer casualties.

Three sources in one place, and one answer to the only question that matters at the start of a task: which agent do I run. Cost per task from CursorBench, quality drift from AI Stupid Level, availability and price from GitHub Copilot. The effort level is always named — without it a benchmark number means nothing.

How to read it

Cyan — the numbers line up. Take it.

Violet — a gap. You are paying a lot for a little.

Effort — always named on the model, never implied.

Today's verdict

Three roles, three models

Is the upgrade worth it

The distance between roles

Cyan means the price gap is small next to the quality gap — take the dearer model, even for simple work. Violet means you are paying a lot for very little — stay on the cheaper one.

Quick pick

Task → agent

The jobs that actually land here. On the right, the model and the effort level you start with — so you do not have to compare benchmarks yourself.

TaskRoleStart with Score$ / task
Cost vs quality

Where the value sits

Every dot is a model at one effort level: average cost per task across (log scale), CursorBench score up. The line joins the models nothing else beats on price and score at the same time.

Architect Worker Scout everything else
AI Stupid Level

Model drift over time

The same model under the same name can simply get worse. The sparkline shows individual benchmark runs from the last 7 days — thin line raw, thick line smoothed. Single runs swing hard, so Δ7d compares the average of the first fifth of the window with the average of the last fifth. Effort is not published here — these runs use each provider's default setting.

ModelScore7 days Δ7dmin–maxTrend
GitHub Copilot

What we have at work, and what it costs

Prices per million tokens from GitHub's documentation — this is the part that comes out of the budget. Models without a CursorBench score are not benchmarked end to end: their price is known, their quality is not.

ModelCategoryInput / 1M Cached / 1MOutput / 1MCursorBench