Claude Benchmarks

Models
Effort

7 lines shown

Suite cost (USD)

  • sonnet-5 (high)
  • fable-5 (high)
  • haiku-4-5
  • opus-4-8 (high)
  • opus-5 (high)
  • fable-5-1 (high)
  • opus-5-5 (high)
Suite cost (USD)
Data table
Claude Code CLI versionsonnet-5 (high)fable-5 (high)haiku-4-5opus-4-8 (high)sonnet-5 (medium)opus-5 (high)opus-5 (low)opus-5 (max)opus-5 (medium)opus-5 (xhigh)fable-5 (medium)opus-4-8 (medium)fable-5 (max)fable-5 (xhigh)opus-4-8 (max)opus-4-8 (xhigh)sonnet-5 (max)sonnet-5 (xhigh)fable-5-1 (high)fable-5-1 (medium)fable-5-1 (max)fable-5-1 (xhigh)opus-5-5 (high)opus-5-5 (medium)
2.1.1960.39959———————————————————————
2.1.2020.3074411.5854730.1386590.7877060.292861———————————————————
2.1.2040.435911.5893450.1369380.74730.423905———————————————————
2.1.2050.4019361.4129470.1370070.684710.399699———————————————————
2.1.2060.4158581.4523280.1340950.7118290.411794———————————————————
2.1.2070.3131611.1691420.1517350.4944310.292131———————————————————
2.1.2080.3015291.0528040.1372170.4733210.271329———————————————————
2.1.2100.2823381.0525050.1358790.4788240.261952———————————————————
2.1.2110.3043121.0425010.1351690.4413840.257755———————————————————
2.1.2120.3049671.0375620.1318740.4490240.266073———————————————————
2.1.2140.2982631.003270.1409860.4706040.270236———————————————————
2.1.2150.2896740.989780.1292310.4033630.265379———————————————————
2.1.2160.2698610.9324450.1435140.4199730.250677———————————————————
2.1.2170.293840.9706180.1380430.4054170.256946———————————————————
2.1.2180.2842980.9935260.1284580.4029940.265344———————————————————
2.1.2190.2864481.0402090.1328450.4081530.261634———————————————————
2.1.2200.2768640.9874170.1255540.401210.2644210.6578460.4326750.8150550.5661230.762884——————————————
2.1.2210.2860490.9773720.1319990.401626—0.65512—0.873209————————————————
2.1.2220.290281.0981880.1319290.4436370.2685450.646637——0.56267—0.8922970.380577————————————
2.1.2230.2960990.9979150.1360630.4195630.263880.654669——0.563948—1.0194590.417344————————————
2.1.2240.2921221.0852810.1318630.4418560.2649690.662316——0.574012—0.9313410.397382————————————
2.1.2250.294781.0847270.1367270.4247060.262170.68277——0.582747—0.9453280.392409————————————
2.1.2260.2989911.0595510.1430380.4353420.2621720.688851——0.593155—0.9245630.423615————————————
2.1.2270.2964321.0745690.1320820.4746790.2600690.72558——0.567964—0.9923870.43579————————————
2.1.2280.3094031.069730.1336590.4221490.2533760.679035——0.573721—0.9122150.413112————————————
2.1.2290.302241.0693660.1361530.4526090.2597790.683558——0.564781—0.9091370.424274————————————
2.1.2310.3016371.0460280.1349290.4428980.2613130.709022——0.574353—0.9463310.407458————————————
2.1.2320.2967411.0461690.1269880.4481570.2633670.636067——0.593776—0.9346240.429926————————————
2.1.2330.2994941.0393480.1310540.3970760.2635680.660987——0.608725—0.8930370.423469————————————
2.1.2340.2838711.0497890.1236640.4156980.2677990.627984——0.55549—0.8841930.415903————————————
2.1.2350.289561.0271510.1282070.4508770.2511890.621802——0.533325—0.8820080.436997————————————
2.1.2360.2875150.9571250.1329530.4545120.2732480.631873——0.516582—0.9239760.445696————————————
2.1.2370.2959281.0621910.1316690.4293560.2542270.624628——0.512943—0.9230610.42662————————————
2.1.2380.2739720.9590720.1351260.4184910.266260.620326——0.505544—0.9027590.421965————————————
2.1.2390.2770271.0058880.1322680.440070.2569250.668432——0.521478—0.8630740.412674————————————
2.1.2400.2966051.0500340.1294660.464848—0.627648——————————————————
2.1.2410.2848861.0196150.1347390.4329110.2644020.634396——0.495076—0.891140.440323————————————
2.1.2430.2813471.0657040.1287770.4390550.2587680.639663——0.510795—0.8950020.418189————————————
2.1.2450.2776181.0786320.1276990.4374610.2751640.62811——0.506187—0.9388240.4392————————————
2.1.2460.2929351.0263440.1283290.4304040.2625550.598752——0.512927—0.901360.445008————————————
2.1.2470.2746491.0693610.128530.4528960.2620640.612011—0.8818190.5096410.7594010.9444050.4315131.4688292.0450820.7101670.5495910.4272450.319708——————
2.1.2480.2419690.8507890.1160720.3722710.2210590.537779—0.8072330.4456080.6747770.7865640.3799492.8427551.9037060.6623010.4578670.3874190.280323——————
2.1.2500.2433620.9459360.1199910.3729470.226320.55282—0.8285570.4480030.6961820.7608130.3543012.8294691.8714190.6509260.457870.3540360.281948——————
2.1.2510.2505430.8707220.1180750.3855710.2351050.53371—0.7870980.4154210.723680.7844720.3599073.0176971.8570150.6708610.4339310.368420.282411——————
2.1.2520.250620.8914680.121860.3729190.2360970.527568——0.405826—0.7754540.361293————————————
2.1.2570.256590.871960.1149440.3804570.2252490.525918——0.439004—0.7387260.363476————————————
2.1.2580.2471040.8190510.118790.383242—0.535461——————————————————
2.1.2590.2482020.8399040.1180030.3863570.2331960.526803——0.430426—0.7473310.360691————————————
2.1.2600.2480780.8362410.1174410.3671970.2238460.554445——0.415243—0.7873310.365427————————————
2.1.2610.2424460.8676810.12130.370430.2223790.554465——0.44107—0.746020.368507——————0.561440.466341————
2.1.2630.2465180.9285230.120430.3830420.2290810.580929——0.412345—0.7176380.366587——————0.5733910.454946————
2.1.2650.2518230.8802320.1219110.3864370.2277460.601206——0.429159—0.8067860.345241——————0.5429450.468057————
2.1.2660.2515990.8250640.1182250.3800740.2265310.582992——0.439395—0.7902820.358274——————0.5311190.47197————
2.1.2670.2601240.9137930.1238040.3804950.2235340.584541——0.463294—0.7566740.359009——————0.5362310.471379————
2.1.2680.2475740.8471960.1206010.4025120.2333320.572696——0.467918—0.748910.371418——————0.5718520.479094————
2.1.2690.2475320.8831610.1254960.376190.2258840.595832——0.46909—0.7923190.367078——————0.5494460.480273————
2.1.2700.2480770.8722740.1206720.363450.2177740.606586—0.7883270.4498360.693980.7790550.3839921.3029870.9709250.6123350.4617920.3590380.2803610.5811840.4766682.0297291.025853——
2.1.2710.2487050.8750750.1170430.3797880.2287410.607335——0.468559—0.781330.364044——————0.6026640.471878————
2.1.272———0.401459—0.605811——————————————————
2.1.2800.2522070.8822130.1163820.3653320.2308380.602693——0.439992—0.7506660.373698——————0.5720170.47293——0.2623470.215071
2.1.2810.24830.8668070.1209140.3752790.2231120.622124——0.478544—0.71620.374442——————0.5839110.470826——0.2708790.2318
2.1.2820.241220.8856780.119630.3816610.223760.622266——0.483949—0.7233190.361705——————0.6094780.481824——0.2763170.225929
highnightly-2.1.282

Data table
Metricfable-5 (high)fable-5-1 (high)haiku-4-5opus-4-8 (high)opus-5 (high)opus-5-5 (high)sonnet-5 (high)
Cost per session (USD)mean0.1051180.0770570.0144060.0481720.0739660.0343940.029528
Cost per session (USD)median0.1106450.0784770.0156740.0538940.07540.0302880.032776
Output tokens493564692460562452442
Input tokens666266868
Cache-read tokens504365180262280463976271042956110829
Cache-write tokens7138829827171200806843
Input + cache tokens509455277562999465686714546887111439
API calls3333434
Tool calls2222323
Thinking blocks1030101
Latency (s, wall-clock)10.57710.4158.2869.22210.6847.5066.945
Suite cost (USD)No published column carried Suite cost (USD) for this cycle.
mediumnightly-medium-2.1.282

Data table
Metricfable-5 (medium)fable-5-1 (medium)opus-4-8 (medium)opus-5 (medium)opus-5-5 (medium)sonnet-5 (medium)
Cost per session (USD)mean0.0930720.0646480.0456810.0586920.0299030.027036
Cost per session (USD)median0.0981570.0625810.050650.0564840.0247240.028805
Output tokens385509459455448354
Input tokens6666668
Cache-read tokens5004848471461284923346274107496
Cache-write tokens2598152291007467745
Input + cache tokens5040452466465685007846773111409
API calls333334
Tool calls222223
Thinking blocks000001
Latency (s, wall-clock)9.1539.1529.6029.4286.0756.74
Suite cost (USD)No published column carried Suite cost (USD) for this cycle.

Claude Code CLI 2.1.282 Analysis

Measured 2026-09-25 21:01 UTC from this cycle’s benchmark runs, written up from those measurements alone, then checked by an independent reviewer. Where the wording and the tables disagree, the tables are correct.

Verdict

At the high reasoning setting, one task costs this much per session right now, cheapest to most expensive:

  • haiku-4-5 $0.0144
  • sonnet-5 $0.0295
  • opus-5-5 $0.0344
  • opus-4-8 $0.0482
  • opus-5 $0.0740
  • fable-5-1 $0.0771
  • fable-5 $0.1051

At the medium setting:

  • sonnet-5 $0.0270
  • opus-5-5 $0.0299
  • opus-4-8 $0.0457
  • opus-5 $0.0587
  • fable-5-1 $0.0646
  • fable-5 $0.0931

Across the last 10 cycles at the high setting, each family moved like this, and stayed inside these dollar bands:

  • fable-5: no clear direction, $0.1035 to $0.1139
  • fable-5-1: drifted up, $0.0707 to $0.0771
  • haiku-4-5: no clear direction, $0.0138 to $0.0148
  • opus-4-8: no clear direction, $0.0460 to $0.0490
  • opus-5: drifted up, $0.0667 to $0.0740
  • opus-5-5: too few cycles to say, $0.0296 to $0.0344
  • sonnet-5: no clear direction, $0.0275 to $0.0314

Across the last 10 cycles at the medium setting:

  • fable-5: no clear direction, $0.0909 to $0.0987
  • fable-5-1: no clear direction, $0.0626 to $0.0646
  • opus-4-8: no clear direction, $0.0435 to $0.0469
  • opus-5: no clear direction, $0.0527 to $0.0604
  • opus-5-5: too few cycles to say, $0.0242 to $0.0305
  • sonnet-5: no clear direction, $0.0267 to $0.0282

This cycle moved from the previous CLI version (2.1.281) to the current one (2.1.282), and the families with a same comparison moved very little:

  • haiku at the high setting rose 0.8%
  • sonnet at the high setting fell 1.2%
  • sonnet at the medium setting fell 0.8%

In all three cases the individual runs behind those averages overlap heavily, so none of the moves stands out from normal run-to-run variation. Something may have shifted; this measurement cannot see it.

The spread between families is far larger than anything that moved this cycle: at the high setting, haiku-4-5 costs 86.3% less per task than fable-5.

Which cost figure is which – setting high

The same sessions summarized three ways. They disagree on purpose, and none is the corrected version of another.

  • Mean $/session: multiplies out to real spend, which makes it the budgeting number, and it is what the Cost Trends chart plots.
  • Median $/session: shrugs off an unusually cheap or expensive run, but it does not sum to real spend.
  • Suite cost: the price of one full pass over every prompt in the catalog, the per-prompt medians added up, so it answers what a whole run costs rather than what one session costs.
family models compared mean $/sess mean delta % median $/sess median delta % suite $ suite delta % prompts
haiku haiku-4-5 0.0143 -> 0.0144 +0.8% 0.0141 -> 0.0157 +10.8% 0.1209 -> 0.1196 -1.1% 9
sonnet sonnet-5 0.0299 -> 0.0295 -1.2% 0.0327 -> 0.0328 +0.2% 0.2483 -> 0.2412 -2.9% 9

Suite cost uses only the prompts both cycles ran, so the two sides always add up the same catalog.

Which cost figure is which – setting medium

The same sessions summarized three ways. They disagree on purpose, and none is the corrected version of another.

  • Mean $/session: multiplies out to real spend, which makes it the budgeting number, and it is what the Cost Trends chart plots.
  • Median $/session: shrugs off an unusually cheap or expensive run, but it does not sum to real spend.
  • Suite cost: the price of one full pass over every prompt in the catalog, the per-prompt medians added up, so it answers what a whole run costs rather than what one session costs.
family models compared mean $/sess mean delta % median $/sess median delta % suite $ suite delta % prompts
sonnet sonnet-5 0.0273 -> 0.0270 -0.8% 0.0316 -> 0.0288 -8.8% 0.2231 -> 0.2238 +0.3% 9

Suite cost uses only the prompts both cycles ran, so the two sides always add up the same catalog.

What these terms mean

New to this report? Here is the vocabulary used below, in plain words. Skip ahead if you already know it.

  • Session (also “task”): one complete run of one prompt by one model, start to finish. Every cost in this report is per session.
  • Family: a model line, such as opus or sonnet. One family can cover more than one version, for example opus-4-8 and opus-5.
  • Cache, warm and cold: the tool reuses recent context instead of sending it again, which is far cheaper. A session that gets to reuse it is “warm”; one that has to send everything fresh is “cold”. Whether a given session lands warm or cold is partly luck, and that luck moves the cost.
  • Cache luck (or cache-hit luck): the cost swing caused by how many sessions happened to land warm rather than cold, rather than by any real change in the tool or the model.
  • Blended cost: cost measured across every session, warm and cold alike. This is the real money figure, and it is what the cost feed shows. “Blended” says which sessions are counted, not how they are averaged; see “Which cost figure is which” above for the averaging.
  • Warm-only cost: cost measured across warm sessions only. Removing the cold ones removes most of the luck, so this is the fair test of whether cost actually changed.
  • Pooled: all sessions for a family added together into a single number, rather than split out per prompt.
  • Mean cost per session: the ordinary average across every session. It multiplies out to real spend, so it is the budgeting figure and the one this report leads with.
  • Median cost per session: the middle session once they are lined up by cost. It ignores an unusually cheap or expensive run, but it does not add up to real spend.
  • Suite cost: the price of one full pass over every prompt in the catalog, found by taking each prompt’s median and adding those up. It is a total for the whole catalog, not an average per session, so it is a larger number than either figure above and is not comparable to them.
  • Cell: one family and one prompt combination, for example haiku on the rename-refactor prompt. Cells hold few sessions, so a single cell is noisy.
  • n: how many sessions sit behind a number. A larger n is more trustworthy.
  • p-value, and “significant”: the same prompt run twice costs different amounts, so a p-value asks whether two sets of runs can be told apart at all. Line up every run from both cycles: if one cycle’s runs sit consistently higher, p is near 0 and the difference is called “significant”; if the runs are mixed together, p is near 1 and the two sets look like the same thing measured twice. Precisely, it is the chance of seeing a gap at least this lopsided if the two cycles were truly identical. A high p is not weak evidence of a change, it is no evidence either way.
  • API-equivalent cost: dollars calculated at published API rates. It is not necessarily dollars billed, because a subscription plan may already cover the usage.

Pooled blended comparison

This is the real per-task spend: every session counted, including the unlucky ones that start with a cold cache. It is the mean dollars per session a user actually pays, and it is the figure the cost feed carries. Each setting is compared against its own previous run at the same setting.

Setting high, nightly-2.1.281 to nightly-2.1.282:

family models compared n (prev -> cur) prev $/sess cur $/sess delta % Mann-Whitney p
haiku haiku-4-5 53 -> 53 0.0143 0.0144 +0.8% 0.919
sonnet sonnet-5 53 -> 53 0.0299 0.0295 -1.2% 0.929

Setting medium, nightly-medium-2.1.281 to nightly-medium-2.1.282:

family models compared n (prev -> cur) prev $/sess cur $/sess delta % Mann-Whitney p
sonnet sonnet-5 53 -> 53 0.0273 0.0270 -0.8% 0.874

Only families with a matching run in the previous cycle at the same setting can appear here, which is why the two settings list different families.

Pooled warm comparison (significance control)

Whether a change is believable is a separate question from what it cost. The warm view drops the sessions that had to build their cache from scratch, so a lucky or unlucky draw of cold starts cannot masquerade as a change in price. Comparing the individual warm run costs from each cycle, every p-value below sits far above the 0.05 line, so none of this cycle’s moves can be told apart from ordinary run-to-run variation.

Setting high:

family models compared n (prev -> cur) prev $/sess cur $/sess delta % Mann-Whitney p
haiku haiku-4-5 45 -> 45 0.0132 0.0134 +2.1% 0.878
sonnet sonnet-5 45 -> 45 0.0273 0.0267 -2.1% 0.884

Setting medium:

family models compared n (prev -> cur) prev $/sess cur $/sess delta % Mann-Whitney p
sonnet sonnet-5 45 -> 45 0.0247 0.0245 -1.0% 0.865

What the reasoning setting costs

The same release was measured twice, once at each reasoning setting. This is the only place the two are held against each other. The change column measures medium against high, so a negative percentage means medium is the cheaper of the two, and the comparison is of the individual session costs from each setting.

model high mean $ medium mean $ change p n
fable-5 0.1051 0.0931 -11.5% 0.111 53/53
fable-5-1 0.0771 0.0646 -16.1% 0.199 53/53
opus-4-8 0.0482 0.0457 -5.2% 0.460 53/53
opus-5 0.0740 0.0587 -20.6% 0.050 57/53
opus-5-5 0.0344 0.0299 -13.1% 0.374 53/53
sonnet-5 0.0295 0.0270 -8.4% 0.250 53/53

Medium is the cheaper setting for all six models measured both ways:

  • fable-5 is 11.5% cheaper at medium
  • fable-5-1 is 16.1% cheaper at medium
  • opus-4-8 is 5.2% cheaper at medium
  • opus-5 is 20.6% cheaper at medium
  • opus-5-5 is 13.1% cheaper at medium
  • sonnet-5 is 8.4% cheaper at medium

None of those six differences clears the 0.05 line, including opus-5 at p=0.050, so for every model here the gap between the two settings cannot be told apart from run-to-run variation. The consistent direction across all six is worth watching, but this cycle’s data cannot confirm it.

haiku-4-5 was measured at the high setting only, so it has no cross-setting figure.

Notable cells (secondary)

At the high setting, one prompt moved more than the rest: sonnet on rename-refactor, which asks the model to rename a function everywhere it appears, covering its definition, its export, and every call site across several files. The figures behind that cell:

  • warm cost fell 18.5%
  • 6 runs before and 6 runs after
  • p=0.347, well above the 0.05 line

This is one cell out of 18 compared at once, and with that many simultaneous comparisons roughly one in twenty will look striking by chance alone, so a single cell is never a finding on its own. Treat it as something to check again next cycle. At the medium setting, no cell moved beyond the noise threshold.

Caveats

  • Two views of cost are reported, and they answer different questions. The headline figure is the real per-task spend across every session, cache luck included, and it matches the dashboard. The warm-only figure strips out that random cache luck and is the control for whether a change is real. Judge “did cost change” by the warm comparison, not by the size of the headline move: a headline move with no significant warm move may simply be cache luck.
  • The two cycles being compared ran on different CLI versions (2.1.281 against 2.1.282). Measuring that version-to-version effect is the point of the benchmark, but it also means the difference cannot be pinned on the model itself without a separate test that holds the CLI version fixed.
  • The per-cell movers are one lens over 18 comparisons made at the same time. At the usual 0.05 threshold, roughly one in twenty such comparisons will look significant purely by chance, so a single standout cell is not a result.
  • Token capture for Haiku is known to be incomplete, so every Haiku cost and token figure should be read as a floor rather than an exact amount.
  • Timings are wall-clock and include the time the model spends thinking, and the costs here are calculated API-equivalent amounts rather than dollars actually billed.

Verification appendix

  • Verifier verdict: PASS after 2 round(s).
  • A first draft was rejected and revised; the objections were resolved.