Highlights

Model Analytics Highlights

Distilled 2026-09-25 21:04 UTC from the published section write-ups alone, then checked by an independent reviewer. Each section’s own page carries the detail and the tables behind these figures.

What this shows

The Model Analytics area collects standing cost studies of coding CLIs. Each section is its own study, with its own prompt catalog and its own rate card, tracking what one task costs and whether that cost is moving between releases. The sections do not share a method or a prompt set, so each entry below carries its own figures and its own limits.

Claude Benchmarks

What a unit of work costs on the Claude Code CLI: one frozen prompt catalog run repeatedly, priced at the provider’s API list rates, and reported as mean dollars per session at the high and medium reasoning settings. At the high setting the section reports what one task costs now, and how far apart the families sit:

  • haiku-4-5 costs $0.0144 per session
  • fable-5 costs $0.1051 per session
  • haiku-4-5 costs 86.3% less per task than fable-5

Across the last 10 cycles at the high setting, fable-5-1 and opus-5 drifted up, fable-5, haiku-4-5, opus-4-8 and sonnet-5 show no clear direction, and opus-5-5 has too few cycles to say; across the last 10 cycles at the medium setting, fable-5, fable-5-1, opus-4-8, opus-5 and sonnet-5 show no clear direction, and opus-5-5 again has too few cycles to say. Moving from the previous CLI version to the current one, sonnet-5 fell 1.2% at the high setting and fell 0.8% at the medium setting; in the section’s warm comparison, comparing the individual warm run costs from each cycle, every p-value sits far above the 0.05 line, so none of this cycle’s moves can be told apart from ordinary run-to-run variation, and something may have shifted that this measurement cannot see. Token capture for Haiku is known to be incomplete, so every Haiku cost figure should be read as a floor rather than an exact amount.

Codex Benchmarks

What a unit of work costs on the Codex CLI, built the same way as the Claude study but on its own fixed set of prompts, priced at OpenAI’s published list rates; because the prompts are different prompts, the section says the two sets of results should not be read as a like for like comparison. High and medium reasoning are reported separately and are not interchangeable, since the same model at a different setting is a different cost. The section reports sol as the dearest model at both settings, and at high reasoning gives these figures:

  • luna costs $0.0032 per session
  • sol costs $0.0557 per session
  • luna costs 94.2% less than sol

Across the last 10 cycles at high reasoning, sol drifted down inside $0.0528 to $0.0834, while luna and terra show no clear direction; across the last 10 cycles at medium reasoning, luna, sol and terra all show no clear direction. Against the previous CLI version, terra rose 8.5% at high reasoning and fell 2.3% at medium reasoning; comparing the individual session costs of the two cycles, none of the six version-to-version moves can be told apart from ordinary run-to-run variation, and that test is not controlled for cache state, because the study cannot tell which sessions started with a warm cache, which makes it weaker evidence than the same test run on warm sessions alone.


Verification appendix

  • Verifier verdict: PASS after 3 round(s).
  • A first draft was rejected and revised; the objections were resolved.