How the Claude Code CLI Harness Influences Model Behavior and Drives Cost
A frozen suite of nine coding tasks, run against every new Claude Code release, separates what the CLI changes from what the model changes.
Everyone else zigs.
A home base for the full range of AI work — model analytics, hands-on builds, and writing across the field, all in one coherent place.
24 benchmark cycles·9 models tracked·832 benchmark sessions·9 frozen tasks
Analysis
A flagship dashboard tracking how models behave and evaluate over time — real data rendered as themeable, accessible charts.
Build
Tools, projects and experiments — the things actually shipped.
New builds are on the way.
A frozen suite of nine coding tasks, run against every new Claude Code release, separates what the CLI changes from what the model changes.