AI - Anthropic release Claude Haiku 5.5
Claude Haiku 5.5 is here, less than a year after Haiku 4.5, and Anthropic is calling it their cheapest, fastest, and most capable small model yet. That sounds like a pretty significant upgrade if it holds up in real use 🤣
Haiku 5.5
Claude on X
Official announcement
Haiku 5.5 is designed for high-volume, cost-sensitive tasks.
It pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work.
What is new?
According to Anthropic, Haiku 5.5 is intended for:
- Repetitive text work such as summaries, compaction, database queries and classification.
- Speed-sensitive uses such as live customer support and browser use.
- Narrowly scoped coding subagent tasks alongside Sonnet 5.5 or Opus 5.5.
- Adjustable effort, so developers can tune the balance between cost and intelligence.
Anthropic says it costs around 75% less to run than Haiku 4.5 on average. Its API pricing starts at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens. For prompts over 100K tokens, the rates are five times higher.
The distinction Anthropic makes is important: Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding, while Haiku is aimed at narrower tasks that might otherwise be too expensive to run at scale.
The model-routing idea makes sense. You do not necessarily need the biggest model to search a codebase, summarize information or classify documents, especially when those jobs are happening at high volume. The question is whether the quality is good enough that a more capable model does not need to redo the work afterwards 🤣
I have not had a chance to try it yet, so we will have to see how much of an upgrade it is in practice 😊
Relevant X posts
Haiku 5.5 costs 75% less to run than Haiku 4.5 and is the first Haiku model with adjustable effort.
Benchmark results
Haiku 5.5 vs Opus 5:
- GDPval: 1620 vs 1596
- AA-Briefcase: 1578 vs 1562
- HLE: 45.9% vs 53%
- Terminal-Bench 4.0: 39.2% vs 46%
Those numbers are striking, but this is one person’s summary of benchmark results, not an independent evaluation of how the models compare in everyday use 🤣
Combining Haiku with larger models
Opus plans, Sonnet edits, Haiku reads, Fable reviews at decision points.
This is one proposed workflow, not a universal recipe: use Haiku for codebase exploration and research, Sonnet for edits and tests, and keep Opus for planning and review. The post argues this can avoid spending Opus tokens on work that does not need it.
It also makes the cost caveat more concrete:
Anyone still running one model for everything is paying $4 per million input tokens to check whether a file exists. Haiku 5.5 does it for $0.10.
The actual bill still depends on how much you use it and how long your prompts are. A lower token price does not automatically mean every workflow will be cheap, particularly if it relies on high volumes or long-context agent loops.
Token use can multiply with agents
Anthropic’s engineering team measured multi-agent systems at about 15x the tokens of a plain chat.
So Haiku itself is cheaper per token, but an agent swarm can consume far more tokens overall. The post also points out that Claude Code subagents inherit the main model by default, so a subagent may be billed at Opus rates unless you explicitly route it to Haiku. In other words, the savings depend on both the workflow’s token usage and which model actually handles each subtask.