Claude Sonnet 5.5: Benchmarks, Pricing and When to Use It
Anthropic released Claude Sonnet 5.5 on September 28, 2026. The price did not move: $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. On Anthropic's own benchmarks it comes within a few points of Claude Opus 5.5, which costs twice as much. That makes it the natural default for teams running AI agents in production. Two things get in the way. Five breaking changes mean you can't just swap the model ID. And an independent test found it can cost more per task than Sonnet 5.
What is Claude Sonnet 5.5, in one sentence?
Claude Sonnet 5.5 is Anthropic's mid-priced model, released September 28, 2026, that scores close to Claude Opus 5.5 at half the per-token price.
It is the second model in the Claude 5.5 family, after Opus 5.5. Sonnet 5 is now listed as legacy. The main specs, from the model documentation:
- Context window: 1M tokens, with up to 128K output tokens per request
- Knowledge cutoff: June 2026 (Sonnet 5 stops at January 2026)
- Input: text and images
- Model ID:
claude-sonnet-5-5, oranthropic.claude-sonnet-5-5on Amazon Bedrock - Available until: at least September 28, 2027
How does Sonnet 5.5 compare with Sonnet 5 and Opus 5.5?
The benchmark rows are Anthropic's own results from the launch announcement. Specs and prices come from the model documentation.
| Claude Sonnet 5 | Claude Sonnet 5.5 | Claude Opus 5.5 | |
|---|---|---|---|
| Price per million tokens (input / output) | $2 / $10 | $2 / $10 | $4 / $20 |
| Context window / max output | 1M / 128K | 1M / 128K | 1M / 128K |
| Knowledge cutoff | Jan 2026 | Jun 2026 | Jun 2026 |
| Terminal-Bench 4.0 (agentic coding) | 10.3% | 70.6% | 66.4% |
| OSWorld 2.1 (computer use) | 57.0% | 80.1% | 81.8% |
| GDPval-AA v2.1 (knowledge work, Elo) | 1449 | 1844 | 1846 |
| Default effort on the API | high | high | medium |
The Terminal-Bench number is the one to question. A jump from 10.3% to 70.6% in a half-step release is unusual, and Anthropic ran the test itself. Check it against your own tasks before you plan around it. On knowledge work and computer use, Sonnet 5.5 lands within two points of Opus 5.5.
Independent testing mostly agrees. Artificial Analysis puts Sonnet 5.5 second on its Intelligence Index at maximum effort. It scored 56, two points behind Opus 5.5. On three of Artificial Analysis's work benchmarks, including GDPval-AA, the two models were roughly even.
What breaks when you move from Sonnet 5?
Changing the model ID is only the first step. Anthropic's what's new page lists five breaking changes. Several hit patterns that enterprise agent loops depend on.
- Forced tool use is gone. Setting
tool_choiceto "any" or to a named tool now returns a 400 error. Extraction pipelines often force a tool call to get structured output. Anthropic's replacement is automatic tool choice with strict tool use, or structured outputs. - The "disabled" thinking setting is gone. Requests that send it get a 400 error. To turn off up-front thinking, use
between_tools, which only works at high effort or below. - Thinking blocks are bound to the model and the conversation. No other model can read Sonnet 5.5's reasoning. Accounts created on or after August 31, 2026 also get a 400 error when a request replays that reasoning after earlier messages were edited. Keep conversations append-only.
- The older computer use tool is rejected on the Claude API and Google Cloud. Integrations have to move to the new computer use toolset. Bedrock still accepts the old tool.
- Some advisor pairings fail. Sonnet 5 and older Opus models can no longer advise a Sonnet 5.5 executor.
One more change fails silently. The notes the model writes between tool calls now arrive as thinking blocks. With the default display setting, their text is empty. A product that shows those notes to users will go quiet between steps.
Effort levels were also recalibrated. Medium effort on Sonnet 5.5 thinks less than medium effort did on Sonnet 5. Anthropic suggests testing effort again from scratch. Its starting point for well-specified agent tasks is medium, moving to high for harder ones.
Is Sonnet 5.5 actually cheaper? Read the fine print
The list price is unchanged, so any saving has to come from using fewer tokens. Anthropic says Sonnet 5.5 "costs up to 30% less per task" than Sonnet 5. It also says output is more than 30% faster. Customer quotes on the launch page agree. Slack reports about 14% fewer output tokens. Zendesk says tickets were processed 20% faster. Anthropic picked these quotes, so treat them as vendor examples.
Artificial Analysis found the opposite at maximum effort. Sonnet 5.5 used about 193,000 output tokens per task on its index. That is roughly 60% more than Opus 5.5 or Sonnet 5 at the same setting. Its cost per task came to $7.60, about 50% higher than Sonnet 5's. On Terminal-Bench 4.0 it scored 64%, below the 70.6% Anthropic reported.
Both can be true. Anthropic's claim says "up to" and "for most work." Artificial Analysis ran the model at maximum effort, where it thinks longest. Your cost per task depends on the effort setting and the job. Measure it on your own work before you count on a saving.
Knowledge is where Sonnet 5.5 falls behind. On the AA-Omniscience test it answered 54% of questions correctly. Opus 5.5 answered 66%. Sonnet 5.5 did have a lower hallucination rate, 47% against 59%. If a step depends on what the model already knows, Opus 5.5 is the safer pick. If you supply the documents, the gap matters less.
Where does Sonnet 5.5 fit in an enterprise agent workflow?
Anthropic pitches Sonnet 5.5 at well-scoped everyday work and Opus 5.5 at complex tasks that need careful judgment. The independent numbers back that split. In practice it looks like this:
- Sonnet 5.5 as the default for clearly specified multi-step work, such as pulling records or preparing a report from data you provide.
- Opus 5.5 for steps that need broad knowledge or open-ended judgment.
- Claude Haiku 4.5 ($1 / $5 per million tokens, 200K context) for high-volume sorting and routing.
The thinking-block change shapes this design. If a task moves from Sonnet 5.5 to Opus 5.5 halfway, Opus cannot see what Sonnet reasoned. Pass state between steps as explicit, structured context. Don't count on one shared conversation history.
Refusals need a path too. Sonnet 5.5 flags a declined request as a refusal and names the category. Anthropic's server-side fallback retries some categories on Sonnet 5. An agent that reads a refusal as an empty answer will skip work without telling anyone. Send refusals to a person or a fallback step.
The bottom line
If you run Sonnet 5 today, upgrade. The price is the same and the model is much stronger on Anthropic's tests. Plan it as a migration. Fix the breaking changes first, then retest effort settings and measure cost per task on real work. Keep Opus 5.5 for steps that lean on the model's own knowledge.
How do you keep agent memory when you switch models?
Sonnet 5.5's thinking blocks are bound to the model and the conversation. A hand-off to Opus 5.5 drops them. On accounts created on or after August 31, 2026, editing earlier messages can also get a request rejected. Any reasoning the agent did not put into a message or a file is lost. The same thing happens on a larger scale when a session ends or you switch from Claude Code to Codex: the next session starts with only what was written down.
memU keeps what the agent concludes, written as a shared wiki of Markdown files that any model can read. A scheduled background task picks up new session logs from hosts such as Claude Code and Codex. The agent decides what is worth keeping and writes it as memory or skill Markdown. memU itself makes no LLM or chat calls: it stores, embeds and retrieves the Markdown the agent prepared. A standing instruction in CLAUDE.md or AGENTS.md tells the agent to retrieve before it answers. All hosts share one backend, so another host, on any model, can retrieve what an earlier one stored.
To try the hosted version, get an API key from memu.so and send your agent one message:
Read https://memu.pro/SKILL.md, follow its instructions to install and configure memU, API Key is
YOUR_MEMU_API_KEY.
The self-hosted setup and the support status for each agent are in the memU GitHub repository.
FAQ
How much does Claude Sonnet 5.5 cost?
$2 per million input tokens and $10 per million output tokens, the same as Claude Sonnet 5. Cache reads cost $0.20 per million tokens. The Batch API is 50% off.
Is Claude Sonnet 5.5 better than Claude Opus 5.5?
On agent and knowledge-work benchmarks the two are close, and Sonnet 5.5 costs half as much per token. Opus 5.5 is more accurate when the answer depends on the model's own knowledge. At maximum effort Sonnet 5.5 can also use more tokens per task, which narrows the price gap.
Can Claude Opus 5.5 read Sonnet 5.5's thinking?
No. Anthropic's documentation says no other model reads Sonnet 5.5's thinking blocks. When a conversation moves to another model, the API drops those blocks before the model sees them. Write anything worth keeping into a message or an external store before the hand-off.
Will my Claude Sonnet 5 code work with Sonnet 5.5 unchanged?
Possibly not. Five changes now return errors, including forced tool use and the "disabled" thinking setting. Check your integration against Anthropic's migration guide before switching.
What is the context window of Claude Sonnet 5.5?
1M tokens, with up to 128K output tokens per request. The Message Batches API allows up to 300K output tokens with a beta header.