Claude Opus 5.5: Pricing, Benchmarks vs GPT-6 Astra, Banked Reset and Breaking Changes
Claude Opus 5.5 costs $4/$20 per million tokens, 40% less than Opus 5, and leads Terminal-Bench 4.0 at 66.4%. Full model card, plan changes, the banked rate-limit reset and 4 breaking API changes.
Claude Opus 5.5 is Anthropic’s new flagship model, released on 22 September 2026 as the first member of the Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5: $4 per million input tokens, $20 per million output tokens, cache reads at $0.20. This is the written companion to my 3-minute explainer video, with every published number in one place, the plan changes (including the banked rate-limit reset), the four breaking API changes, and the sources at the end.
Watch the 3-minute explainer: Claude Opus 5.5 Explained on YouTube. Made with ZArchitect.
Key facts (quick answer)
| Fact | Claude Opus 5.5 |
|---|---|
| Release date | 22 September 2026 |
| Model id | claude-opus-5-5 |
| Price (input / output per 1M tokens) | $4 / $20 (Opus 5: $5 / $25) |
| Cache read / write | $0.20 / $5 (Opus 5: $0.50 / $6.25) |
| Fast mode | $8 / $40, up to 2.5× faster output |
| Context / max output | 1M tokens / 128K tokens |
| Default effort | medium (Opus 5: high); thinking adaptive, always on |
| Terminal-Bench 4.0 | 66.4% (GPT-6 Astra 57.9%, Fable 5.1 55.8%, Opus 5 52.3%) |
| GDPval-AA v2.1 | 1846 Elo (Fable 5.1 1735, GPT-6 Astra 1542) |
| Availability | Claude API, Amazon Bedrock, Google Cloud Vertex AI, Microsoft Foundry |
| Plans | higher 5-hour limits; banked rate-limit reset on Pro, Max, Team, Enterprise |
| Next models | Sonnet 5.5 and Haiku 5.5 in the coming weeks |
TL;DR
- Price: $4 in / $20 out per million tokens (Opus 5: $5 / $25). Cache reads $0.20 (was $0.50). Cache writes $5 (was $6.25). Batch at half price. Net effect on typical workloads: about 40% cheaper, plus 30% faster output.
- Plans: higher five-hour usage limits on Pro, Max, Team and seat-based Enterprise, and a banked rate-limit reset you save and spend when you choose.
- Benchmarks: leads every published agentic-coding, computer-use and knowledge-work benchmark. GPT-6 Astra keeps AutomationBench and Terminal-Bench-Science.
- Breaking changes: thinking always on, no forced tool choice, thinking blocks bound to the model, old computer-use tool retired.
- Safety: best score to date on Anthropic’s automated behavioral audit, 85% fewer containment-boundary attempts than Opus 5, Fable-class safeguards on cyber and biology.
- Next: Sonnet 5.5 and Haiku 5.5 “in the coming weeks”.
How much does Claude Opus 5.5 cost? The model card, Opus 5 → Opus 5.5

| Opus 5 | Opus 5.5 | |
|---|---|---|
| Input / output per 1M tokens | $5 / $25 | $4 / $20 |
| Cache read / write | $0.50 / $6.25 | $0.20 / $5 |
| Batch (half price) | $2.50 / $12.50 | $2 / $10 |
| Fast mode | — | $8 / $40, up to 2.5× faster |
| Context · max output | 1M · 128K | 1M · 128K |
| Effort default | high | medium |
| Thinking | optional | adaptive, always on |
Model id on the Claude Platform: claude-opus-5-5. Available today on the Claude API, Amazon Bedrock, Google Cloud Vertex AI and Microsoft Foundry. Fast mode is a research preview in Claude Code and on the Claude Platform.
The effort default matters more than it looks. Opus 5 ran at high by default; Opus 5.5 runs at medium. If you compare the two at their defaults you are comparing different effort settings, so re-run your own effort sweep before you conclude anything about quality or cost.
What is the banked rate-limit reset? Plans and limits
Two changes for subscribers landed with the model:
- Five-hour usage limits go up on Pro, Max, Team and seat-based Enterprise plans.
- A banked rate-limit reset. When you hit a limit, you no longer have to wait it out: you get a reset you can save and use whenever you choose. This mirrors the reset OpenAI added to its plans, and it is the single most useful change for anyone who works in bursts.
How does Claude Opus 5.5 compare to GPT-6 Astra and Fable 5.1? Benchmarks, with Anthropic’s own caveats

| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 (Main) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 | 57.8% | 51.8% | 46.6% | — | 41.7% |
| GDPval-AA v2.1 (knowledge work, Elo) | 1846 | 1735 | 1708 | 1542 | 1588 |
| AutomationBench (business workflows) | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity’s Last Exam (with tools) | 67.7% | 65.6% | 63.6% | 57.2% | — |
| Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| OSWorld 2.0 (computer use, partial) | 81.8% | 80.7% | 74.0% | — | — |
| Chartography (with tools) | 89.0% | 88.4% | 83.4% | — | — |
Things to keep in mind when you read that table:
- All Claude results use adaptive thinking at max effort unless noted. Terminal-Bench 4.0 is reported at xhigh effort for Opus 5.5 and at high for GPT-6 Astra, as reported by OpenAI.
- Opus 5.5 was evaluated with its production safeguards on. When a safeguard intervened, cyber tasks were handled by Opus 4.8 and biology tasks by Opus 5, which likely lowers its scores.
- AutomationBench was run by Zapier without fallback models, so safeguard interventions counted as failures.
- Anthropic’s own words: at this level, “benchmark margins have become a less reliable guide to real-world differences”, and in their own use the gap to Fable 5.1 is narrower than the scores suggest.
Where the advantage is unambiguous is efficiency. At default effort on FrontierCode, Opus 5.5 beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal-Bench 4.0 it matches Astra at about 40% of the cost. On GDPval at medium effort it beats Astra at max effort for about a fifth of the cost per task.
The tester stories
- A 680,000-line code migration finished in under a day.
- Asked to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times; Opus 5 made smaller gains and changed the app’s behaviour along the way.
- HAProxy translated from C to Rust: Opus 5.5 finished in 9.5 hours against 12 for Fable 5.1, at 51% less cost, both passing nearly all of HAProxy’s regression tests.
- Company earnings reports from a copy of the web where the release was hard to find: 16 of 18 Opus 5.5 reports cleared a bar where any invented figure fails. Neither Fable 5.1 nor Opus 5 cleared it in any attempt.
- A merger analysis in Excel plus an executive deck: Opus 5.5 in 63 minutes versus 93 for Opus 5, at 50% less cost, with fewer errors.
How does Opus 5.5 write? Like a colleague
The most common complaint about Opus 5 was the writing: long, dense, the key point at the end. Opus 5.5 puts the most important information first, uses less jargon and follows the writing rules you give it. Anthropic’s side-by-side of a billing bug explanation shows the difference: Opus 5.5 opens with the dollar impact and the offending commit, then explains the half-open interval that turned into a closed one. Testers summed it up as “it writes the way I do”. That is a safety feature as much as a convenience: work you can follow is work you can check.
The most secure coding agent so far
- A classifier screens every action before it runs.
- An open-source sandbox that security teams can audit.
- Code review that catches vulnerabilities before merge.
- Prompt injection: matches or beats Opus 5 in every setting tested, and ties Fable 5.1 for the lowest success rate on Gray Swan’s benchmark.
- Preserved thinking (the anti-distillation safeguard from Fable 5.1) applies to Opus 5.5 for API accounts created on or after 31 August 2026: you cannot edit Claude’s prior context to extract its reasoning. Check your integrations against the preserved-thinking docs.
- Zero data retention is still available. EU AI Act watermarking applies as with Fable 5.1.
What breaks when you switch to claude-opus-5-5? The four breaking changes

- Thinking can no longer be disabled.
thinking.typeofdisabledorenablednow returns 400. Useeffortto control how much it thinks. - No forced tool choice.
tool_choicevaluesanyandtoolreturn 400. Useautowith strict tool definitions. - Thinking blocks are bound to the model that produced them and to the conversation.
computer_20251124is retired on the Claude API and Google Cloud. Use the current toolset.
And two defaults to re-check: effort is now medium (Opus 5 ran at high), and fast mode is priced separately at $8 / $40.
How safe is Claude Opus 5.5? The safety numbers
- Automated behavioral audit across nearly 2,000 scenarios: the best results of any Claude model to date, and the strongest on most measures of honesty.
- In a new containment evaluation, Opus 5.5 attempted to cross boundaries about 85% less often than Opus 5 or Claude Mythos 5.1; every attempt was low severity and self-reported.
- Outside pre-release evaluation by METR and Frontier Design.
- Because it matches Mythos 5.1 in biology and cyber capability, it ships with Fable-class safeguards; most cyber tasks are re-routed to Opus 4.8, and biology research needs the Life Sciences Verification Program. A Cyber Verification Program expansion is coming.
- Anthropic is candid that Opus 5.5 “often suspects it is being evaluated”, which limits how much any pre-release test can promise.
This is also the first release since Anthropic called for pacing the frontier, so read the safety section of the launch post as policy as much as engineering.
The demo: from one sentence to a playable game
The part everyone asks about. One prompt asked for an animated SVG of a teacher teaching. One more sentence, “make a playable flying airplane game”, produced Dusk Flyer: fuel, birds, balloons, scoring, hold to climb, let go to dive, crash and fly again. The video shows the recording full frame; it is a screen capture of Claude Artifacts, shown for illustration.
My take
The price cut is real and the efficiency claims are the part I trust most, because they are the easiest to verify on your own workload. The benchmark table is a win, but Anthropic’s own caveat about margins is the right way to read it, and GPT-6 Astra still owns two rows. If you are on a subscription, the banked reset changes how you can work more than any benchmark does. If you are on the API, the four breaking changes are a 30-minute migration; the effort default is the one that will silently change your bills and your quality if you ignore it.
Sources
- Anthropic, “Introducing Claude Opus 5.5” (22 September 2026): https://www.anthropic.com/claude-opus-5-5
- Claude Platform docs, “What’s new in Opus 5.5”: https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5
- Opus 5.5 System Card, linked from the launch post
- Benchmarks are vendor-reported; GPT-6 Astra and GPT-5.6 Sol figures as reported by OpenAI, AutomationBench by Zapier.
Independent explainer by Anass Kartit. Not affiliated with Anthropic. Video made with ZArchitect.
Newsletter
Get the next post and game in your inbox
One email when I publish something new: measured write-ups on AI, local models and cloud, plus games like NEON RUN.