Anass Kartit
← Writing / / 9 min read / Updated

Claude Opus 5.5: Pricing, Benchmarks vs GPT-6 Astra, Banked Reset and Breaking Changes

Claude Opus 5.5 costs $4/$20 per million tokens, 40% less than Opus 5, and leads Terminal-Bench 4.0 at 66.4%. Full model card, plan changes, the banked rate-limit reset and 4 breaking API changes.

claude opus 5.5claude opus 5.5 pricingclaude opus 5.5 benchmarksopus 5.5 vs gpt-6 astraanthropicclaude api breaking changesbanked rate limit resetclaude codellm pricing 2026
Claude Opus 5.5 Explained: Pricing, Benchmarks & Breaking Changes
Watch: Claude Opus 5.5 Explained: Pricing, Benchmarks & Breaking Changes

Claude Opus 5.5 is Anthropic’s new flagship model, released on 22 September 2026 as the first member of the Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5: $4 per million input tokens, $20 per million output tokens, cache reads at $0.20. This is the written companion to my 3-minute explainer video, with every published number in one place, the plan changes (including the banked rate-limit reset), the four breaking API changes, and the sources at the end.

Watch the 3-minute explainer: Claude Opus 5.5 Explained on YouTube. Made with ZArchitect.

Key facts (quick answer)

FactClaude Opus 5.5
Release date22 September 2026
Model idclaude-opus-5-5
Price (input / output per 1M tokens)$4 / $20 (Opus 5: $5 / $25)
Cache read / write$0.20 / $5 (Opus 5: $0.50 / $6.25)
Fast mode$8 / $40, up to 2.5× faster output
Context / max output1M tokens / 128K tokens
Default effortmedium (Opus 5: high); thinking adaptive, always on
Terminal-Bench 4.066.4% (GPT-6 Astra 57.9%, Fable 5.1 55.8%, Opus 5 52.3%)
GDPval-AA v2.11846 Elo (Fable 5.1 1735, GPT-6 Astra 1542)
AvailabilityClaude API, Amazon Bedrock, Google Cloud Vertex AI, Microsoft Foundry
Planshigher 5-hour limits; banked rate-limit reset on Pro, Max, Team, Enterprise
Next modelsSonnet 5.5 and Haiku 5.5 in the coming weeks

TL;DR

  • Price: $4 in / $20 out per million tokens (Opus 5: $5 / $25). Cache reads $0.20 (was $0.50). Cache writes $5 (was $6.25). Batch at half price. Net effect on typical workloads: about 40% cheaper, plus 30% faster output.
  • Plans: higher five-hour usage limits on Pro, Max, Team and seat-based Enterprise, and a banked rate-limit reset you save and spend when you choose.
  • Benchmarks: leads every published agentic-coding, computer-use and knowledge-work benchmark. GPT-6 Astra keeps AutomationBench and Terminal-Bench-Science.
  • Breaking changes: thinking always on, no forced tool choice, thinking blocks bound to the model, old computer-use tool retired.
  • Safety: best score to date on Anthropic’s automated behavioral audit, 85% fewer containment-boundary attempts than Opus 5, Fable-class safeguards on cyber and biology.
  • Next: Sonnet 5.5 and Haiku 5.5 “in the coming weeks”.

How much does Claude Opus 5.5 cost? The model card, Opus 5 → Opus 5.5

Opus 5 versus Opus 5.5 pricing and settings

Opus 5Opus 5.5
Input / output per 1M tokens$5 / $25$4 / $20
Cache read / write$0.50 / $6.25$0.20 / $5
Batch (half price)$2.50 / $12.50$2 / $10
Fast mode$8 / $40, up to 2.5× faster
Context · max output1M · 128K1M · 128K
Effort defaulthighmedium
Thinkingoptionaladaptive, always on

Model id on the Claude Platform: claude-opus-5-5. Available today on the Claude API, Amazon Bedrock, Google Cloud Vertex AI and Microsoft Foundry. Fast mode is a research preview in Claude Code and on the Claude Platform.

The effort default matters more than it looks. Opus 5 ran at high by default; Opus 5.5 runs at medium. If you compare the two at their defaults you are comparing different effort settings, so re-run your own effort sweep before you conclude anything about quality or cost.

What is the banked rate-limit reset? Plans and limits

Two changes for subscribers landed with the model:

  1. Five-hour usage limits go up on Pro, Max, Team and seat-based Enterprise plans.
  2. A banked rate-limit reset. When you hit a limit, you no longer have to wait it out: you get a reset you can save and use whenever you choose. This mirrors the reset OpenAI added to its plans, and it is the single most useful change for anyone who works in bursts.

How does Claude Opus 5.5 compare to GPT-6 Astra and Fable 5.1? Benchmarks, with Anthropic’s own caveats

Agentic coding benchmarks: Terminal-Bench 4.0, FrontierCode, CursorBench

BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 Sol
Terminal-Bench 4.0 (agentic coding)66.4%55.8%52.3%57.9%37.3%
FrontierCode v1.1 (Main)54.4%50.3%48.0%53.3%47.5%
CursorBench 4.057.8%51.8%46.6%41.7%
GDPval-AA v2.1 (knowledge work, Elo)18461735170815421588
AutomationBench (business workflows)40.0%31.4%26.9%41.4%28.8%
Humanity’s Last Exam (with tools)67.7%65.6%63.6%57.2%
Terminal-Bench-Science 0.158.7%52.6%29.0%64.6%22.4%
OSWorld 2.0 (computer use, partial)81.8%80.7%74.0%
Chartography (with tools)89.0%88.4%83.4%

Things to keep in mind when you read that table:

  • All Claude results use adaptive thinking at max effort unless noted. Terminal-Bench 4.0 is reported at xhigh effort for Opus 5.5 and at high for GPT-6 Astra, as reported by OpenAI.
  • Opus 5.5 was evaluated with its production safeguards on. When a safeguard intervened, cyber tasks were handled by Opus 4.8 and biology tasks by Opus 5, which likely lowers its scores.
  • AutomationBench was run by Zapier without fallback models, so safeguard interventions counted as failures.
  • Anthropic’s own words: at this level, “benchmark margins have become a less reliable guide to real-world differences”, and in their own use the gap to Fable 5.1 is narrower than the scores suggest.

Where the advantage is unambiguous is efficiency. At default effort on FrontierCode, Opus 5.5 beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal-Bench 4.0 it matches Astra at about 40% of the cost. On GDPval at medium effort it beats Astra at max effort for about a fifth of the cost per task.

The tester stories

  • A 680,000-line code migration finished in under a day.
  • Asked to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times; Opus 5 made smaller gains and changed the app’s behaviour along the way.
  • HAProxy translated from C to Rust: Opus 5.5 finished in 9.5 hours against 12 for Fable 5.1, at 51% less cost, both passing nearly all of HAProxy’s regression tests.
  • Company earnings reports from a copy of the web where the release was hard to find: 16 of 18 Opus 5.5 reports cleared a bar where any invented figure fails. Neither Fable 5.1 nor Opus 5 cleared it in any attempt.
  • A merger analysis in Excel plus an executive deck: Opus 5.5 in 63 minutes versus 93 for Opus 5, at 50% less cost, with fewer errors.

How does Opus 5.5 write? Like a colleague

The most common complaint about Opus 5 was the writing: long, dense, the key point at the end. Opus 5.5 puts the most important information first, uses less jargon and follows the writing rules you give it. Anthropic’s side-by-side of a billing bug explanation shows the difference: Opus 5.5 opens with the dollar impact and the offending commit, then explains the half-open interval that turned into a closed one. Testers summed it up as “it writes the way I do”. That is a safety feature as much as a convenience: work you can follow is work you can check.

The most secure coding agent so far

  • A classifier screens every action before it runs.
  • An open-source sandbox that security teams can audit.
  • Code review that catches vulnerabilities before merge.
  • Prompt injection: matches or beats Opus 5 in every setting tested, and ties Fable 5.1 for the lowest success rate on Gray Swan’s benchmark.
  • Preserved thinking (the anti-distillation safeguard from Fable 5.1) applies to Opus 5.5 for API accounts created on or after 31 August 2026: you cannot edit Claude’s prior context to extract its reasoning. Check your integrations against the preserved-thinking docs.
  • Zero data retention is still available. EU AI Act watermarking applies as with Fable 5.1.

What breaks when you switch to claude-opus-5-5? The four breaking changes

Breaking changes, safety and what is next

  1. Thinking can no longer be disabled. thinking.type of disabled or enabled now returns 400. Use effort to control how much it thinks.
  2. No forced tool choice. tool_choice values any and tool return 400. Use auto with strict tool definitions.
  3. Thinking blocks are bound to the model that produced them and to the conversation.
  4. computer_20251124 is retired on the Claude API and Google Cloud. Use the current toolset.

And two defaults to re-check: effort is now medium (Opus 5 ran at high), and fast mode is priced separately at $8 / $40.

How safe is Claude Opus 5.5? The safety numbers

  • Automated behavioral audit across nearly 2,000 scenarios: the best results of any Claude model to date, and the strongest on most measures of honesty.
  • In a new containment evaluation, Opus 5.5 attempted to cross boundaries about 85% less often than Opus 5 or Claude Mythos 5.1; every attempt was low severity and self-reported.
  • Outside pre-release evaluation by METR and Frontier Design.
  • Because it matches Mythos 5.1 in biology and cyber capability, it ships with Fable-class safeguards; most cyber tasks are re-routed to Opus 4.8, and biology research needs the Life Sciences Verification Program. A Cyber Verification Program expansion is coming.
  • Anthropic is candid that Opus 5.5 “often suspects it is being evaluated”, which limits how much any pre-release test can promise.

This is also the first release since Anthropic called for pacing the frontier, so read the safety section of the launch post as policy as much as engineering.

The demo: from one sentence to a playable game

The part everyone asks about. One prompt asked for an animated SVG of a teacher teaching. One more sentence, “make a playable flying airplane game”, produced Dusk Flyer: fuel, birds, balloons, scoring, hold to climb, let go to dive, crash and fly again. The video shows the recording full frame; it is a screen capture of Claude Artifacts, shown for illustration.

My take

The price cut is real and the efficiency claims are the part I trust most, because they are the easiest to verify on your own workload. The benchmark table is a win, but Anthropic’s own caveat about margins is the right way to read it, and GPT-6 Astra still owns two rows. If you are on a subscription, the banked reset changes how you can work more than any benchmark does. If you are on the API, the four breaking changes are a 30-minute migration; the effort default is the one that will silently change your bills and your quality if you ignore it.

Sources

Independent explainer by Anass Kartit. Not affiliated with Anthropic. Video made with ZArchitect.

Newsletter

Get the next post and game in your inbox

One email when I publish something new: measured write-ups on AI, local models and cloud, plus games like NEON RUN.

Your email is never shown or shared. What is stored and how to leave