Claude Opus 5.5 for Small Business: The Number That Matters Is the Cache Price

Claude Opus 5.5 costs 20% less per token than Opus 5, but its cached input now matches Sonnet 5. What that changes for small business AI automations.
Chart showing Claude Opus 5.5 cached input price matching Sonnet 5
On this page

Anthropic released Claude Opus 5.5 on September 22, 2026, two months after Opus 5. The headline is price: $4 per million input tokens and $20 per million output, down from $5 and $25. Anthropic says it costs about 40 percent less to run than Opus 5 on typical workloads and generates output more than 30 percent faster.

The line we care about more is further down Anthropic’s pricing documentation. Cached input on Opus 5.5 now costs $0.20 per million tokens. That is the same as Sonnet 5.

For small businesses running AI inside their marketing and sales systems, that one number changes more than the headline cut.

What Anthropic actually shipped

From Anthropic’s announcement and its pricing page:

Opus 5Opus 5.5Sonnet 5Haiku 4.5
Input, per million tokens$5$4$2$1
Output, per million tokens$25$20$10$5
Cached input (cache hit)$0.50$0.20$0.20$0.10
Batch input / output$2.50 / $12.50$2 / $10$1 / $5$0.50 / $2.50

Beyond price, Anthropic reports Opus 5.5 ahead of Opus 5 across its benchmark table, including 40.0 percent on AutomationBench against Opus 5’s 26.9 percent, and says it performs at the level of its larger Fable 5.1 model on most work. The model is available on the Pro, Max, Team, and Enterprise plans with higher five-hour usage limits, and through the API on AWS, Google Cloud, and Microsoft Azure. TechCrunch and 9to5Mac both note that Sonnet 5.5 and Haiku 5.5 are due in the coming weeks.

Why the cache price matters more than the headline

Most small business AI automations send the same long block of text on every call: brand voice rules, service descriptions, pricing, FAQs, qualification criteria. The customer’s actual message is a small part of each request. That long, repeated block is exactly what prompt caching discounts.

So we ran an illustration. Take a lead-reply assistant with a 20,000-token instruction block, 1,000 tokens of fresh input per inquiry, and a 500-token reply, handling 1,000 inquiries a month. Using the list prices above, here is what it costs when the cache stays warm, meaning each call finds the instruction block already cached:

ModelCached inputFresh inputOutputMonthly total, cache warm
Opus 5$10.00$5.00$12.50$27.50
Opus 5.5$4.00$4.00$10.00$18.00
Sonnet 5$4.00$2.00$5.00$11.00
Haiku 4.5$2.00$1.00$2.50$5.50

Two things jump out. Opus 5.5 is about a third cheaper than Opus 5 on this shape of workload, before counting any token savings from shorter answers. And the gap to Sonnet 5 is now mostly the output price, because the part of the bill that used to punish Opus, the big repeated context, costs the same on both.

The catch nobody mentions: small businesses often have a cold cache

Illustrative monthly cost of a lead reply assistant on four Claude models, warm versus cold cache
A cache that is always written and never read costs more than not caching at all. The warm-cache discount only exists if your traffic keeps the cache warm.

Anthropic’s prompt caching documentation says the default cache lives for 5 minutes, refreshed each time it is used. A thousand inquiries a month is roughly one every 40 minutes if they arrive evenly. For a dental clinic in Cebu, in the central Philippines, or a small US accounting firm, inquiries often arrive far enough apart that the 5-minute cache has expired, and the call pays to write the cache again instead of reading it.

Here is the same workload with every call finding a cold cache and paying the 5-minute cache write price:

ModelCache writeFresh inputOutputMonthly total, cache cold
Opus 5$125.00$5.00$12.50$142.50
Opus 5.5$100.00$4.00$10.00$114.00
Sonnet 5$50.00$2.00$5.00$57.00
Haiku 4.5$25.00$1.00$2.50$28.50

Look at Opus 5.5 there. With no caching at all, that 20,000-token block would cost $80 a month at the normal $4 input price. A cache that is always written and never read costs 25 percent more than not caching. The warm-cache discount only exists if your traffic keeps the cache warm.

What to do about it: busy workflows (a chatbot on a high-traffic site, a batch of CRM records processed together) get the full benefit. Low, spread-out traffic should either skip caching for that call, try the 1-hour cache (written at 2x the input price, so it pays off after two reads, per Anthropic’s pricing page), or be batched: collect inquiries for a few minutes and process them together.

These are illustrations from list prices, not a client bill. Your real traffic pattern decides which table you live in, so pull a week of usage and look at cache reads versus cache writes before trusting any calculator, ours included.

What this changes from our Opus 5 advice

Diagram splitting an AI workflow between Opus for judgment and Sonnet for drafting
Judgment steps get cheaper to keep on Opus. Drafting stays on the cheaper models, where output price dominates.

When Opus 5 came out, our read in Claude Opus 5 for small business was that most small business workloads should move down the ladder to Sonnet or Haiku, and use Opus only for the judgment calls. We still think that is the right shape. What changes is where the line sits.

  • Judgment steps get cheaper to keep on Opus. Lead qualification, routing an inquiry to the right service, deciding whether a complaint needs a human. These are short-output, long-context tasks, which is the exact profile the cache price rewards.
  • Drafting stays on the cheaper models. Long replies, blog outlines, and summaries are output-heavy. Output on Opus 5.5 is still twice Sonnet 5’s price.
  • Batch jobs are worth a second look. Anything that does not need an instant answer (overnight CRM enrichment, tagging last week’s inquiries, first-pass review of ad copy) gets the 50 percent batch discount on top.

Anthropic also published customer claims about token efficiency. Box, for example, says Opus 5.5 used a third of the tokens Opus 5 did on its task. Vendor-selected quotes are marketing until you reproduce them. If even part of that holds for your prompts, the output side of the table above shrinks too.

What it does not change

A 40 percent score on AutomationBench is a big jump from 26.9 percent. It also means most of the tasks in that benchmark still were not completed. We will keep saying it with every release: better models make unsupervised automation less wrong, not safe. Keep the human checkpoint on anything that sends money, makes promises to a customer, or changes records in bulk. We covered the readiness side of this in why most small businesses are not ready for AI sales agents, and nothing in this release moves those conditions.

It also does not fix a bad prompt structure. Caching works on the prompt prefix: Anthropic’s docs say a change at any level invalidates the cache from that point on. So the discount only covers the part of your prompt that stays identical from call to call, and only if that part comes first. If your automation inserts the customer’s name or today’s date at the top of the instruction block, the cache never gets a hit and you pay full input price every time.

And if you use Claude through a Pro or Team subscription rather than the API, the per-token math above is not your bill. For you, the practical change is faster answers and higher usage limits.

The test we are running next

Here is the plan, written out so you can run the same test on your own pipeline:

  • Take 20 real past inquiries from one client pipeline, with the correct routing decision already known.
  • Run them through Sonnet 5 and through Opus 5.5 at a low effort setting, with the same cached instruction block.
  • Compare three things: routing accuracy, total tokens per inquiry, and cost per correct decision.

We will not know the answer until we run it, and we will not know how Opus 5.5 compares to Sonnet 5.5 until Sonnet 5.5 ships. If the cheaper model is within a point or two on accuracy, it stays. If Opus 5.5 catches the edge cases Sonnet misses, the cache price makes it affordable to switch just that step.

If you run your own version of this test, send us your numbers. Comparing notes across different pipelines is how this stops being guesswork.

Frequently asked questions

How much does Claude Opus 5.5 cost?

Anthropic lists Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens, down from $5 and $25 for Opus 5. Cached input costs $0.20 per million tokens, and the Batch API halves input and output prices to $2 and $10.

Is Claude Opus 5.5 better than Opus 5?

On Anthropic’s published benchmarks, yes, across every test in its comparison table, including AutomationBench at 40.0 percent versus 26.9 percent for Opus 5. Anthropic also says it runs about 40 percent cheaper on typical workloads and generates output more than 30 percent faster. Test it on your own tasks before switching production workflows.

Should a small business use Claude Opus 5.5 or Sonnet?

Use Opus 5.5 for short, judgment-heavy steps with a long repeated context, such as lead qualification or routing, where its cached input now costs the same as Sonnet 5, as long as traffic keeps the cache warm. Keep long drafting and summarizing on Sonnet or Haiku, since Opus 5.5 output still costs twice Sonnet 5’s rate.

When will Claude Sonnet 5.5 and Haiku 5.5 be released?

Anthropic said at the Opus 5.5 launch that Sonnet 5.5 and Haiku 5.5 will arrive in the coming weeks, without a specific date. If you are deciding how to split work between models, it is reasonable to wait for them before rebuilding the cheaper tiers of your automations.

Over to you

We write up model tests like this one as we run them, including the ones where the new model loses. Join the list if you want the results when they are in, or talk to us about your automations if you want help deciding which model belongs where.

About the author

Picture of Benison David Sanchez

Benison David Sanchez

Co-founder and CEO of BDGS Digital, leading AI digital marketing. Benison leads digital marketing, client acquisition, and brand strategy, and makes sure every campaign connects back to a larger growth system.

Every article follows our editorial policy and corrections policy. Meet the team behind BDGS Digital.

Next step

Need Help Implementing This for Your Business?

Tell us about your setup and we will tell you honestly whether it is worth doing now, and how we would approach it.

On this page
Scroll to Top