Claude Opus 5 for Small Business: What Actually Changed


Anthropic shipped Claude Opus 5 on 24 July 2026 at $5 per million input tokens and $25 per million output tokens, the exact price Opus 4.8 cost before it. The company's framing is that Opus 5 comes close to the frontier intelligence of Claude Fable 5 at half the price. Almost every write-up that followed was about coding benchmarks, which is useful if you run an engineering team and close to useless if you run a dental clinic in Quezon City with two people answering inquiries between patients.
The release still matters to that clinic. Just not for the reason the headlines gave.
Price held flat. Fable 5, the model above Opus 5, still costs $10 per million input and $50 per million output. Sonnet 5 sits at $2 and $10, Haiku 4.5 at $1 and $5, per the Claude pricing docs. So the ladder now spans 50x from bottom to top on output, and every rung got better without getting more expensive.
The benchmark Anthropic led with is Frontier-Bench, a software engineering evaluation. The one a business owner should read is Zapier's AutomationBench, which measures whether a model can finish a business task start to finish. Anthropic reports Opus 5 passing roughly 1.5 times as many tasks as the next-best model at the same cost per task, and passing more tasks than any other model even at its lowest effort setting.
That last clause is the entire release, compressed into one line.
Zapier's CEO Wade Foster described Opus 5 taking a raw account-health workbook and running a churn-prevention sequence end to end: flagging accounts at risk, alerting the right owner, summarising for retention ops. He says prior models did not pass and Opus 5 hit 100 percent. That is a vendor quote inside a launch post, so treat it as a claim rather than a measurement. But look at the shape of the task. Flag, route, notify, summarise. That is most of what a small business automation does all day.
Confirmed: better scores, unchanged price, an effort dial with five settings (low, medium, high, xhigh, max), thinking on by default, and a 1M token context window. See What is new in Claude Opus 5 for the full list.
Implied by the coverage: everybody should move up to Opus 5.
Before you accept that, read the footnote under Anthropic's own Frontier-Bench chart. It says Opus 4.8 stood in as the fallback whenever a safety classifier refused a request from Opus 5 or Fable 5. Some portion of both scores is therefore work an older model did, and Anthropic does not say how large that portion is. Publishing the footnote at all is more than most vendors bother with. It is still a reason to run your own test instead of buying a chart.

Your existing prompts are probably costing you money now. Anthropic's Opus 5 prompting guide tells you to delete verification instructions. If your prompt says "include a final verification step" or "double-check your answer before responding," remove it. Opus 5 verifies its own work, and those instructions cause over-verification with no gain in quality. The same applies to subagent delegation, which the model reaches for more readily than earlier ones and which multiplies cost on small tasks. Both of those lines sit in half the prompt templates people copied off LinkedIn in 2025.
The effort dial is the real product. Low and medium effort produce strong quality at a fraction of the tokens and latency of the higher settings, and Anthropic's guidance is blunt about it: if you carried effort defaults over from a prior model, re-run an effort sweep on your own evaluations. Nobody covering this release led with that, and it is the single line most likely to change your invoice.
Work you priced out last year deserves a second look. Anthropic's own worked example puts roughly 10,000 support conversations at about $37 on Haiku 4.5, averaging around 3,700 tokens each. Prompt caching drops repeated input to 10 percent of the standard rate. The Batch API takes 50 percent off both directions for anything that does not need an answer this second. Half the automation builds we quoted in 2024 and lost on budget would cost a fraction of that today. Overnight review of every inbound inquiry from the past month is now a rounding error, not a project.
Here is the part that keeps getting lost. A fitness coach we worked with was running intake by hand. Inquiries arrived, somebody saw them eventually, and eventually was frequently after the person had already booked with somebody else. What we built in GoHighLevel was not clever, and it is the same shape as most lead generation systems we put in: a proper intake form, an internal notification the second it fires, and an automated reply to the person who filled it in. Their client acquisition went up. We cannot publish the numbers, most of our agreements do not allow it.
Not one step in that build needs a frontier model. It needs a form that fires reliably.
Your bill can rise on identical work. Three reasons stack here. Thinking is on by default on Opus 5 where it was not on Opus 4.8. Claude 4.7 and later models use a newer tokenizer that produces roughly 30 percent more tokens for the same text. And Opus 5's default responses run longer than prior Opus models, which Anthropic documents and suggests you fix with an explicit conciseness instruction. Same headline rate, same volume of work, bigger invoice.
The model was almost certainly not your bottleneck. In five years and 100-plus clients, the thing blocking growth has hardly ever been the intelligence of the text generator. It has been a form that does not notify anyone, a CRM with four spellings of the same company name, an ads account pointing at a page that does not load on a phone, a follow-up sequence nobody turned back on after the holidays. A better model applied to a broken pipe gives you a more articulate leak.
The competition moved harder on price than Anthropic did. On 30 July 2026, OpenAI cut GPT-5.6 Luna by 80 percent to $0.20 and $1.20 per million tokens, and Terra by 20 percent to $2 and $12. If your constraint is cost per task on high-volume, well-defined work, that is the more consequential announcement of the month, and it got a fraction of the attention.

Skip the benchmarks. Build a twenty-record test set out of your own history: twenty real inquiries, twenty real support threads, twenty real quote requests. Write down what a pass looks like before you run anything, in one sentence, in plain language. "Correctly identifies the service requested and does not invent a price."
Then run the same twenty through four configurations: Haiku 4.5, Sonnet 5, Opus 5 at low effort, Opus 5 at default. Score pass or fail yourself. No partial credit. Log what the automation costs to run somewhere you will actually look at again, because the pass rate only means something next to what the automation costs to run.
Our expectation, and we would like to be argued out of it, is that most small business workloads pass on Sonnet 5 or below, and that the cases needing Opus 5 are the ones where the model has to decide something rather than draft something. Routing a lead is a decision. Writing the reply is drafting. Those two jobs have been sitting in the same automation, on the same model, for most of the builds we have inherited.
Whether the effort dial holds up under real load. Anthropic's numbers come from bounded tasks with clean pass conditions, and business work is neither. Whether low effort stays reliable when a workflow runs 400 times a week rather than five is a question the benchmarks cannot answer, and neither can we yet.
We also do not know how the automatic fallback behaves in practice. When a safety classifier flags a request in Claude.ai, Claude Code, or Claude Cowork, it now routes to Opus 4.8 by default rather than failing. Quiet substitution is better than a hard error for most users. It is worse for anyone trying to keep a stable output format across thousands of runs.
Probably not for most automations. Opus 5 is the strongest option for tasks where the model has to make a judgment call, such as routing, triage, or reviewing a document against rules. For drafting replies, tagging leads, summarising calls, or classifying inquiries, Sonnet 5 at $2 and $10 or Haiku 4.5 at $1 and $5 will usually pass your test at a fraction of the cost.
$5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8, as of its release on 24 July 2026. Fast mode runs at twice that. Prompt caching cuts repeated input to 10 percent of the standard rate, and the Batch API takes 50 percent off both input and output for work that can wait.
Effort controls how much the model thinks before answering, across five levels from low to max. Anthropic reports that low and medium produce strong quality at a fraction of the tokens and latency of higher settings. For a small business, it is the most direct control you have over your monthly bill, and defaults carried over from an older model are worth re-testing.
It can. Thinking is on by default on Opus 5, Claude 4.7 and later models use a tokenizer that produces roughly 30 percent more tokens for the same text, and Opus 5 writes longer responses by default. Same rate card, larger invoice. Add an explicit conciseness instruction and drop the effort level where quality holds.
It depends entirely on the task and the volume. After OpenAI's 30 July 2026 price cut, GPT-5.6 Luna at $0.20 and $1.20 per million tokens is dramatically cheaper than anything in the Claude lineup for high-volume, well-defined work. Test both on twenty of your own records before committing, because the gap between vendor benchmarks and your actual pass rate is usually large.
If you have already run an effort sweep on real business tasks rather than coding benchmarks, we want to see the numbers. Our reading is that the ladder matters more than the top rung, and that most people are about to pay Opus prices for Haiku work. We could be wrong about where the line sits.
We write up builds like this one as we do them, including the parts that fail. The same discipline applies to search: we wrote up what Google actually requires to rank in AI Overviews by reading the documentation, not the vendor pitches.
If your CRM, follow-up, and ads are meant to run as one system and it keeps breaking at the handoff, that is the problem worth solving before you change models. Book a strategy call and we will map what you have now and where it is leaking.

Written by
Benison David Sanchez is the CEO and AI Digital Marketing Strategist of BDGS Digital, a technology, CRM, automation, and growth company based in the Philippines and serving clients globally. He leads digital marketing and client acquisition for the company, and over five years has worked with more than 100 businesses across ecommerce, service, coaching, accounting, and technology, building connected growth systems where every campaign ties back to a larger system rather than running in isolation. He spends most of his week inside GoHighLevel workflows, which is where this post came from.
View all articles by BenisonWant help building this?
BDGS Digital builds connected growth systems across CRM, automation, marketing, and technology. Explore our services or book a strategy call.
Book a Strategy CallLET'S TALK
Book a strategy call and we'll map how your CRM, automation, marketing, and technology can work together as one connected system.