Claude Sonnet 5.5: What Anthropic's Faster Mid-Tier Model Means for Your Business
On 28 September Anthropic released Claude Sonnet 5.5, the newest version of its mid-tier AI model, and priced it exactly the same as the model it replaces: $2 per million input tokens and $10 per million output tokens. That sounds like a quiet update. It is not. The company claims the new model generates output more than 30% faster than Sonnet 5 and, on many tasks, finishes the whole job at up to 30% lower cost, simply because it needs fewer tokens and fewer tool calls to get there.
For business owners, this launch matters more than most frontier model announcements. The expensive top-tier models make headlines, but mid-tier models like Sonnet are what actually run inside customer support systems, internal tools, coding assistants and document workflows, because that is where the economics work. When the workhorse tier gets noticeably better without getting more expensive, the maths changes for every product that has AI in it, including the ones you may be planning right now.
So what did Anthropic actually ship, how honest are the numbers, and what should you do about it? Let's take a closer look.

What exactly Anthropic released
Claude Sonnet 5.5 is the middle model in Anthropic's line-up, sitting under the flagship Opus 5.5 and above the small Haiku models. It became available on 28 September through the Claude apps, Claude Code and the Claude Platform API, and also through AWS, Google Cloud and Microsoft Azure, so most companies can reach it through the cloud contract they already have.
Pricing stays where Sonnet 5 was: $2 per million input tokens, $10 per million output tokens, with cached input at $0.20 per million. For comparison, Opus 5.5 costs $4 and $20, exactly double. Anthropic also said a new Haiku 5.5 is coming in the following weeks, which will refresh the budget tier as well.
One practical detail worth noting: the model ships with five adjustable effort levels, from low to max. The default is medium in the consumer apps and high on the API. That means two teams can get quite different speed and cost profiles from the same model, depending on how they configure it. If you test it, test it at the setting you would actually run in production.
The benchmark numbers, in plain language
The headline result is Terminal-Bench 4.0, a test of how well a model handles real work in a command-line environment. Sonnet 5.5 scored 70.6%, against 10.3% for Sonnet 5 and 66.4% for the much more expensive Opus 5.5. A mid-tier model beating its own flagship on an agentic benchmark is not something we saw a year ago.
On GDPval-AA, which measures performance on economically valuable knowledge work such as documents, spreadsheets and presentations, Sonnet 5.5 scored 1844 versus 1846 for Opus 5.5. The gap between the middle tier and the top tier has, on this measure at least, almost disappeared. On OSWorld, a computer-use benchmark, it reached 80.1%, again just behind Opus.
Independent testing adds a useful correction to the marketing. Artificial Analysis gave Sonnet 5.5 an intelligence index of 56, a big jump from Sonnet 5's 38, but measured its cost per task in their harness at $7.60, higher than Sonnet 5's $5.09, because the new model thinks in longer reasoning chains by default. In their runs, Opus 5.5 scored 58 at $5.98 per task. In other words, the savings are real on many workloads but not automatic on all of them. Your own tests on your own tasks are the only numbers that count.
Same price per token, cheaper per task
The most interesting part of this release is the argument behind it. Anthropic is not competing on price per token this time. It is competing on price per finished task, which is the number that actually appears in your monthly bill.
The logic is simple. If a model works faster, uses fewer tokens to reach an answer and makes fewer calls to external tools along the way, the same job costs less even though the price list has not moved. According to Anthropic's own testing, that adds up to as much as 30% per task. The app-building platform Lovable reported that coding tasks needed about a third fewer tool calls and roughly half as many shell executions compared with the previous model.
This is a trend worth watching well beyond Anthropic. As AI moves from chat into agents that carry out multi-step work, efficiency per task is becoming the real battleground. For buyers, it means the price table on a vendor's website tells you less than it used to. Two models with identical token prices can produce very different invoices.

Where Claude Sonnet 5.5 sits against OpenAI and Google
The mid-tier market is now genuinely crowded. Google's Gemini 3.8 Flash undercuts everyone at roughly $0.75 per million input tokens and $3.75 for output, while OpenAI keeps its own middle options close to Sonnet's range. Sonnet 5.5's answer is not to be the cheapest but to be the cheapest model that can reliably handle serious agentic work: debugging, document production, browser and computer use, longer chains of actions.
The launch timing also says something. It came the same week OpenAI cancelled a planned model over safety test failures and Nvidia introduced hardware for containing autonomous agents. The industry's attention is shifting from raw capability to models that behave predictably at a price companies can sustain, and Sonnet 5.5 is aimed precisely at that gap.
For most businesses the practical conclusion is boring but useful: the days when one obvious model choice existed are over. The sensible approach is to keep your integration flexible enough to swap models, then let benchmarks and your own bills decide.
What early adopters are seeing
Anthropic published early customer numbers, and they line up with the efficiency story. Box reported results 2.4 times faster with 12% fewer total tokens. Zendesk saw ticket processing speed up by about 20%. Slack measured roughly 14% fewer output tokens for the same work, and Atlassian said its Rovo agents run up to 30% faster.
Two figures stand out. Base44, which builds full applications from natural language, said the new model needed 3.6 iterations to reach a working build where Opus 5 needed 7.7. And the trading firm Balyasny reported that a research answer that used to consume 497,000 tokens now takes about 121,000. Numbers from a vendor's launch page always deserve some scepticism, but the pattern across ten different companies is consistent: fewer steps, fewer tokens, faster results.

Stronger safety rails, and one catch for developers
Sonnet 5.5 is the first mid-tier Claude model to carry the same cybersecurity safeguards as the flagship, with Anthropic reporting 99.43% recall on blocking harmful cyber requests. Requests such as discovering vulnerabilities in compiled binaries are now refused by design.
The catch is that stricter filters sometimes catch legitimate work. Security researchers and developers doing penetration testing or malware analysis may hit false positives, and API customers have to opt in to fallback behaviour for blocked requests. If your product operates anywhere near security tooling, this deserves a proper test before you switch models in production, not after.
For everyone else, tighter guardrails on the model your customer-facing features run on is a quiet benefit. It reduces the chance of your own product being misused through prompt tricks, which is a risk many businesses still underestimate.
Practical takeaways for businesses building digital products
First, re-run your model comparison if you made it more than a couple of months ago. If you are paying for a flagship model to power support flows, internal assistants or document generation, there is a fair chance a mid-tier model now does the same job at half the token price. That is worth an afternoon of testing.
Second, measure cost per completed task, not cost per token. Set up a small evaluation with twenty or thirty real tasks from your own workload and compare total spend and completion quality across two or three models. The Artificial Analysis results show why: a model can be cheaper per token and dearer per task, or the other way round.
Third, build for portability. Keep model calls behind your own thin service layer so that switching providers is a configuration change, not a rewrite. The mid-tier leaderboard has changed three times this year already, and it will change again.
And finally, if you have been putting off an AI feature because the economics did not work, check the numbers again. Faster, cheaper mid-tier models are exactly what turns a marginal business case into a viable one.
Final notes
Claude Sonnet 5.5 is not a revolution, and that is rather the point. It is the workhorse tier quietly reaching a level that matched the flagship six months ago, at half the price and with better manners. The launches that change budgets tend to look like this one: no dramatic demo, just better economics per task.
Our advice is simple. Do not take the 30% figure on faith, test it on your own workload, and make sure your product can change models without pain. If you want help evaluating what the new generation of mid-tier models could do for your product, or building one around them, the Davydov Consulting team is always happy to talk it through.





Comments