Anthropic's Sonnet 5.5 Matches Opus Until You Read the Token Bill
Anthropic says Claude Sonnet 5.5 is 30% faster and 30% cheaper per task. An independent benchmarking firm pushed it to the limit and got a different answer on cost. Here's what's actually going on.
Is Anthropic's new model actually cheaper? Depends who you ask, and how hard you push it.
JUST IN: Claude Sonnet 5.5 landed September 28 at the exact same list price as Sonnet 5. Anthropic says it runs over 30% faster and costs up to 30% less per task. Sounds like free money. It isn't.
Same Price, Different Bill
Anthropic's pitch is clean. Same sticker price, more speed, lower cost per completed job. On paper that's a straight win for anyone burning tokens at scale.
Then an independent benchmarking firm ran the model to its limit, and its conclusion on cost didn't match the marketing. Push Sonnet 5.5 hard and it burns through more tokens getting to the same answer.
That's the catch nobody's leading with. Token count is the real invoice. Price per token is just the sticker.
So the 30% number holds up fine on light prompts. On heavy multi-step reasoning, it gets a lot messier.
Why Token Discipline Is the Whole Game
Here's the thing about model pricing. A cheaper rate per token means nothing if you need 40% more of them.
Sonnet 5.5 now gets close to Opus-level output. That's the story under the story. You're getting flagship-adjacent reasoning at mid-tier pricing, which is a genuinely big deal for anyone building agents that run all day.
But it comes bundled with a catch. The model thinks longer. It writes more. It spends more getting there.
And just like that, the discount gets eaten alive.
This matters because the whole industry is shifting from chat to agents. Agents don't send one prompt. They loop. They plan, retry, and re-read context dozens of times per job. Token efficiency isn't a footnote in that world. It's the P&L.
What the Benchmarkers Say
Traders are watching this one closely, and not the crypto kind. The people tracking it are devs and infra teams running production agent workloads.
According to the benchmarking firm, Sonnet 5.5's token efficiency breaks down on harder tasks. Simple prompts? Excellent. Complex multi-step reasoning? Expensive.
My take: Anthropic's "up to 30% less per task" is carrying a lot of weight in that one word.Up to.Every lab does this. It's the escape hatch that lets marketing survive contact with reality.
Does that make Sonnet 5.5 a bad model? No. It makes it a model you've to measure, not trust. The speed gain is real. The Opus-adjacent quality is real. The cost claim is conditional, and conditions matter more than headlines when you're paying the bill.
What to Watch Next
Watch the next four to six weeks. Independent evals on real agent workloads, not sanitized benchmarks, will settle this argument. If token counts stay elevated on multi-step tasks, the cost advantage quietly disappears for the heaviest users.
Also watch OpenAI and Google. If either ships comparable capability with tighter token discipline, Anthropic's speed story gets a lot less interesting in a hurry.
And watch Opus pricing. If Anthropic cuts it, that tells you everything about where Sonnet 5.5 really sits in the lineup.
My advice? Run your own workload before you migrate. Your prompt mix isn't theirs, and their benchmark isn't your bill.
The market's verdict on Sonnet 5.5 won't come from a press release. It'll come from the invoice. This changes things.