Anthropic’s New Model – Claude Sonnet 5
Anthropic released a new model called Claude Sonnet 5. This is the biggest change in the Sonnet series for years. The company made big promises about this model. It was supposed to have fewer errors. It was meant to plan better than older models. The truth is very different from these promises.
We tested this model on many real tasks. The results are very disappointing. The model does not work as it should. In some tests it is worse than older versions. This is a big letdown for all users who waited for it.
Let us compare it with earlier models. For example Claude Opus 4.8 works much better. Many companies that use AI in business waited for this model. But it seems better to stay with older versions. The team at AI w Biznesie tested this model carefully. We do not see a reason to switch to it.
This model was meant to be a smart choice for business users. But the reality does not match the marketing. Users get worse results for similar costs. This makes the model hard to recommend for any serious work.
Benchmarks – Good Numbers, Bad Reality
On paper Sonnet 5 looks decent. The test results are quite good. For example in agentic coding tests the model scores 63.2 percent. That is better than Sonnet 4.6 which had 58.1 percent. But Opus 4.8 has 69.2 percent so the gap is still large.
In other tests the model does not look bad. In Terminal Bench 2.1 it scores 80.4 percent. That is very close to the Opus result. In tests for using a computer the model gets 81.2 percent. But numbers do not show the full truth. In daily use the model works much worse than expected.
Test scores can trick us. They measure specific skills in a lab setting. Real work is different and harder. The model fails in ways that tests cannot catch. This is a common problem with new AI models today.
Real World Rankings
More important are real life tests. For example in Cursor Bench which tests coding models Sonnet 5 is only 13th place. That is a very weak result. Other models like Opus are much higher in this ranking. This shows the model cannot help coders well.
The model was meant to be faster and cheaper than Opus. But in practice this is not true. Users report that the model runs slow. The results are also weaker than promised. This destroys the whole reason to use this model for work.
The team at AI w Biznesie tests many AI models for clients. Our experts saw that Sonnet 5 does not keep its promises. Compared to other models on the market it is weak. This is a problem for companies that need good tools for their work.
Cursor Bench is a trusted test for coding skills. Being 13th place is very bad for a new model. Users who code every day need reliable tools. Sonnet 5 fails to give them what they need for their projects.
Price and Tokens – A Hidden Trap
Anthropic lowered the price for using this model. The starting price is 2 dollars per million input tokens. Tokens are small pieces of text that the model reads. The price for output tokens is 10 dollars per million. This discount lasts until August 2026.
After that time the price will go up. It will be 3 dollars for input tokens and 15 dollars for output. But that is not all. In the fine print there is a key fact. Sonnet 5 uses a new kind of tokens. These are the same tokens that Opus 4.7 uses.
This change is very important for users. It changes how much you pay for the same work. Many people will not see this hidden detail. They will think they are saving money but they are not. The cost savings are not real for most tasks.
More Tokens, Higher Costs
The new tokens are different from before. The same text can have 1.0 to 1.3 times more tokens than earlier. This depends on what the text is about. So the price may look lower but costs go up. Some questions cost more than they seem to cost.
Anthropic did this on purpose. They wanted prices to look good to buyers. But in practice there are no savings. Users pay the same amount as for Opus. They only get worse results for their money. This is not a good deal for anyone who needs AI tools.
Let us compare closely. Sonnet 5 is only 72 cents cheaper than Opus 4.8 Max. That is a very small difference in price. For that small saving it is better to pay a bit more. You get a much better model for a little extra money. At AI w Biznesie we tell clients to check costs very carefully.
The token change is clever marketing but bad for users. It hides the real cost of using the model. Business owners need to know the true price before they choose. This hidden trap can hurt budgets and slow down work projects.
Real Tests – The Model Fails in Practice
We ran several tests to check the model in real use. We wanted to see how it handles true tasks. The results are mixed but mostly weak. The model wastes many tokens and gives poor results back to users.
MacOS Clone Test
The model made a working clone of the MacOS system. This was done in the World of AI Benchmark tool. It took about 40 minutes to finish the task. The model used a very large number of tokens in max mode. This was not efficient work at all.
Some features do work on the clone. There is a top bar and light or dark mode. There are also wallpapers and apps like Safari. You can use Messenger, Calendar and Notes too. But the Maps app is much worse than in the Opus model. The model is not efficient at all for this kind of task.
The team at AI w Biznesie tested the clone carefully. We saw that the model wastes resources. It could do the same job with fewer tokens. But it chose to use more and give less back. This is a bad sign for anyone who needs to save money on AI tools.
Minecraft Game Test
The model tried to make a game like Minecraft. It used new textures which is interesting to see. Water works in the game and there are different characters. You can place blocks in the world as well. But there is no animation for breaking blocks at all.
There is no inventory system in the game. The world does not go on forever as it should. The score is only 6.5 out of 10 points for quality. The game has glitches and does not meet expectations. The model did not handle this task very well for a new release.
AI w Biznesie tests models for many different uses. This one fails at game creation tasks. It cannot make a simple game that works well. Users who need game tools should look at other models instead of this one.
Website Frontend Test
The model made a start page for a SAS company. It looks nice but does not work correctly. Functions like scroll are not active on the page. Some parts of the page are also missing or broken. This shows the model cannot make interactive sites well.
We compared it with the GLM 5.2 model for this task. That model makes better frontend pages than Sonnet 5. Sonnet 5 is worse than its competition for this work. This is a problem for companies that build websites for their clients.
Businesses need pages that work for their users. A nice look is not enough for a good website. The model fails to give working parts that people need. AI w Biznesie recommends other models for building websites and apps.
SVG Graphics – The Worst Test
We tested making SVG graphics with this model. SVG is a format for drawing pictures with code. We asked for a BMW M4 CS car drawing. The model could not make a correct car shape at all. The proportions and shape are all wrong for the car.
The car does not look like a BMW at all in the result. Even in high quality mode the result is very weak. Opus 4.8 makes this task much better than Sonnet 5. It has correct proportions and a good shape for the car. Sonnet 5 is very bad in this creative area of work.
Even models like GLM 5.2 do better than Sonnet 5 here. This is a big disappointment for people who make graphics. The team at AI w Biznesie saw this as a major failure. Creative tasks need models that can understand shapes and style well.
Token Efficiency – The Worst Model Ever
The main benefit of Sonnet models was efficiency. They were meant to be faster and cheaper than Opus. But Sonnet 5 is very different from this goal. It is the least efficient model Anthropic has ever made. It uses a huge number of tokens for its work.
In max mode the model uses almost as many tokens as Opus. But it gives worse results than Opus does for the same work. This destroys the whole reason to use this model type. Users lose both time and money when they choose this model.
Maybe Anthropic rushed to release this model early. Perhaps they wanted to react to changes in the US government. But this does not explain such weak quality from a big company. The model is not useful for daily work tasks at all. It is better to wait for newer versions to come out later.
At AI w Biznesie we warn our clients about this model. The efficiency is the worst we have seen from Anthropic. Other models on the market give more value for less money. Business users should avoid this model for their important projects right now.
Summary – Do Not Use This Model
Claude Sonnet 5 is a failure for Anthropic. The model is not a step forward for users. It is worse than what we already have from the company. For example Opus 4.8 is much better in every way. The price difference is small but the quality gap is huge for all users.
The biggest surprise is the new token system for the model. The model uses more tokens and gives worse results back. This makes using the model not worth the cost. In creative tasks like SVG the model is very weak indeed. Other models are much better for this kind of work.
Our advice is simple and clear for all users. Do not use Sonnet 5 for your work. Stay with Opus 4.8 for better results every time. The model looks cheaper but only on the surface. In practice the costs are similar to Opus but results are worse for users.
Anthropic may release a better model in the future. Maybe that will be Fable 5 or Opus 5 for users. But for now Sonnet 5 is the worst model from this company. At AI w Biznesie we test many models for our clients. This one we do not recommend to anyone at all.
If you need AI tools for your work choose proven options. Models like Opus 4.8 are trustworthy and reliable for users. They give better results and are more efficient for work. Sonnet 5 is a waste of both time and money for everyone who tries it.
We hope Anthropic learns from this mistake soon. The company can do better for its users in the future. Until then smart buyers will choose other models for their work. AI w Biznesie will keep testing and reporting on new tools as they come to the market.
No responses yet