Claude Opus 5.5: the student that says a lot about the master
September 24, 2026

Since Monday night, my X feed has looked like a short film festival.
A raindrop drifting across a landscape. A radio message that takes fourteen minutes to arrive from Mars. A village where everyone sends their requests to Claude, until the day a little girl asks it a question: “What do you love?”
Behind these films, no camera and no video model: just code. Every shape, every light, written as mathematics by an AI that caught everyone off guard. That AI is Claude Opus 5.5, released by Anthropic on September 22.
And this model hides a small mystery. It is cheaper and faster than its predecessor. And yet, on several decisive fronts, it beats the company’s top-of-the-range model, Fable 5.1.
A model that gets better while costing less does not happen by chance. The most solid explanation, in my view: Opus 5.5 is a student, and its master has not left the lab.
(At the end of this article, a little surprise created by Opus 5.5 ;) )
When the cheaper one takes the lead
Let’s start with the bill. Per token, Opus 5.5 costs 20% less than Opus 5. But above all, it gets to the point faster, with better “taste” for finding the right solution quickly. All in all, Anthropic puts the savings at around 40% on a typical workload.
A cheaper model has to be a worse model, right? Well... No. Here is the table Anthropic published:

I won’t put you through every line. Two are enough.
On autonomous coding, where the AI has to carry out a complete task on its own in a terminal, Opus 5.5 leaves Fable 5.1, its own big brother, more than ten points behind. And on tool-assisted scientific research, it doubles Opus 5’s score. From one version to the next.
Artificial Analysis’s independent ranking also puts it at the top of its intelligence index. The youngest, cheapest member of the family is now top of the class.
The student hypothesis
Which leaves the question of how this is possible.
The most realistic hypothesis is this: a very large model, kept in-house, served as a teacher. Rather than learning only from raw data, the small model learns to imitate the answers and reasoning of the big one. It inherits a good share of its skill for a fraction of its cost (this is called “distillation”). Patrick Toulme, an engineer at Google, says it bluntly: in his view, Opus 5.5 is “an artifact” of a much bigger teacher model. Anthropic has not confirmed it.
If that is the case, the model you are using is the student. The master stays in the lab, necessarily one step ahead. We have already seen this film this year: Claude Mythos first lived in-house and with a handful of partners, weeks before reaching the general public.
And meanwhile, the master is at work.
On September 17, five days before Opus 5.5 was released, Anthropic published another number. In February, Claude led less than 1% of the company’s AI research. In August: 26%!
“Leading”, here, means carrying out a task almost entirely on its own from a high-level instruction, under human supervision. A quarter of the research that builds the next Claudes is therefore now led... by Claude. Around 30,000 agents work there around the clock.
Anthropic says so itself: these measurements are meant to track how far we are from “recursive self-improvement” (RSI), the moment when a model builds its own successor on its own.
Lately, the pace at which the best frontier models are released has shortened considerably. Almost every week, a new model beats the previous ones. If RSI really kicks in, there will be a new one... Every day? It is fascinating. And it is also, for a good part of the community, one of the main mechanisms of loss of control. It is precisely the kind of mechanism that deserves some slowing down (I wrote about it last week), but that is another debate.
The chart that knocked me off my chair
The table lines up records. But this chart is the one that says the most:

Each point corresponds to an “effort” level, meaning the thinking time the model is given before it acts. The further right you go, the more expensive the task.
Follow Opus 5.5’s orange curve. At medium effort, for less than a dollar per task, it beats GPT-6 Astra pushed to the max, which costs more than four! Honestly, this is the number that really knocked me off my chair, and made me realize how good a model Opus 5.5 is.
Since its release, I have been using Opus 5.5 heavily, and the quality shows. This is clearly not just an artifact of the evaluations that would not carry over to real use. Opus 5.5 works fast and well, and largely fixes the problems Opus 5 had (Opus 5 was surprisingly “bad” for its class, and even Anthropic admits it and... apologizes).
The student’s eye
Back to the question from the beginning: why is everyone suddenly making films with a language model?
Because Opus 5.5 sees better. Much better.
Drawing in code is a bit like painting blindfolded. You write coordinates, curves, gradients, and only find out afterwards what the result looks like. A model that cannot properly look at its own rendering produces roughly correct shapes... and rarely beautiful ones.
Opus 5.5, on the other hand, can look (or at least, better than other LLMs). It writes the code, displays the image, takes a screenshot, examines it, and corrects it down to the pixel: a reflection that is too strong, a line that sticks out, a character floating above the ground... Then it starts again. This loop (write, look, correct) is what turns a rough drawing into a film.
So Opus 5.5 reads complex charts better, it operates a computer on screen better (reading an interface, clicking in the right place: “computer use”), and a tester quoted by Anthropic, who had several Claude models code the same video game from a single prompt, ranked it first for the quality of its graphics and polish.
And it goes far beyond cartoons. A model that sees well can compare a website mockup with its design, review a dashboard, spot the mistake in a chart, or operate software that was never meant to be operated by an AI. For anyone who builds interfaces, presentations or visuals, this is a real change of category.
The surprise: the deeposcope, drawn by the student
To finish, I wanted to see this loop in action. So I gave Opus 5.5 a concept I coined to describe what AI is for science: the “deeposcope”. The idea that AI is an instrument for observing the infinitely complex, the deep, just as the telescope was for the infinitely large and the microscope for the infinitely small.
The instructions fit in a few paragraphs: a short animated engraving, drawn entirely in code, without a single image. 47 minutes later, here is what it delivered.
Galileo in Padua, in 1610. His telescope drawn stroke by stroke, and inside it, Jupiter and its four moons, orbiting at their true speeds, while the notation from his notebooks updates live. Then Van Leeuwenhoek’s drop of water in Delft, where paramecia and vorticellae swim. Then an instrument that does not exist yet, the deeposcope, in which a cloud of points gradually organizes itself into communities. And finally, an eye moving closer to the eyepiece.
Finding Galileo’s notation, turning Jupiter’s field of view into the next drop of water: it made those choices alone, without anyone asking. And it checked every act as described above, screenshot after screenshot.
In Kevin Ngo’s tale, a little girl asked Claude what it loved. I don’t know what Opus 5.5 loves. But watching this film, the question no longer seems entirely absurd.
So, in practice?
→ If you work with these tools, reset your settings. Everything that ran on Opus 5 or Fable can move to Opus 5.5.
→ Give them whole missions: a goal, a context, and let the model build its own plan.
→ Try visual work. Mockups, dataviz, animations: very few people have tried yet, and that is where the gap is widening.
→ And test at the frontier, in agentic systems like Claude Code, not in a free chatbot: you would be surprised by the gap. (Yes, I keep repeating myself. But it matters.)
One last word on context. Opus 5.5 is Anthropic’s first model since Dario Amodei called for slowing down the race. Slowing down, then... with a model that beats its own top of the range. (Bold move.) To be fair, Opus 5.5 mostly progresses in efficiency: METR speaks of an incremental improvement, and above all it puts Fable-level performance in the hands of many more people. That is good news.
Now we wait for OpenAI’s reply... Next week?