In partnership with

News of the day

1. Anthropic's Sonnet 5.5 outperforms GPT-6 Astra on specific benchmarks like Terminal-Bench, offering a compelling balance of cost and capability for defined tasks. → Read more

2. Alibaba's Qwen team unveils Qwen-Audio-3.1, a full-duplex voice model for agents, alongside significant price cuts on its audio stack. → Read more

3. Anthropic introduces a disciplined workflow for AI agent improvement, using Claude Code for evaluation and iterative refinement to ensure genuine progress. → Read more

4. AI model GPT-6 Astra assists in solving a complex fusion energy math problem, proving the existence of stable plasma equilibria. → Read more

Our take

Hi Dotikers!

The AI industry has spent two years selling us a simple story: bigger flagship, bigger price, bigger brain. Claude Sonnet 5.5 just poked a serious hole in that narrative, and honestly, it was about time.

Here is the headline number: Anthropic's mid-tier model, priced at $2 per million input tokens and $10 for output, scored 70.6% on Terminal-Bench 4.0. GPT-6 Astra, OpenAI's $10/$50 flagship, sits at 57.9% on the same benchmark. That is a 13-point gap in favor of the model that costs five times less. Turns out the intern just outperformed the executive, and the intern works faster too.

But the real story is more interesting than a leaderboard upset. Sonnet 5.5 does not beat Astra everywhere, and independent testing from Artificial Analysis confirms Opus still leads on factual and science-heavy tasks. What we are watching is capability fracturing by type of work. Sonnet shines when the job is well-scoped with a verifiable finish line: fixing a defined bug, producing a document, running agentic tasks in a terminal. Opus and Astra keep their edge on ambiguous, long-horizon problems that demand sustained judgment.

This is genuinely good news for anyone building with these tools. The question "which model is smartest?" was always a lazy one. The right question is "which model fits this specific task at this specific price?" And on that front, Sonnet 5.5 just made a lot of expensive API bills look hard to justify.

One caveat worth flagging: at maximum effort settings, Sonnet burned through roughly 193K output tokens per task in early testing, which can quietly erase the bargain. Read the fine print before you let it rip.

The era of one model to rule them all is ending. Choose your tools like a professional, not like a fan.

Alex.

Where Quantitative Thinkers Compete, Learn and Grow

The International Quant Championship (IQC) is one of the world's largest quantitative research competitions, bringing together 156,000+ participants globally.

Participants have the opportunity to develop quantitative research skills, challenge themselves alongside peers from around the world and connect with a global community of quantitative thinkers.

Meme of the day

Reply

Avatar

or to participate