Sakana AI's Fugu Ultra v2 Is a Model Made of Models: It Just Outscored Claude Opus 5 on One Benchmark
Sakana AI's Fugu Ultra v2 doesn't use a single frontier model in its own agent pool, yet it just outscored Claude Opus 5 on a real coding benchmark. Here's what it actually is, and what it costs.
Sakana AI shipped Fugu Ultra v2 on September 11, splitting what was previously a single Fugu release into two: a cheaper Fugu Max and this higher-end Ultra v2. The split itself isn't the interesting part. What matters is what Fugu actually is: not one trained network competing with GPT or Claude directly, but what Sakana itself describes as a multi-agent system packaged and priced as a single model.
That distinction matters more than it sounds. Sakana says Fugu Ultra v2 reaches its benchmark scores without Fable 5, Fable 5.1, or GPT-6-Astra anywhere in its own agent pool, meaning it isn't quietly routing hard questions to Anthropic's or OpenAI's frontier models and taking credit for the answer. Whatever it's demonstrating is coming from how it orchestrates its own, smaller components, not from borrowing someone else's.
Built on Two Papers, Not One Big Bet

Sakana isn't presenting Fugu Ultra v2 as a single research breakthrough. The company points to two papers behind it: TRINITY, on coordinating multiple large language models, and Conductor, on agent orchestration specifically. That framing matters, because it means Fugu's performance is a claim about a coordination method, not about one model finally getting big enough to win. If the method holds up, it should keep improving as the underlying LLM pool gets better, independent of whether any single lab ships a bigger model next quarter.
Fugu Max is the cheaper sibling launched alongside Ultra v2, aimed at the same orchestration approach at a lower cost for workloads that don't need the top-end scores.
The Numbers That Actually Matter
Across eight cited benchmarks, Fugu Ultra v2 lands best or joint-best on five: GDP.pdf, Chartography, SWEFish, DeepSWE, and Toolathon. Two of those are worth sitting with.

On DeepSWE, a benchmark built around real-world software engineering rather than isolated coding puzzles, Fugu Ultra v2 scores 74.3, a genuinely relevant number if you're evaluating it as a coding agent rather than a chatbot. On Chartography, it scores 48.3 against Claude Opus 5's 27.3, a wide enough gap that it reads as a real result, not a rounding error or a benchmark quirk.
Sakana is positioning this squarely at complex multi-step reasoning, autonomous research, and full-stack software development, not general conversation. If you're shopping for a coding agent rather than an assistant, this is the kind of release that's supposed to matter more than its name recognition suggests right now.
The Pricing Has a Catch Worth Knowing
Fugu Ultra v2 is priced at $5 per million input tokens, $30 per million output tokens, and $0.50 per million cached input tokens, for contexts up to 272,000 tokens. Cross that threshold and the rate jumps to $10, $45, and $1.00 respectively.
That's worth knowing before you build anything that might occasionally run a long context window. Plenty of models charge one flat rate regardless of how much context gets used; Fugu Ultra v2 doesn't, and that step-up is easy to miss if you're only comparing the headline price on a chart.
Already Wired Into Real Tools
Fugu Ultra v2 isn't just a benchmark exercise. It's already available through OpenRouter, and Sakana lists integrations with Vercel, OpenCode, Creao, and Merge on its own launch page. That's a meaningfully different signal than a lab publishing a paper and a leaderboard score. Actual integrations mean developers are already routing real workloads through it, not just running it against a benchmark suite once for a launch post.
Why This Is Worth Watching, Not Just Reading About
The bigger story here probably isn't Fugu itself. It's the approach: a smaller lab getting frontier-competitive results by orchestrating multiple smaller components well, instead of training one enormous model and hoping scale gets it there. If that keeps working as well as this week's numbers suggest, it's a real alternative path for labs that can't afford to train the next GPT-scale model from scratch, and worth tracking regardless of whether Fugu specifically becomes the model people actually adopt.
Have you run Fugu Ultra v2 on a real coding task yet? We're curious whether the DeepSWE number holds up outside a benchmark. Let us know in the comments.
Related News
The RAM Crisis of 2026: Why Every Laptop and Phone in India Is Quietly Getting More Expensive (Or Worse)
DRAM contract prices have surged by up to 98% in a single quarter as AI data centers consume the world's memory supply, and the fallout is already hitting laptop and smartphone prices across India — through outright hikes and, in some cases, quiet RAM downgrades.

