Other Blogs
Claude Opus 5 landed, so I pointed the eval suite at it. 23 scenarios, replayed through the actual chat agent with the real CAD tools, nothing mocked. The methodology is here if you want the details.
I ran it twice, because the first run made me suspicious. Same scenarios, same prompts, same afternoon. The only thing I changed between them was the reasoning effort: one line of config.
That one line mattered more than the choice of model.
| Metric | Opus 5 (low) | Opus 5 (high) | Fable 5 (low) | Gemini 3.6 Flash |
|---|---|---|---|---|
| Deterministic checks | 84% | 69% | 81% | 85% |
| Code errors | 0 | 5 | 2 | 8 |
| Total cost | $13.99 | $19.39 | $12.17 | $5.23 |
| Avg duration | 116s | 275s | 94s | 107s |
Not "more expensive for a small gain." Worse. Low effort was 28% cheaper, 2.4x faster, and went from five code errors to zero, while the judge's adherence score barely moved (0.92 to 0.90, which is noise at this sample size).
The failure pattern explains it. At high effort, 30 of the 33 failed checks were cost or duration budget violations rather than wrong answers. The model wasn't getting the geometry wrong. It was thinking itself past the wall clock. Two multi-turn scenarios blew straight through the 10 minute per-generation timeout and got killed mid-run. At low effort, no timeouts, and only one cost failure left.
I've hit this before. Back in the GPT-5 sweeps, medium reasoning beat high for about half the cost. Same shape, except this time the top of the ladder didn't just fail to earn its tax, it actively lost.

My first instinct was that Fable 5 had walked away with it, since Fable's low-effort run was cheaper and faster than Opus 5 at high. That comparison was junk. I was putting one model's low effort against another's high.
Matched properly, at low effort on both, they're near parity: Opus 5 costs 1.15x and takes 1.22x the wall time, and scores slightly better on deterministic checks with zero code errors against Fable's two. Worth noting Fable's list price is double Opus 5's, so Opus is still burning more tokens per job. Just not enough to change the answer.
Gemini 3.6 Flash stays the default for now. It's roughly 2.7x cheaper than Opus 5 at low effort with a comparable pass rate, and cost per generation is the number that decides this for a product where people iterate.
The thing tugging the other way is the error column. Gemini threw 8 code errors across the suite, Opus 5 at low effort threw zero. A code error isn't just a retry on our side, it's a user watching a spinner while the agent repairs itself. I'm not switching on one 23 scenario run, but that gap is what I'll be watching.
Every run, including the ugly ones, is logged on the evals page. If you want to see what the models actually produce, the 3D viewer will open any result you export.