On Terminal-Bench 2.1 the routed fleet scores 82.0% — the best published result on Opus 4.8, and 5.2 points above running one model on every unit.
Audited rules, no test reads, no web
Native lanes
And anything OpenAI-compatible
A single agent run is a few hundred decisions of wildly different worth. Most developers point their best model at all of them.
Loom predicts which model each unit of work needs, then spends what it saved on verifying the answer.
The router is a loopback endpoint your existing harness already knows how to talk to. No provider key ever enters your shell.
On Terminal-Bench 2.1 the routed fleet scores 82.0% — the best published result on Opus 4.8, and 5.2 points above running one model on every unit.
Turning the verify→repair loop on doubled the pass rate on identical tasks — 0.333 to 0.667. That loop is what the routing savings buy.
Same weights, different scaffolds: stock Claude Code scores 78.9% and Terminus 2 scores 74.6%. The harness, not the model, sets the ceiling.
1— A cheap answer is never accepted for being cheap. An adversarial verifier runs the real checks, and anything that fails climbs a rung and is re-run higher.
2— Lanes are configuration, not code. Pin a slot — or the whole fleet — to one vendor and you keep the metering, the verification and the receipts.
3— It fails toward the expensive model. Missing key, upstream 500, malformed stream: the unit falls up the ladder, never off it.
Modelled at equivalent output quality, escalations included.
Start saving with intelligent model routing
Download Loom