Announcing 82.0% on Terminal-Bench 2.1 — the best Opus 4.8 result

Understanding is the product Understand every change.

Loom runs a fleet of coding agents on one goal and gives you back the part that matters: what was intended, what actually landed, and what a rival model could not break.

$ pnpm loom:run “ship metered billing”

Part of the startup programs at

CombinatorStartup School/AWS Activate/Microsoft for Startups/OpenAI/Anthropic

Used by engineers at

The run
Vercel
Stripe
The reportDatabricks
ramp

Conductor. Run a fleet on one goal

Fig.1

01

Route

Every role gets the cheapest model that can actually do that role’s job — and the frontier still wins the hard rows.

Open the calculator

1.1A model per role, not one model per project.

1.2Your keys or ours — GPT, Claude and DeepSeek on one fleet.

1.325× cheaper on the parts that never needed the frontier.

Fig.2

02

Verify

The checks are authored red before the work starts, and a rival model is paid to try to break the result.

See the review

2.1A red baseline is authored before any code lands.

2.2A different model runs the gates and files refutations.

2.3No model grades its own homework.

Fig.3

03

Weave

Six branches clear one gate and weave into one trunk. Nothing waits for a merge day, because the gate runs first.

How the weave works

3.1The gate runs before the merge, not after it.

3.2A branch that fails rebases itself and re-enters the queue.

3.3Six sessions land as one reviewable pull request.

Fig.4

One goal is staged across six roles over thirty seconds. Two builders claim a lease and start almost together; a third stages in mid-run; the verifier authors its checks early; the tester waits for the verifier; and the critic is held until last, because there is nothing to refute until the rest have finished.

04

Conduct

One goal in. Loom stages a fleet around it — builders first, then the roles whose whole job is to check them.

See the app

4.1One goal in, a staged fleet out — never six at once.

4.2Dispatch is a lease: one task, one worker, never two.

4.340+ agents across six roles, supervised end to end.

Fig.C — drawn live, one clock

Six in flight. One goal, six Claude Codes

A repo is a maze with more than one way in and more than one way out. One session takes them one at a time — finish, reset, start the next on a cold context. Conductor dispatches six Claude Code sessions onto six routes in parallel and lands them on one line.

  • 01Six real sessions. Not six tabs — six Claude Code processes under one conductor, dispatched from one brief.
  • 02Own worktree, own lease. Each session writes in isolation, so two runners can never take the same corridor.
  • 03Chorus merges them back. Six routes land as intents on one canon line — byte-identical in any order.

The app. The shipped components, running right here

Four surfaces, live in this page, on one 48-second run: a brief filed, the diff reviewed a file at a time, four slices merged, and the graph that recorded all of it. The line under each frame says which beat you are watching.

Fig.5 — live

Loom ConductorReal UI
Loading the real app…
05

The shell

A brief gets filed, the fleet is dispatched, the gates flip. One surface, mid-mission, with nothing staged for the camera.

Download for macOS

5.1The same components the desktop ships, compiled for the browser.

5.2Session rail, conductor and inspector on one surface.

5.3Every session is a real Claude Code process under the hood.

Fig.6 — live

Review deckReal UI
Loading the real app…
06

The review deck

One file per card. Swipe through what the fleet changed, with the intent for every hunk beside it — approve, hand back, or let it land.

Read the record

6.1One card per file — you never review a wall of diff.

6.2Approve, hand back, or let it land — your call, per file.

6.3The intent is committed with the code, not lost in a thread.

Fig.7 — live

Chorus — mergeReal UI
Loading the real app…
07

The merge engine

Six writers, one file, no conflicts. Chorus resolves before the merge rather than making you resolve after it.

How Chorus works

7.1Six session threads clear one gate and land as one trunk.

7.2Nothing merges that has not cleared the gate first.

7.3A rejected thread rebases itself and re-enters the queue.

Fig.8 — live

Mission graphReal UI
Loading the real app…
08

The mission graph

One node per session, one task tag clamped onto each thread. Drill into any node and the transcript that produced it opens beside it.

Graph engineering

8.1The graph is the plan and the record at the same time.

8.2A stuck session can be steered without stopping the fleet.

8.3A refuted node goes back green only after it earns it.

LOOM CONDUCTOR — THE VERIFIED FLEET

One prompt. The fleet ships it. Verified.

Loom stages a fleet around your goal — builders first, then a verifier that writes the checks, a tester that runs the gates, and a critic that tries to break the result. Every role gets the right model: DeepSeek where it’s tiny and 25× cheaper, GPT where it’s tricky, Claude Fable where it’s hard.

Best Opus 4.8 result on Terminal-Bench 2.1 · Certificates, not vibes · Apple Silicon · 12 MB · auto-updates

scroll — watch one prompt become a shipped mission

● DISPATCH → SESSION_1 · THE RATING ENGINE DISPATCHEDRECORD No. 0519-1 · LEASE 26-07-17 · 0x5F19
CLAUDE FABLE-5
ANTHROPIC · FRONTIER · BUILDER — “CLAUDE FABLE WHERE IT’S HARD”
Long-horizon edits, exact arithmetic paths, holds the whole billing model in its head. The hard core goes here.
BENCHMARKS · PUBLIC BOARDS · JUL 2026
SWE-BENCH VERIFIED▃▅▆▇█84.2
TERMINAL-BENCH 2.1▂▄▅▇█82.0
GPQA DIAMOND▄▅▆▇█88.1
AIME 2026▅▆▇██96.4
MMLU-PRO▃▄▆▆█89.9
$5.00 / $25.00PER 1M TOKENS IN · OUT
CONTEXT200K MAX OUT64K SPEED68 TOK/S CUTOFFJAN 2026 RELEASEDJUN 2026 PARAMSUNDISCLOSED EST. THIS ROLE$1.84 WORKTREEsessions/s1
api.anthropic.com/v1/messages · RPM 4K · TPM 2M · TEXT + VISION + TOOLS · PROMPT CACHE 0.1× READ · STOP-FENCE ON · INTENT ANNOTATIONS LIVE
● DISPATCH → SESSION_2 · THE CRITIC, HELD TO BREAK IT DISPATCHEDRECORD No. 0519-2 · LEASE 26-07-17 · 0x5F2A
GPT-5.6 SOL
OPENAI · FRONTIER · CRITIC — “GPT WHERE IT’S TRICKY”
Adversarial by temperament. Tries to reproduce failures the builders can’t see — and every refutation must reproduce.
BENCHMARKS · PUBLIC BOARDS · JUL 2026
TERMINAL-BENCH 2.1▄▅▆▇█88.8
BROWSECOMP▃▅▆██90.4
GPQA DIAMOND▄▆▆▇█94.6
FRONTIERMATH T1-3▅▆▇▇█89.0
SWE-BENCH PRO▁▃▄▅▆64.6
$5.00 / $30.00PER 1M TOKENS IN · OUT
CONTEXT1.05M MAX OUT128K LONG CTX>272K: $10/$45 CUTOFFFEB 2026 RELEASEDJUL 2026 PARAMSUNDISCLOSED EST. THIS ROLE$1.86 WORKTREEsessions/s2
api.openai.com/v1/responses · TEXT + VISION + TOOLS · CACHE READ 0.1× · CACHE WRITE 1.25× · REFUTATIONS MUST REPRODUCE · HELD FOR LAST
● DISPATCH → SESSION_3 · THE MIGRATION DISPATCHEDRECORD No. 0519-3 · LEASE 26-07-17 · 0x5F3B
DEEPSEEK-V4
DEEPSEEK · OPEN WEIGHTS · BUILDER — “DEEPSEEK WHERE IT’S TINY AND 25× CHEAPER”
Mechanical work in 10k-row batches. Same gates as everyone else — at a twenty-fifth of the output price.
BENCHMARKS · PUBLIC BOARDS · JUL 2026
SWE-BENCH VERIFIED▂▃▅▆▇76.8
TERMINAL-BENCH 2.1▁▃▄▅▆71.2
GPQA DIAMOND▂▄▅▆▇79.4
AIME 2026▄▅▆▇▇91.2
MMLU-PRO▃▄▅▆▇84.1
$0.44 / $0.87PER 1M TOKENS IN · OUT
CONTEXT128K MAX OUT32K SPEED61 TOK/S CUTOFFSEP 2025 RELEASEDAPR 2026 PARAMS685B MoE · 37B ACT EST. THIS ROLE$0.19 WORKTREEsessions/s3
api.deepseek.com/v1 · RPM 6K · TPM 10M · TEXT + TOOLS · OFF-PEAK −50% · SAME GATES, SAME CHECKS · 25× CHEAPER OUTPUT

THE RECORD · WHY EDITS CARRY INTENT

Every edit files
its intent.

The fleet annotates the code as it writes it — every changed line carries the reason it exists. Then the engine reads those intents back and grades the work against them. Not vibes. A grade.

test_greet.py§ 3 · annotated 10:18 AM[LIVE] INTENTS
3/3§ 1§ 2§ 3
INTENT
§ Pulls in pytest for its raises assertion helper and imports greet from th…
1import pytest 2 3from greet import greet
§ Locks the happy path exactly as the README promises: a normal name and the literal "world" both must round-trip through the "Hello, {name}!" template…
6def test_normal_name(): 7 assert greet("Yiming") == "Hello, Yiming!" 8 10def test_default_world(): 11 assert greet("world") == "Hello, world!"
§ Enforces the rejection half of the contract, and covers whitespace-only input separately from the empty string so the implementation cannot pass by a bare truthiness check — it must strip befor…
14def test_empty_name_raises(): 15 with pytest.raises(ValueError): 16 greet("") 17 19def test_whitespace_name_raises(): 20 with pytest.raises(ValueError): 21 greet(" ")
✓ 3/3 EDITS MATCH THEIR FILED INTENTGRADE A · ADMITTED TO MERGE

THE ENGINE RE-READS EVERY INTENT AT MERGE TIME · DRIFT FAILS THE GATE

claude-fable-5
builder · $5 / $25
gpt-5.6-sol
critic · $5 / $30
deepseek-v4
builder · $0.44 / $0.87

The fleet is working.

3 SESSIONS STAGED · CHECKS AUTHORED FIRST · NO MODEL GRADES ITS OWN HOMEWORK

The position

No model grades its own homework.

LoomThe verification engine

The evidenceThe run

One model, two harnesses on Terminal‑Bench 2.1: 78.9 on Claude Code, 82.0 on this one. The tasks it still failed are published with the run.

MingLLMThe position we build from

Graph engineering. Control flow that is software

Fig.10

  • 2022 — Prompt. Few-shot, chain-of-thought. The instruction. One turn, tuned by hand.
  • 2025 — Context. RAG, memory, tool defs. What the model can see.
  • February 2026 — Harness. Sandbox, permissions, AST gates. What the model can touch.
  • June 2026 — Loop. Plan, act, verify, repair. How one agent iterates.
  • July 2026 — Graph. Nodes, edges, state, conditions. Where control goes next. You are here.
10

The ladder

Prompt, context, harness, loop, graph. Five rungs of who decides the path — and the one this is standing on.

Read the run

10.1A harness is the floor a graph runs on, not the ceiling.

10.2Loom scored 82.0% on Terminal-Bench 2.1 as a harness alone.

10.3The same model on Claude Code scores 78.9%.

Fig.11

11

The cast

In a graph, “agent” stops being one thing. Six kinds of node do the work — and the ones that make a run trustworthy mostly aren’t models.

See the wiring

11.1Session, router, judge, human gate, function, policy.

11.2Only two of the six are a model call.

11.3Each one already ships as a part of Loom.

Fig.12

12

The wiring

A loop becomes a graph in four moves: give each step a kind, decide where control goes next, thread one typed value through, then put the condition in code with a bound the executor enforces.

Read the spec

12.1A loop is a graph with one node and an edge back to itself.

12.2It works right up until the agent is the only thing that knows why it stopped.

12.3Every decision here points at the code and the state value that made it.

Honest caveat. None of this is new. Graph-structured agent runtimes have shipped since 2024, and the ancestry runs back through DAG schedulers, state machines and Airflow. A graph does not replace a loop either — a cycle is a loop, so a graph contains one. No benchmark comparing the two exists that we can find, and we are not going to pretend otherwise. The label is contested: some people mean workflow graphs, others mean knowledge graphs. We mean workflow graphs, and only that. What we claim is narrower and checkable — every decision on this page points at the code and the state value that made it. Sources: ExplainX 2026-07-18 · LangChain, “3 years of graph engineering” · AIBuilderClub hype check · SmartScope on containment.

> > > > > >

CHORUS · THE MERGE ENGINE

Six writers. One file.
No conflicts.

Every session works in its own worktree; Chorus lands their changes as intents, re-runs the gates on each landing, and advances the canon — byte-identical in any order. Below: the pipeline itself — six threads through one gate into one canon line. It’s why the product is named Loom.

FIELD BASELINE: 27.7% OF AGENT PRS LAND WITH MERGE CONFLICTS (107K PRS, ARXIV 2604.03551). EVERY OTHER TOOL ISOLATES WRITERS OR PICKS ONE WINNER — CHORUS MERGES THEM. DRAWN LIVE, NOT A VIDEO.

>

The delta. Price the fleet against your habit

Fig.14 — interactive

Pick the model you run today and how many tokens you burn a month. The Conductor re-prices the same work routed across the roster — small jobs to the cheap seats, big builds to the front-runner. Or skip the guessing and drop a real transcript.

01 · THE MODEL YOU RUN TODAY

04 · OR STOP GUESSING

DROP A SESSION TRANSCRIPT .JSONL FROM CLAUDE CODE, OR ANY LOG — OR CLICK TO CHOOSE. COUNTED IN YOUR BROWSER; NOTHING UPLOADS UNLESS YOU SEND IT.

02 · TOKENS PER MONTH 30M

THE NUMBER ON YOUR PROVIDER’S USAGE PAGE. 5M ≈ A LIGHT WEEK · 100M ≈ A HEAVY MONTH.

03 · WHAT DOES THE WORK LOOK LIKE?

THE READOUT
QUALITY70.182.0+11.9 PTS
SPEND / MO$180$158−12%

CHEAPER AND SHARPER — NO TRADE.

MODELED · LIVE CATALOG
>

82.0TERMINAL-BENCH 2.1

MEASURED, NOT CLAIMED

The best Opus 4.8 result on Terminal-Bench 2.1 — 82.0% with the same model that scores 78.9% on Claude Code. The whole run is public: every task, every transcript, every failure.

READ THE RUN →

Research. The runs and the reports

Sam Altman: there's a betting pool for the first year there is a one-person billion-dollar company.
2026

Will you make this the year it happens?

THE FILM

For those who dared.

Built for people who bet on themselves.

BEFORE YOU ASK

Questions, on the record.

Which models does the fleet run?+

Any of them. Claude runs the harness; the docket routes each task to Claude, GPT, DeepSeek, or Gemini by price and benchmark — your API keys or Loom credits.

Does my code leave my machine?+

Loom is a desktop app. Sessions run in local worktrees on your hardware; model calls go to the providers you configure. Loom Cloud is optional and self-hosted on your own mini-server.

What does it cost?+

The app is a free 12 MB download. Bring your own keys and pay providers directly, or use Loom credits — tiers are credit limits, nothing else. See pricing.

What do I need to run it?+

A Mac with Apple Silicon. The app auto-updates from a signed feed. Windows and Linux aren’t available yet — we won’t pretend otherwise.