Skip to main content
Matt Sears

I built an orchestrator for my agents. They outgrew it.

In June I wrote about building shipyard, an autonomous engineering loop on Claude Code. In July I wrote that everyone else had built the same one and that a chunk of my plumbing had quietly become a native feature.

In September I went and deleted it.

What an orchestrator was for

Shipyard existed because a May-2026 model couldn't be handed an issue and trusted to finish it. So the harness did the managing. It picked which issue to work. It refined the raw text into something dispatchable first. It decided an epic was too big and sharded it into ordered sub-issues. It told the worker what order to do things in, and it checked the worker's homework afterwards.

That's a lot of scaffolding, and all of it earned its place at the time. It burned down a backlog of a couple thousand issues.

Then the models got good enough that the managing stopped helping.

Everyone is arriving here at once

Robert C. Martin described building an orchestration harness for "a squad of agents" on 4 August — analysts, reviewers, implementers, reviewers, the lot, driven by a state machine with a simulator over it. He signed off: "Don't tell me software engineering is dead."

Five weeks later, he wrote:

And while I was heads-down getting that to work, the agents got a LOT better. So much so that when I came up for air, the need for my harness was obviated.

Anthropic's own team says the same thing with a number: they cut about 80% of the Claude Code system prompt, and now ship a different system prompt per model, because only the frontier ones can go without the training wheels. Boris Cherny, who built Claude Code, tells people to delete their CLAUDE.md, their skills and their hooks every six months and see what the model does without them.

Nobody coordinated this. It's just what happens when the thing underneath your abstraction gets better than the abstraction.

What we deleted

Four changes, about 11,000 lines:

  • A whole second dispatch system, kept as a "reversible alternate." It only ever existed because its primitive couldn't isolate a worker into its own git worktree — something the platform now does for free.
  • The issue refiner. It classified and rewrote raw bug reports into dispatchable specs in a pass before the work started. A current model does that while reading the issue, in the same breath as understanding it.
  • The epic decomposer. Same story: a model asked to work a large issue shards it as part of planning it. It doesn't need a separate agent and a config knob to be told.
  • A context-bloat linter whose entire job was policing the size of my own prompt corpus — which is a fairly good sign the corpus had a problem.

Plus half a megabyte of design notes that were sitting in the directory the orchestrator reads specs from, one careless file read away from eating a context window. Those moved rather than died.

What's still worth keeping

I expected to keep going and take out most of the rest. I didn't, and the reason is the useful half of this story.

The orchestration was obsolete. The verification wasn't.

There's a script in there whose only job is answering "did CI actually pass?" It exists because gh run list --commit <sha> silently matches zero runs if you hand it an abbreviated SHA — and exits 0. So the obvious three-line version of that check passes on the empty set. It reports green having looked at nothing at all.

There's another that writes state to a temp file and renames it, so a process killed mid-update leaves the previous state intact instead of a corrupted one. Another that caches GitHub reads because the loop was asking the same question four times a session.

None of those care how smart the model is. They're not managing the model at all — they're catching the world being weird: an API that returns success for a query that matched nothing, a process that dies mid-write, a rate limit. Each one is a thing that went wrong once, compressed into a check that won't let it go wrong again, and a better model doesn't make any of them less true.

If you're doing this to your own harness, that's the sorting question I'd use: is this compensating for the model, or for the world? The first kind expires. The second kind doesn't.

The line I'd draw

My July post said the plumbing converged and the judgment didn't. I'd sharpen that now.

The parts of a harness that go obsolete are the parts that were compensating for the model: do this before that, don't forget the other thing, here's the shape of the output I want. Those were always a workaround, and they expire the moment the workaround isn't needed.

The parts that stay are the parts that encode a reason. And those aren't scaffolding at all — they're scar tissue, and you don't remove scar tissue because the patient got healthier.

I built an orchestrator for my agents. They outgrew it. — Matt Sears