Three Agents, One Working Tree
A fleet running wide in one repository is a concurrency problem, and you already own the playbook. Isolation, claimed work, idempotent reruns, and where the analogy gives out.
By Jason Waldrip, with Claude Opus 5 07:40 8 min read
Three agents are running. One is fixing a bug in billing, one is adding a field to the onboarding form, one is partway through a dependency bump you approved without thinking about it. You come back and the test suite is red, and when you read the failure it doesn’t belong to any of them. You open all three diffs. Every one is something you would have approved on its own.
So you go work on the prompt. Everybody does. You add a rule to the guide in the repo, something about checking whether a file has changed before rewriting it, and you feel a little better because it’s a real instruction and it’s not wrong.
It helps at the margin. It doesn’t fix it, and the reason it doesn’t fix it is worth more than the afternoon it costs to find out.
nothing in there was reasoning badly
Go back and read the three diffs again with a different question. Not which one is wrong. Ask instead what each agent believed about the repository at the moment it started.
All three read the tree. All three read it at slightly different times. Each one then spent a while producing careful work against the state it had loaded, while the state it had loaded was quietly stopping being true. The billing fix was correct against a file the dependency bump was already rewriting. Nobody hallucinated. Nobody skipped a step. Each of them did competent work on a stale read.
You have seen this failure. You have probably fixed it. It just had a different name the last time, and it wasn’t wearing a chat interface.
the shape I already knew
At GigSmart I built a gig marketplace, and the interesting part was never the feature list. Workers and shifts move against each other in real time, which means the actual product is thousands of things happening at once without knocking each other over. I reached for Elixir because process isolation is the floor there rather than something you assemble later out of hand-written supervisors and queues you babysit. One slow thing shouldn’t be able to reach the things beside it.
That’s a runtime decision, and I’ve argued the general version of it before: scale is a design, not a dial. What sits underneath both is one line I keep on the distributed systems page, because it’s the whole discipline compressed: decide where state lives before you write a line.
I had that on a page on my own website while running a fleet of agents against a single checkout on my laptop.
the memo nobody sent
Somewhere in the last two years, without any particular day marking it, the job changed. You stopped being the person who writes the code. You became the person who operates a scheduler.
That’s not a promotion and it isn’t a metaphor I’m stretching for effect. It’s a literal description of what you do now. You admit work into a pool of workers, those workers run concurrently, they contend for a resource, and you are responsible for what happens when two of them want the same thing. There is a name for every part of that sentence and none of the names come from machine learning.
Which is exactly why it got past me. A checkout on a laptop doesn’t present itself as a system design decision. It’s just where the code is. Nothing about opening a third terminal feels like you’re widening a write path, and there’s no dashboard anywhere that turns red when you do.
the playbook transfers almost entirely
The useful thing about correctly naming the problem is that the answers were worked out decades ago by people solving it in much harder conditions.
- Isolation first. Git has shipped worktrees for years: multiple branches checked out in separate directories, backed by one repository. One tree per agent turns three writers on shared state into three writers on their own state, and the contention moves to the merge, where you have tools and a human. It is the cheapest change on this list and it takes about a minute.
- Make work items atomic and claimed. An agent should take a unit of work, own it, and be the only thing that owns it. At Phoneware I keep a workboard in the monorepo that the fleet actually tracks against, and its real job isn’t planning. It’s making sure two workers don’t pick up the same thing, and that a thing picked up has a visible owner.
- Assume the retry. A run gets interrupted, context runs out, you kill it because it wandered. If rerunning the same task twice produces two migrations or two config blocks, that’s not an agent problem, that’s a missing idempotency guarantee. Work should be safe to repeat.
- Name the serial section out loud. Some work was never going to run wide. A schema migration, a release, anything with an external clock on it. I’ve written about the things you can’t parallelize, and the failure is trying anyway rather than scheduling around it.
None of that is novel. That’s the point, and it’s the part I find genuinely good news. The tooling under a fleet is boring infrastructure that the industry already knows how to build, which means the constraint is recognizing what you’re looking at rather than waiting for somebody to invent a fix.
where the analogy gives out
I’d rather hand you the limits than a clean ending.
A thread is reproducible enough to chase. You can put the race under a debugger, run it enough times, and eventually watch it happen. You cannot replay an agent. Run the same task twice and you get two different reasonable attempts, so a race you hit this morning may never present the same way again. Detection after the fact is mostly off the table, which raises the price of structure and lowers the value of watching carefully. Prevention isn’t the better option here so much as the only one that compounds.
The second difference is the one that actually worries me. A thread that loses a race tends to fail loudly. It throws, it deadlocks, it corrupts something a checksum catches. An agent that loses a race writes fluent, well-organized, plausible code on top of a premise that stopped being true twenty minutes ago, adds a test that passes against its own assumptions, and tells you it’s done.
A crash announces itself. A confident wrong answer waits for you in review, formatted nicely, with a reasonable commit message.
That’s the cost of the whole arrangement, and it lands on the same place everything else does: you have to be able to tell. Volume is free now, verification is the job, and this is one more room where that bill comes due. The difference is that this particular failure is cheap to design out and expensive to catch by reading.
the thing worth taking
I don’t think many people are going to get burned badly by one agent. The model is good, the diff is readable, you’re paying attention because there’s one thing to pay attention to.
You get burned at three. And the reason you get burned at three is that going from one to three feels like nothing. No purchase, no migration, no design review. You open two more terminals on a Tuesday because the work is there, and you’ve quietly moved from a sequential system to a concurrent one without any of the ceremony that decision would have gotten if it had involved a database.
So before you widen the fleet again: what exactly is the shared resource, and what happens when two workers reach it at the same time?
If you can’t answer that, you haven’t got a prompting problem yet. You’ve got a scheduler you haven’t designed.
Written with Claude Opus 5, and said so. The Elixir and marketplace calls are from my time at GigSmart; the workboard and monorepo setup are how I run the fleet at Phoneware. The worktree pattern is plain git, not anything proprietary. The three-agent scene at the top is the shape of the failure rather than a transcript of one afternoon.