im·a·cto

From the log

Back to the log

Review Is the Rate Limit

Generation got cheap and approval didn't, so the narrow stage moved and I spent a year widening the wide one. On queue depth, demand-driven pipelines, and the difference between throttling the producer and actually widening review.

There are two finished posts sitting in a queue on this site right now. Written, formatted, images generated, ready to render. Each one waiting on exactly one thing, which is me reading it and saying yes or no. Nothing is broken. Nobody is blocked. And the thing that produced them will cheerfully produce another tomorrow without ever asking how the first two went.

I set the generation side up on purpose. Scheduled, grounded in my own material, a guide in the repo so it stays in my voice instead of drifting into LinkedIn. It works, which is the part I was proud of.

What I never did was size the other half. I have been getting paid to fix that exact mistake in other people’s companies for twenty years.

the stage nobody sizes

Every pipeline has a narrowest stage, and that stage sets the throughput of the whole thing. Work you do upstream of it doesn’t raise output. It just gets the work to the narrow stage sooner, where it waits.

This is the most reliable piece of systems thinking I own. I apply it to build pipelines, deploy paths, marketplaces, onboarding flows, hiring loops. Then generation went to roughly free, and I spent a year happily widening a stage that was already the widest one I had.

So ask the question properly. For most of what I build now, the narrow stage is one person deciding whether the output is any good. That decision was never the fast part, and it is the one part of this whole shift that didn’t get cheaper.

what I actually did at CommercialTribe

At CommercialTribe I replaced a weeks-long release cadence with automated testing and deploys that shipped in hours, across onsite engineers and a remote team in Argentina. That’s the line on my works page and it’s accurate, but the way I usually tell it buries the interesting part.

I didn’t make anybody write code faster. Nobody’s typing was ever the problem. Releases were the problem, and releases were slow because shipping required a specific set of humans to agree, in a specific window, on a specific afternoon. So I went at the narrow stage directly: automate the checks a person was doing by hand, and leave in the loop only the judgment that genuinely needed a human attached to it.

Weeks to hours is the difference between a team that learns and one that guesses. But notice which project I picked. I went after the approval stage. I did not hire two more engineers and let the releases stack up behind the same afternoon.

a growing queue is not output

I’d like to tell you the two drafts are harmless. Queues are too well understood for me to be vague about it.

When work arrives faster than it gets served, the queue does not settle at some comfortable larger size. It grows until something gives, and the wait on any single item grows along with it. That’s not a productivity style. It’s arithmetic.

The items don’t wait neutrally either. A draft written against this month’s thinking goes stale. A branch written against a moving codebase quietly accrues conflicts. DORA keeps “work in process limits” in its delivery capabilities for this reason, and describes it as limiting how much people are working on so a small number of high-priority things actually get done. Read that as inventory management, because that’s what it is. Work in progress ages, and aging inventory depreciates.

Two drafts, zero decisions.
A pile of pending work reads like output on a calendar. Nothing in it has been judged yet, which was supposed to be the whole value I add.

There’s a quieter cost under the obvious one. The queue let me not decide. If each of those posts had forced a yes or no on the day it landed, I’d have two answers now. Instead I have a stack, and a stack feels like having done something.

the topology I’d reject in a design review

At GigSmart I chose Elixir because a real-time marketplace is a concurrency problem before it’s anything else: fault isolation by default rather than bolted on afterward. That call still holds up, and the piece of that ecosystem I think about most often isn’t supervision trees. It’s demand.

In GenStage, Elixir’s library for staged pipelines, the consumer tells the producer how much it can take, and the producer is not permitted to exceed it. The documentation is blunt about the mechanism: “Once demand arrives, the producer will emit events, never emitting more events than the consumer asked for. This provides a back-pressure mechanism.” The consumer pulls. The producer cannot push.

The same idea runs through TCP windows, bounded channels in Go, and every work queue I’ve ever put a depth limit on. An unbounded producer feeding a fixed-rate consumer is a known-bad shape. If a team walked me through that design I’d send them back inside of five minutes.

It is also, precisely and embarrassingly, how I had my own writing configured. A scheduled producer with no knowledge of consumer state, pointed at one reviewer with a calendar.

backpressure when the consumer is you

  1. Bound the queue in the system, not in your intentions. Pick the number of items allowed to sit awaiting your judgment. When it’s full, generation doesn’t run. A limit you enforce by remembering to feel bad about it is a preference, and preferences don’t apply backpressure.
  2. Let review pull the work. The generator shouldn’t be asking what day it is. It should be asking whether there’s room. Scheduling by calendar is push. Scheduling by queue depth is pull, and pull is the entire mechanism.
  3. Watch age, not count. Three items that turn over weekly is a working pipeline. One item sitting for a month is a stall. Count alone cannot tell those apart, and the age of the oldest item can.
  4. A no drains the queue exactly as well as a yes. This is the one I keep getting wrong. Rejection is service. Deferral is the only available response that leaves the item in the system, and it’s the one that feels least like a decision, which is why my queue has a depth of two instead of zero.

where I’d stop trusting the analogy

Two places, and the second one matters more.

Judgment isn’t a fixed service rate, and the items aren’t interchangeable. A short take I can read in a minute and a load-bearing architectural call both count as “one item,” and any model that treats them as the same unit will walk you straight into a bad conclusion. Backpressure in software assumes a uniform consumer. I’m not one.

The bigger limit: throttling the producer is only one of the two available fixes, and it’s the lazier one. The other is the CommercialTribe move, which is to make the review stage genuinely wider. Better specs up front so there’s less to adjudicate at the end. Smaller diffs. A written standard for what a yes actually requires, so the decision stops getting re-litigated from scratch every time. Volume being free is not an argument for restricting volume on principle. It means you now have to choose, deliberately, which of those two levers you’re pulling.

Throttle when the consumer is the irreducible part, the judgment that can’t be handed off to anything. Widen when the consumer is doing work that a system should have been doing all along. Mix those up and you get either a bottleneck dressed as discipline, or a rubber stamp dressed as a process.

back to the two drafts

I don’t think the mistake was generating too much. The mistake was building a producer with no idea whether anyone downstream was reading, then measuring the arrangement by whether it produced. The repair is small. The generator has to ask first.

So, your turn. What in your setup produces faster than you can judge, and where does the overflow go? If the answer is that it waits its turn, go find the oldest one and look at the date on it.

That’s the number I had been avoiding too.

Jason Waldrip is a fractional CTO and CAIO through The Bushido Collective, working with founders drowning in AI-generated code and teams scaling past the leadership that got them here. Flat retainers, measured on what ships and what gets prevented. If your generation side is running and you can’t say what your review capacity is, that’s a conversation worth having. Work with me.

Written with Claude Opus 5; the frame, the CommercialTribe and GigSmart calls, and the judgments about what matters are mine. The GenStage passage is quoted from its documentation and the work-in-process description from DORA’s capability catalog, both linked above. The two drafts are real, and this one went into the same queue behind them.