You Pinned a Countdown
Your lockfile will hold every package in your build at a fixed version for as long as you want. The one component in your system that makes decisions comes with a published date after which it stops answering, and the version was never the part you needed held still.
08:20 8 min read
The agent that answers text messages for my Starlink rental company has been doing it since July. Nothing about it has changed in weeks. No commits, no config edits, no deploys. It works, which in that business means somebody standing in a parking lot asks whether a dish can reach them by Friday and gets a straight answer.
Its model has a retirement date. I didn’t pick the date. I can read it on a documentation page, and when it arrives, requests to that model fail.
Nobody is behaving badly in that sentence. It’s just that I’ve taken on a dependency which doesn’t work like any dependency I’ve taken in twenty years, and I spent months applying the wrong habit to it without noticing.
the responsible thing
You pinned it. Of course you did. You’ve been pinning versions since the first time a patch release ate an afternoon, and by now it’s reflex: put the exact string in the config, commit the lockfile, let the upgrade be a decision somebody makes on purpose rather than something that happens to you on a Tuesday.
So the model id in your repo says claude-sonnet-4-5-20250929 or whatever the equivalent is on your provider, and it does not say latest, and you feel fine about that. You’ve written about the person behind every other line in that manifest, or at least you’ve read about it. This one looks like the same discipline applied to a newer kind of thing.
It isn’t. A version on npm is immutable and permanent by policy, and the one famous time somebody pulled eleven lines back out, it was a scandal that ended with the registry force-restoring the package. Permanence is the promise. The whole reason a lockfile means anything is that the thing it points at will still be there.
Anthropic publishes the other arrangement openly. Models move through four states: active, then legacy, then deprecated with a retirement date attached, then retired. The page is blunt about the last one. Requests to retired models will fail. Claude Sonnet 4 and Opus 4 retired on June 15, 2026. Opus 4.1 followed on August 5. Customers with active deployments get at least sixty days of notice, which is real notice and more than most things that break your week will give you.
It goes further than the model, too. On that same page, temperature, top_p, and top_k are deprecated on Opus 4.7 and later, and setting them to a non-default value returns a 400. Those are knobs. Some of you tuned production behavior with those knobs and wrote a comment above the number explaining why it was 0.2.
None of this is a complaint. Retiring models frees capacity for better ones, and the notice period is published, and the replacement is named for you in a table. I’d rather have a vendor with a documented lifecycle than one that quietly swaps things underneath me. But a documented lifecycle is still a lifecycle, and it means the thing you pinned is not a version in the sense you’ve used that word your whole career. It’s a reservation.
The consumer side moves faster than that. When GPT-5 shipped in August 2025, GPT-4o disappeared from ChatGPT the same day, with no deprecation window at all, and people who’d built daily working habits on it found their conversations converted underneath them. Simon Willison wrote it up while it was happening. The backlash was loud enough that OpenAI restored 4o for paid users within days and promised warning next time. The API kept serving it the whole time, which is the part worth holding onto: the interface you’re on determines what you’re owed.
the pin never held the thing you cared about
Here’s the part that took me longer to admit, and it’s the one that actually matters.
Even when the version string holds perfectly still, you haven’t frozen anything you’d recognize as behavior.
In 2023, Lingjiao Chen, Matei Zaharia, and James Zou ran the March and June snapshots of GPT-4 against the same tasks and published what they found. On identifying prime versus composite numbers, the March version scored 84 percent. The June version scored 51. Same product, same name, three months apart.
That paper got argued with, hard, and some of the argument was fair: the prompt format mattered, the task was narrow, part of the movement was formatting rather than reasoning. The more interesting question is why the argument was possible at all. A room full of capable people could not settle, from the outside, whether a system had gotten worse at something. There was nothing to diff. That is the position you are in every day with the component at the center of your product.
So ask your pipeline what it actually checks about that agent’s reply. It checks that the JSON parsed. It checks that the tool call arrived with arguments of the right shape. Maybe a snapshot test asserts the string, which means it breaks on every harmless rewording and tells you nothing on the day the reasoning slips. The compiler I trust to review everything the fleet writes cannot see inside this at all. Your type system stops precisely where your judgment layer starts.
Green tests and working software were always different claims. This is the widest that gap gets, because here the thing under test doesn’t fail. It answers, confidently, in well-formed JSON, slightly worse than it did last quarter, to a customer in a parking lot.
pin the behavior instead
You can’t pin the model. You can pin what you’re willing to accept from it, and that turns out to be the artifact that solves both halves of this.
There’s a voice agent on my rental line, and she has a kill switch. I wrote out what she says when it’s off: Sorry, I can’t help with that over the phone right now. Text this number and I’ll help you there. At the time I thought of that as an outage plan. It’s the same instinct pointed at a slower failure. I decided in advance what unacceptable looks like, and I built the thing that notices, and I wrote down where the traffic goes when it happens.
Scale that up and it has a name nobody finds exciting: a set of cases, with real inputs from your own transcripts, and an expected shape of answer, and a score you can run on demand. Twenty cases beats zero by more than you’d think, and the first ten come straight out of the conversations you already have.
Notice what that one artifact buys you. It’s the detector for silent drift, and it’s also your migration plan. The vendor’s own advice is to test your application against the replacement well before the retirement date, which is excellent advice that costs nothing if you have a suite and is close to impossible if you don’t. Without one, every forced migration is a vibe check performed under a deadline you didn’t set, on the flow that talks to your customers.
- Write down what working means, for one flow. The one that touches a customer, not all of them. Real transcripts, expected shape of answer, your own judgment about which replies are fine. This is the one part that can’t be delegated, because it’s the definition of good, and defining good is the job.
- Make the model id a variable, and run two. If swapping models means a code change and a deploy, your migration is a project. If it’s a config value with a scoreboard behind it, it’s an afternoon.
- Keep the kill switch, and write what it says. Every autonomous surface in front of a customer needs an off position and a sentence to say in it. Decide that sentence on a calm day.
- Put the retirement date on the calendar. Treat it like a TLS expiry, because it’s the same category of problem: a clock you don’t control and can’t compress. Sixty days of notice only helps somebody who reads the notice.
The habit worth breaking is thinking of the model as infrastructure that sits still. Every other dependency in your build is a thing you installed. This one is a service you’re renting, with a term, that changes character inside the term, and the rental agreement is a docs page you have never opened.
So go look at the agent that’s been quietly working since July. Not the code around it, which is fine. The line that names the model. Do you know what date is next to it, and if it answered a customer worse this morning than it did in July, what in your entire system would have told you?
Jason Waldrip is a fractional CTO and CAIO through The Bushido Collective, working with founders drowning in AI-generated code and teams scaling past the leadership that got them here. If there’s an agent in front of your customers and nobody can say what it’s scored against, that’s a good place to start a conversation. Work with me.
Written with AI assistance; the frame and the calls on what matters are mine. WiFi Without Walls is a company I own, and the agents and the kill switch described here are its own. The lifecycle states, the retirement dates, and the parameter deprecation come from Anthropic’s public model deprecations page; the drift figures come from Chen, Zaharia and Zou’s paper rather than a retelling of it.