Your Product Runs on a Model You Don't Control. Plan for the Day It Changes

A startup ships an AI feature in the spring. The prompts are tuned, the demo is sharp, and the first customers are onboarded. Months later an email arrives from the provider saying the model ID in production will stop working on a fixed date. The engineer who wrote the prompts has moved to another project. Nobody knows which outputs customers actually depend on.
This is a scheduling and ownership problem more than a technical one. Anyone building on a hosted language model has a supplier whose roadmap they do not control, and that supplier will retire, reprice, or quietly modify the model your product runs on. This article covers how those changes reach you, what they cost when you are unprepared, and how much protection a startup should buy at each stage.
Three clocks that run outside your company
Model changes arrive on three separate schedules. Most teams watch only the first.
The retirement clock
Providers publish deprecation dates, but the notice periods vary a great deal. Anthropic says it gives at least 60 days of notice before retiring a publicly released model. OpenAI's stated minimum for generally available models is six months. Shorter windows exist. OpenAI notes that models with "preview" in the name can be retired on as little as two weeks. One inference platform documents an initial notice of 14 days before retirement. A developer who tracks deprecations across providers reported that xAI retired eight Grok API models in May 2026 with nine days of notice. That last figure comes from a community tracker rather than a provider's own documentation, so verify it before quoting it. The pattern still holds: your notice period depends on which provider and which tier you chose.
Retirements are also constant. On June 11, 2026, OpenAI told developers that older GPT-5 and o3 snapshots would be removed from the API on December 11, 2026. Retirement is routine housekeeping for providers, not an emergency.
The behavior clock
This one has no announcement. If your code calls a floating alias such as a "latest" model name, the provider can change what sits behind it. Pinning explicit model versions is the standard advice for reproducibility, because a floating alias changes your behavior silently. Support ticket tone shifts, a JSON field gets renamed, a classifier starts refusing borderline inputs. Nothing throws an error. You find out from a customer.
The economics clock
Prices, rate limits, and regional availability are set by someone else. I am not citing a specific change here because the point is structural: your gross margin on an AI feature depends on a price list you cannot negotiate at seed stage. A price cut helps you, and a price rise or a tighter rate limit hurts you, on a schedule you did not choose.
What being unprepared costs
Take a hypothetical seed-stage company with a support triage agent. The team wrote about forty prompts against one model over six months and never recorded what "good" looked like. A retirement notice arrives. Two engineers spend three weeks swapping the model ID, rewriting prompts that no longer behave the same, and manually spot-checking outputs. Nobody can say whether the new version is better or worse, only that it feels different.
The visible cost is engineering time. The larger cost is the roadmap slipping while two of your best people do migration work under a deadline. If a customer's workflow depended on a formatting quirk of the old model, you also find out about that in production.
How much portability to buy
Founders often overcorrect here. Full multi-provider parity, where every prompt works identically on three vendors, is expensive to build and to maintain. It also tends to force you onto the lowest common denominator of features, and most early-stage teams cannot afford that trade.
A better target is a migration budget: how long would it take us to move to a different model, and is that number acceptable? A useful way to set it is by stage.
Prototype or pre-revenue: Pin versions, keep model calls in one place, and accept that a migration will cost a week or two. Portability work beyond this is premature.
Paying customers, one core AI feature: Add an evaluation set and a tested second option. Aim for a migration you can finish in days, with evidence about quality.
AI is the product, with enterprise contracts or SLAs: Run a live fallback path, keep a deprecation calendar with a named owner, and treat provider concentration as a risk you disclose to customers and investors.
The goal is a migration that is boring, predictable, and cheap. Zero switching cost is neither achievable nor worth chasing.
Controls that make a migration boring
Pin every model version in production. Use dated snapshot IDs rather than floating aliases, so a behavior change only happens when you choose it.
Route all model calls through one module. Prompts, parameters, and provider specifics live behind a single interface, so a swap touches one place instead of forty.
Treat prompts and settings as versioned artifacts. Store them in the repo with review history, the same way you store code. This is also where good context design pays off, as covered in context engineering for AI agents, because well-structured context ports better than clever wording tuned to one model.
Keep an evaluation set that defines "good." Fifty to a hundred real cases with expected behavior is enough to start. It turns "this feels different" into a pass or fail result, and it is the gate every candidate model must clear.
Test the fallback before you need it. A second provider you have never run in staging is a hope, not a plan. Run your evaluation set against it every quarter and record the gap.
Put deprecation dates on a calendar with an owner. Subscribe to each provider's deprecation feed and assign one person to review it monthly. Ownership matters as much here as it does for security, a question explored in who owns AI security without a security team.
A quarterly migration drill
Once a quarter, pick your second-choice model and run the full evaluation set against it. Record the score gap, the prompts that broke, and the hours it took. After two or three drills you will know your real migration cost instead of guessing at it, and that number is what you tell a CTO, a board member, or an investor who asks about provider risk.
The drill also surfaces hidden coupling, such as parsing code that assumes one model's output format or a cost model built around one price list. Each one you find while nothing is on fire is a problem you fix on your own schedule.
The question to ask in your next planning review
If your primary model were retired in 60 days, could you name who does the work, what tells you the replacement is good enough, and what it costs? If any of those answers is unclear, that gap is your provider risk.
Providers will keep retiring models, and that is normal. The startups that handle it well treat model dependency as an operating risk with an owner, a measured migration cost, and a rehearsal schedule, rather than something they learn about when the notice arrives.




