This post opens the AI systems series: ten Friday posts on the decisions an architect makes before a model goes into a product. It stands on its own, and it covers where a model sits, what it may change, and how the feature runs without it.
TL;DR
- A model call buys a gain in what the feature can do. It costs a new dependency and a new way to fail, and whoever is on call pays.
- Adding a model is three decisions: where it sits, what it may change, and how the feature runs when it is off.
- The test: can the team notice a bad answer, limit the worst case, and switch the model off without a release? If not, the model does not go in the path yet.
The roadmap says “add AI”
What you get: a picture of what a model in the path looks like, before any definitions.
Here is a scene. It is an illustration, not a case study. Nothing in it is a measured result.
A support agent opens a ticket on a help-desk screen. The roadmap has one line about this screen: “add AI to the help desk.” Afterward, the architecture diagram looks almost the same, with maybe one more box. The agent’s screen looks the same too.
That one box does not answer three questions. What does the screen now depend on? How can it now go wrong? How do we turn it off?
Plain versions help. A dependency is something the screen now waits on, or reads from, that it did not before. A new failure is a way to be wrong or slow that did not exist before. And “the path” is the chain of calls the user is waiting on.
The quick win is real, and it is not the whole cost. Sculley and coauthors argued in a 2015 NIPS paper that developing and deploying machine-learning systems is relatively fast and cheap, while maintaining them is difficult and expensive. They call it dangerous to treat the quick wins as free.
A caution on three sources: Sculley (2015), Breck (2017) and Zinkevich (undated) wrote about machine-learning systems that predict, before today’s model calls. The pattern carries over; the authors did not say so.
I use one sentence as the test for this series: “Put a model in the path only when the failure is one you can operate.” The last section pays it off.
Takeaway: adding a model is three decisions, not one, and a diagram like this one answers none of them.
Where the model sits
What you get: three places a model can sit, in plain words, and what each does to the user’s wait.
Here is how I sort the places a model can sit. I call them seats. No source I found uses this trio.
Request path. The model is called while the user waits. On the help-desk screen, a suggested reply appears as the agent opens the ticket.
Job. The app starts the model, and the result lands later. In the illustration, an overnight run tags tickets, so the tags are there in the morning.
Offline. The model ran before the request, and nothing in the running app starts it. The screen reads a stored result, such as ticket summaries written last week.
Google’s machine learning crash course draws a line between live predictions and ones already computed and stored. Live predictions are heavy on compute and sensitive to delay. You generally cannot check each one before a user sees it. Monitoring can still fire, but the course says that is after the problem has spread.
Stored predictions can be checked before they are pushed, but only stored ones can be served, and updates may take hours or days. That is the course’s wording for how stale a store can get, not a timeout.
Martin Zinkevich’s “Rules of Machine Learning” (Google, undated) names two placements: run the model live, or precompute results offline and store them in a table. Two, not three. So the split between a job and offline is my addition.

Each seat changes who is hurt on a bad day. On the request path, the user waits. Netflix’s engineers wrote that a single service in the active request path adds a failure mode. Their routing service had become a shared dependency whose failure would degrade or disable multiple experiences on their platform, so they moved it off the direct request path. That is one company’s platform.
Offline looks safe because nothing waits. Zinkevich (Rule #10) tells one anecdote, from one product: a table the system reads stops updating, and behavior decays gradually without looking like a hard failure. Tying that to this seat is my point, not his.
Cost and latency as a bill are covered in Cost and Latency of a Model Call on the User Path.
Takeaway: the seat decides who waits and what a bad answer touches. Offline looks safe because nothing waits, so check it on purpose.
What the model may change
What you get: a second decision, separate from the seat: how much the model may do with what it produces.
Back to the help-desk screen. The model has written a reply. What happens next? I use three words: suggest, draft, and write. These words are mine.
- Suggest. The reply appears as a proposal beside the ticket. The agent decides whether to use any of it.
- Draft. The reply arrives already in the reply box. The agent edits it, then sends it.
- Write. The model’s reply is sent or saved directly. No person comes between the model and the customer.
| Label | What the user sees | Who commits the change | What the off position leaves |
|---|---|---|---|
| Suggest | A proposed reply beside the ticket | The agent, by choosing or typing their own | The agent writes the reply, as before |
| Draft | A reply already in the reply box | The agent, after editing and sending | An empty reply box |
| Write | A reply already sent or saved | The model’s output, directly | Replies stop going out on their own and people handle the tickets |
Table: Three permissions for the same model output. The labels are the author’s, and the cells are qualitative.
Others make a similar cut with different words. Google’s People + AI Research (PAIR) guidebook (undated) names no automation, partial automation, and full automation.
In partial automation, the system recommends and the user chooses. In full automation, the system decides on the user’s behalf, and the user should still be able to take control back. PAIR says to start at the lowest level, and not to automate when the user cannot undo the result.
The model also needs a limit on how much it may change. This post uses an action limit as that bound (see the failure section).
This post only names the permission. Which actions must never go automatic is covered in Which Actions Never Go Automatic, and checks outside the model in Guardrails That Are Controls, Not Prompts.
Takeaway: the more the model may write, the more has to be in place before it goes live. Pick the lowest permission that does the job.
How the feature ships without the model
What you get: the third decision, and the plain rule it rests on.
The rule is mine: off returns the non-model result.
What does that look like in each seat? On the request path, the call is skipped or times out, and the screen shows what it showed before the model existed. In a job, the job stops, and the field stays empty or manual.
Offline is the odd one. The app keeps reading the last table, and the table ages. Off is not free there, because stale output looks fine.
Software teams already know this move from feature flags, and I am applying it to a model call. I did not find a guide that prescribes it for models, so treat the off rule as a design rule, not a standard.
Pete Hodgson’s 2017 article on feature toggles (martinfowler.com) says feature flags let a team change system behavior without changing code. In plain terms, a team can switch a behavior off without a release. His convention is that off means the existing, legacy behavior, and he says to test with the toggles off. That is general software, not machine learning.
On launching without the model, Zinkevich writes, “Don’t be afraid to launch a product without machine learning.” PAIR says, “Don’t use AI just because you can.” Both support a version that works with no model.
NIST’s AI Risk Management Framework (January 2023, voluntary) says mechanisms should be in place, with assigned responsibilities, to supersede, disengage, or deactivate AI systems that demonstrate performance or outcomes inconsistent with intended use. It states an outcome, not how to build a flag, a timeout, or a fallback.
Rolling back to an earlier model version, which Breck and coauthors (2017) say teams should practice, is a different move from removing the model. Off removes it.
One more note, and it is about load, not models. The Google SRE book chapter “Addressing Cascading Failures” says: “Remember that the code path you never use is the code path that (often) doesn’t work.” It says to exercise that path on purpose. I borrow one thing: an off position nobody has tried is a guess.
Takeaway: build the off position first, test it on purpose, and a model failure becomes a dull afternoon rather than an incident.
What “a failure you can operate” means
What you get: the series thesis as three things a person on call can do on a Tuesday afternoon.
The three verbs are detect, bound, and remove. They are the test I use in this series, not a framework from a source. Notice the answer is wrong before a customer does. Know the worst thing the model can do. Switch it off without shipping a release.
The three are not equally well sourced.
Detect. Sculley and coauthors say tests are valuable but not sufficient when the outside world changes. They say comprehensive live monitoring with automated response is critical. The stale table above is the same lesson: a bad input that does not look like a failure.
Building test cases is covered in Evals an Architect Can Run without a Research Team.
Bound. Bound is my word for Sculley’s action limits. For systems that take actions, they say it can be useful to set and enforce limits as a sanity check, broad enough not to fire spuriously. A hit sends an alert and a manual look.
Remove. This is the least-sourced of the three, and it is the one this post asks you to build first. The support is Hodgson (general software), the SRE chapter (general service engineering, about load), and NIST, which names the outcome but not a mechanism. None of them says how to build the off position for a model.
With all three, a team can say yes to the roadmap line without crossing its fingers. When a rule, a query, or nothing is the better design is covered in When Not to Add a Model.
Takeaway: if any of the three is missing, the model does not go in the path yet.
3 Questions for Your Architecture Team
- Which seat does this feature put the model in (request path, job, or offline), and how would we know the answer is wrong before a customer tells us?
A good answer names the seat on the diagram, names who sees a wrong answer first, and names a check that is not the user.
- What is the most the model is allowed to change (suggest, draft, or write), and what stops it at that limit?
A good answer names the permission in one word and names a limit that is enforced outside the model’s own output.
- If we turned the model off at 2 p.m. on a Tuesday, what would the user see, who could do it without a release, and when did we last try it?
A good answer shows the non-model result on screen and has a date for the last test.
When all three have an answer, the roadmap line is ready for the diagram, and the failure is one the team can operate.
Sources
- D. Sculley et al., “Hidden Technical Debt in Machine Learning Systems,” NIPS, 2015. https://proceedings.neurips.cc/paper_files/paper/2015/file/86df7dcfd896fcaf2674f757a2463eba-Paper.pdf
- Netflix Technology Blog (Nipun Kumar, Rajat Shah, Peter Chng), “State of Routing in Model Serving,” 2026. https://netflixtechblog.com/state-of-routing-in-model-serving-16e22fe18741
- Google, “Production ML systems: Static versus dynamic inference,” Machine Learning Crash Course, undated. https://developers.google.com/machine-learning/crash-course/production-ml-systems/static-vs-dynamic-inference
- Martin Zinkevich, “Rules of Machine Learning,” Google for Developers, undated. https://developers.google.com/machine-learning/guides/rules-of-ml
- Google People + AI Research, “Patterns,” undated. https://pair.withgoogle.com/guidebook-v2/patterns
- Pete Hodgson, “Feature Toggles (aka Feature Flags),” martinfowler.com, 2017. https://martinfowler.com/articles/feature-toggles.html
- NIST (Elham Tabassi), “Artificial Intelligence Risk Management Framework (AI RMF 1.0),” NIST AI 100-1, January 2023. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
- Eric Breck et al., “The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction,” IEEE Big Data, 2017. https://research.google.com/pubs/archive/aad9f93b86b7addfea4c419b9100c6cdd26cacea.pdf
- Mike Ulrich, “Addressing Cascading Failures,” Site Reliability Engineering (Google), undated. https://sre.google/sre-book/addressing-cascading-failures/

Leave a Reply