7 min read

Pre Deployment Simulation Is the SRE Run Book for AI Agents

Hand-drawn paper-craft path crossing checkpoints and a bridge to rehearse safe AI-agent deployment.

The alert fired at 2:47 AM. A production AI agent, deployed six weeks prior to handle inventory routing for a mid size logistics company, had started rerouting shipments through a deprecated warehouse. Not because it was instructed to. Because a downstream data pipeline updated its schema overnight and the agent's retrieval logic, never stress tested against stale schema references, silently defaulted to a fallback that happened to point at the decommissioned facility. Forty seven orders were misrouted before the on call engineer caught the anomaly in the morning's reconciliation report.

The TPM leading that program had a monitoring dashboard that showed token throughput, latency percentiles, and error rates. It did not show that the agent was making decisions against corrupted state. Post deployment monitoring had done its job: it surfaced the symptoms. What it could not do was prevent the incident, because the conditions that caused it had never been simulated before go live.

This is the moment we are collectively arriving at, and it is uncomfortable.

The Gap: Why Post Deployment Monitoring Fails Agentic AI

Traditional software monitoring was designed for a world where systems process inputs and return outputs. The loop is closed. An API responds with data or an error; a service writes a record or fails. Monitoring that world means watching for anomalous outputs, raised error rates, or degraded latency. It works, and we have built sophisticated tooling around it.

Agentic AI breaks that model in a structural way. An agent does not merely process and return. It acts. It calls tools, writes to databases, triggers downstream workflows, amends its own instructions based on retrieved context, and propagates those changes across subsequent calls. The output of an agent is frequently not text but behavior: a sequence of actions whose consequences unfold over minutes or hours across multiple systems.

This changes the risk profile entirely. When an LLM returns a hallucinated answer, the damage is a bad response on screen. When an agent acts on a hallucinated context, the damage is a rerouted shipment, an incorrect database write, a policy decision made against wrong data. The monitoring we built for the former is structurally inadequate for the latter.

The gap is temporal. Monitoring tells you something went wrong after it has already gone wrong. For agents that take irreversible actions, that lag is not a performance problem; it is a categorical insufficiency. We need to know whether the agent will take the right action in a given scenario before it encounters that scenario in production. That is not a monitoring problem. That is a simulation problem.

What Pre Deployment Simulation Actually Is

Pre deployment simulation is the practice of running an AI agent through a library of realistic operational scenarios in a staged environment that mirrors production behavior, before the agent is live. It is not unit testing. It is not a benchmark evaluation. It is end to end rehearsal.

Static LLM evaluation asks: how does this model perform on this prompt? The output is a score, a preference ranking, a classification accuracy. That is useful for model selection, not for deployment confidence. A model can score in the 94th percentile on MMLU and still route forty seven shipments through the wrong warehouse because it encountered a schema edge case that no benchmark anticipated.

Pre deployment simulation asks a different question: what will this agent do when it hits the real world? It involves constructing scenario libraries that represent the actual distribution of inputs the agent will face in production, including adversarial cases, edge cases, schema variations, upstream failures, and permission boundary conditions. The agent is run through these scenarios in an environment that emulates production dependencies, and every consequential action is logged, evaluated, and traced to its downstream effects.

The SRE analogy is instructive. Before we promote a service to production, we do load testing. We simulate traffic patterns, inject network faults, trigger downstream dependency failures, and verify that the system degrades gracefully. We do not ship a service and wait for the on call alert to tell us it fell over under real load. That would be considered negligent. Yet that is exactly what we have been doing with AI agents: deploying them into production and waiting for monitoring to surface problems that we could have found in rehearsal.

The run book exists because we learned, over decades of production incidents, that documented response procedures reduce mean time to recovery. Pre deployment simulation is the run book for agentic AI: it is the documented procedure for verifying, before go live, that the agent behaves correctly under the scenarios it will actually encounter.

The OpenAI Signal

The market validated this practice on June 16, 2026, when OpenAI formally extended its deployment simulation capabilities to cover agentic coding workflows. This was not a research paper or a proof of concept. It was a production grade feature, documented and shipped, that allowed developers to simulate an agent's tool use sequences against a defined scenario library before deployment.

The significance is not the feature itself. It is what it represents. OpenAI does not ship experimental governance features. When OpenAI formalizes a capability into its deployment workflow, it is a leading indicator that the practice has crossed the threshold from emerging to table stakes. The market leader has determined that simulation is no longer optional for agentic deployments; it is a cost of doing business.

This follows a pattern we have seen before. observability was optional in early microservice architectures. It became mandatory as systems scaled and the cost of production incidents grew. Contract testing was optional when services were stable. It became mandatory as services changed independently and integration failures multiplied. Pre deployment simulation for AI agents is following the same trajectory, and the OpenAI announcement is the inflection point where the trajectory becomes undeniable.

For practitioners, the implication is straightforward: simulation is not a nice to have any more. It is what you specify when you commission an agentic AI deployment, and it is what you verify before you sign off.

The Lyft Case Study

Lyft's Metric Semantic Layer offers the most concrete example of how leading organizations are operationalizing pre deployment governance for AI agents that interact with data. The engineering team built a semantic layer that serves as a governed interface between AI agents and the organization's metric definitions, making sure that every agent that queries Lyft's data does so through a standardized, versioned, access controlled API rather than raw database calls.

This is pre deployment governance at the architectural level. Before an agent can access a metric, that metric must be defined in the semantic layer with documented logic, owner, and access control policy. The agent does not bypass this layer in production because the infrastructure is structured to make raw access impossible. The governance happens at design time, when the metric schema is defined and the agent's data access permissions are scoped, not at monitoring time when the agent has already made an unauthorized call.

What makes this directly relevant to TPMs is the commissioning question. When you commission an AI agent that will interact with organizational data, you are making a decision about what data it can access, what actions it can take based on that data, and who is accountable when it acts incorrectly. The Lyft pattern makes that decision explicit, auditable, and enforceable before the agent is deployed. It shifts governance left: from runtime enforcement to pre deployment specification.

The lesson is not to copy Lyft's semantic layer architecture. It is to recognize that the governance structure for agentic AI access needs to be specified, simulated, and enforced before deployment. Simulation in this context means verifying that the agent's data access behavior conforms to the access control policy under realistic query conditions, not just that it produces correct outputs in isolation.

For TPMs, this means requiring the same rigor of your agentic AI deployments that you would require of any system that touches sensitive data: documented access control policies, simulated verification of boundary conditions, and explicit sign off that the agent's behavior has been tested against the scenarios it will actually encounter.

The TPM Commissioning Checklist

Before signing off on an agentic AI deployment, a TPM should require the following:

  1. A documented scenario library that covers the agent's operational distribution, including edge cases, upstream failure modes, and adversarial inputs. This is not a list of happy path prompts. It is a living corpus of the scenarios the agent will face in production, curated and maintained by the team responsible for the domain.
  2. Simulated execution against the full scenario library in a production mirrored environment, with full action tracing. The agent's entire sequence of tool calls, data accesses, and downstream actions must be logged and auditable. PASS/FAIL criteria for each scenario must be defined before simulation begins, not retroactively.
  3. Data access governance documentation specifying exactly what data the agent can access, through what interface, and under what conditions. If the agent interacts with organizational data, the access control model must be reviewed and signed off before simulation begins.
  4. Rollback criteria and automated circuit breakers defined before go live. When the agent exhibits behavior outside the defined acceptable range in production, what happens? The answer must be documented, tested, and ready to execute before the agent is live.
  5. An accountable human owner for the agent's behavior in production, with defined escalation paths and response SLAs. The agent is a system. Someone owns it.
  6. A post deployment monitoring specification that covers agent specific metrics: action sequences, tool call patterns, downstream propagation effects, and context drift over time. Standard latency and error rate monitoring is necessary but not sufficient.

The Institutional Argument

The case for pre deployment simulation extends beyond any single deployment. Organizations that build simulation capability for agentic AI are building institutional knowledge that compounds across programs. The scenario libraries, simulation frameworks, and governance patterns developed for one agent deployment become reusable assets for the next. The TPM who specifies rigorous pre deployment simulation for their program is not just reducing risk for that program; they are contributing to the organization's ability to deploy agentic AI safely at scale.

There is also a continuity argument. Agentic AI systems are not static. They change as the models that power them are updated, as the data they depend on evolves, and as the tasks they perform are extended. Pre deployment simulation provides the baseline against which those changes are evaluated. Without it, you cannot know whether a model update has altered the agent's behavior in consequential ways. With it, you have a reproducible, auditable test use that makes model rotation a governed event rather than an unmonitored deployment.

Risk management at the program level benefits similarly. When the board asks how you know the agent is safe to deploy, a documented simulation run against a defined scenario library, with traceable results and explicit sign off, is a materially better answer than pointing to a latency dashboard.

Closing

The logistics company whose agent rerouted those forty seven shipments did everything right by the standards that existed six months ago. They had monitoring. They had on call coverage. They had incident response procedures. By the standards that are emerging now, they were underprepared, because no one had asked the agent to run through its edge cases before it hit production.

That standard is changing. OpenAI's adoption of deployment simulation, Lyft's architectural approach to agent data access, and the growing body of production incidents driven by agent action sequences are collectively establishing a new baseline for what responsible deployment looks like. The TPMs who commission and specify simulation as part of the standard deployment contract will reduce program risk, build institutional capability, and set the bar for the rest of the industry.

The run book exists. Now we have to use it.