The Agent OS Race Is On — TPMs Need a Platform Bake Off, Not Another Agent POC

The agent runtime is becoming a platform commitment, not a collection of API calls. In a 48 hour window, Cloudflare and AWS made distinct platform plays, while Red Hat made the operational layer explicit and the stateless MCP discussion sharpened the substrate question. Together, these signals suggest that the category is moving above the model and below the business workflow. For TPMs, the question has changed from "which agent harness should we trial?" to "which operating layer can we live with for the next year?"
The number that matters
The number is 48: three adjacent platform signals landed within roughly two days while the protocol conversation moved toward a simpler stateless model, an early sign that the buying category is forming faster than most roadmaps are prepared to evaluate, not proof that one vendor has won.
The risk is treating this as another short proof of concept. A proof of concept asks whether an agent can complete a narrow task, while a platform commitment asks whether a company can operate context, skills, identity, tools, observability, incident response, and migration without rebuilding the program every quarter.
The framework
"The procurement question is no longer which agent harness do we buy? It is whose agent OS do we commit to?"
That is the TPM framing. The vendor names matter, but the buying decision should be made against the operating layer each vendor is trying to own.
- Separate the substrate from the brand.
Simon Willison's account of Stateless MCP is useful because it describes the protocol change from an implementer's perspective. The stateless direction reduces the ceremony required to build MCP clients and servers. That matters to a TPM because protocol friction compounds across every tool integration, test environment, and security review.
Do not score a platform by the number of MCP servers it shows in a demo. First ask which MCP version it supports, how it handles legacy stateful integrations, and whether the tool boundary remains inspectable when the server delegates to another runtime. The substrate should reduce switching cost, not become a new form of lock in.
- Score the operating layer, not the feature list.
Cloudflare OS makes the most explicit operating system claim. Its description centers on a shared library of company context and skills that people across functions can use. That is a governance and reuse bet: when one person finds a better way to perform a task, the pattern can become available to others.
AWS Bedrock AgentCore makes the clearest runtime play in the brief. Its coordinated material covers a harness for n8n, persistent memory, tool access, code execution, and VPC isolation. The MCP bridge example treats local tools as a platform integration problem rather than an isolated developer trick.
Red Hat's production observability position takes the operations wedge. It treats telemetry as a first class layer for agents, which is exactly where a platform becomes accountable after the demo ends.
These are different entry points into the same buying conversation. Your rubric should ask what each platform makes routine: shared context, reusable skills, controlled tool access, production telemetry, incident investigation, rollback, and migration.
- Use a vendor neutral production vocabulary.
The Data For Science harness model offers a practical vocabulary for comparing products without repeating vendor branding. It describes mission planners, parallel execution, fuel budgets, flight recorders, and after action reviews.
Translate those ideas into your program rubric. Mission planning maps to task decomposition and ownership. Parallel execution maps to concurrency controls. Fuel budgets map to token and tool spend. Flight recorders map to traces and audit evidence. After action reviews map to evaluation and incident learning.
If a vendor cannot explain how its product supports each of those operating needs, treat it as an agent SDK for this decision, even if the demo is polished.
- Pre stage the incident primitive.
The HyperProbe launch is a useful leading signal because it frames read only production debugging as an agent capability. Its pitch is not "give the agent a shell and hope." It is a probe that can inspect the line where a problem occurred while preserving a read only boundary.
That is the kind of primitive a platform bake off should test before a vendor makes it part of the marketing surface. Ask whether the agent can inspect a production failure without receiving broad write access. Ask what evidence it records. Ask how a human can reproduce the investigation and challenge the conclusion.
The future operating layer will need more than an execution environment. It will need safe ways to understand failures in the system it operates.
- Make the commitment reversible.
A platform bake off is incomplete if it measures only first run success. For each candidate, document data egress, context portability, skill portability, MCP compatibility, observability migration, token economics, and the procedure for moving one critical workload away.
Run the same workload against three candidates for 90 days. Use a single evaluation set built from representative tasks and known failures. Record quality, latency, tool reliability, incident time, and the human effort required to correct a bad run. Name the rollback trigger before the first result arrives.
Use the Mobileye support deployment and the LendingTree mortgage deployment as scenario shapes, not as proof that AgentCore wins. Recreate the underlying evaluation questions with your own workload: what evidence is captured, who owns a failure, what data crosses the boundary, and how much human correction is required.
The goal is not to predict the winner. It is to make the cost of being wrong visible before the commitment becomes structural.
Sources for the platform cut and counter evidence
The platform signals are documented in the Cloudflare OS announcement, the AgentCore harness release, the MCP bridge, the Mobileye deployment, and the LendingTree deployment. The protocol and operations context comes from Stateless MCP, Red Hat's observability article, HyperProbe, and the advanced harness model. These are directional product and practitioner signals, not a controlled comparison of platform performance.
What this does not solve
Vendor selection is not architecture. Choosing Cloudflare, AWS, Red Hat, or a combination does not define your data boundaries, service ownership, escalation rules, or model change process. Those still belong in your program design.
A shared substrate does not create adoption. A context library and a skill library can make good practice reusable, but they can also spread a bad workflow faster. Assign owners, review changes, and retire patterns that no longer match the system.
Observability does not equal control. A trace can explain what an agent did after the fact. It does not automatically prevent an unsafe action, establish a correct authorization boundary, or make rollback cheap. Treat telemetry as evidence inside a control loop, not as the control loop itself.
The signal that matters most
The strongest signal is not that three vendors used platform language in the same week. It is that the buying horizon is lengthening. A TPM who evaluates only whether an agent can complete a task will miss the cost of operating the task across context, tools, identity, telemetry, and failure recovery.
Start the bake off with the workload that would be most painful to migrate later. Write its quality floor, latency ceiling, evidence requirement, and rollback condition. Then run it across the candidates before a vendor's product story becomes your architecture.
Send me your agent platform bake off rubric — five categories and the one workload you would use as the migration test. DM me on LinkedIn (Doron Katz). I am collecting working patterns into a public agent operating layer playbook; three examples would let me ship it next month.
Member discussion