Three companies operating at the intersection of AI and physical consequence — Shield AI, Waabi, and General Motors — shared a stage at TechCrunch Disrupt 2026 to discuss a problem that doesn't exist in chatbot benchmarks: how to verify that an autonomous system won't kill someone when the model hallucinates at 30,000 feet or 65 mph. The common thread across defense aviation, long-haul trucking, and factory robotics isn't model architecture. It's the infrastructure built around the model to bound its behavior, test its edges, and create auditable evidence that the system meets a safety case before it ever sees production.
The shift from accuracy to assurance
In supervised learning, you optimize for accuracy on a held-out test set. In mission-critical autonomy, you optimize for bounded failure modes under distribution shift. Nathan Michael, CTO of Shield AI, described Hivemind's development not as a model-training exercise but as a systems-engineering discipline where the neural network is one component among deterministic controllers, formal verification layers, and runtime assurance monitors. The U.S. Air Force's Collaborative Combat Aircraft program didn't select Hivemind because it had the highest mAP on a detection benchmark. They selected it because Shield AI could demonstrate, through a combination of simulation, hardware-in-the-loop, and flight test, that the autonomy stack maintains controlled flight and mission compliance across a defined operational envelope — including sensor degradation, communication loss, and adversarial electronic attack.
This distinction matters for any team moving agentic AI toward production. An agent that books the wrong restaurant reservation is annoying. An agent that commands a flight control surface incorrectly is a Class A mishap. The engineering response isn't "better prompts" or "larger models." It's a shift from validation by sampling to validation by construction: defining the safety properties that must hold, then building the test infrastructure to prove they hold across the state space you care about.
Simulation as a verification primitive, not a training accelerator
Waabi's Raquel Urtasun has been explicit that Waabi World isn't primarily a data generator for model training — it's a stress-testing environment designed to surface the long-tail scenarios that real-world miles will never adequately cover. The simulator encodes the physics of vehicle dynamics, sensor noise models calibrated from production hardware, and the stochastic behavior of other traffic actors. Crucially, it supports counterfactual replay: given a real-world disengagement or near-miss, engineers can re-simulate with perturbed initial conditions, sensor dropouts, or adversarial actor policies to map the boundary of the system's competence.
This approach inverts the typical ML workflow. Instead of asking "does the model generalize?", the team asks "where does the system violate its safety specification?" and builds regression suites around those boundaries. The partnership with Uber targeting 25,000+ robotaxis makes this concrete: each vehicle generates telemetry that feeds back into the simulator's scenario database, creating a closed loop where operational data continuously expands the verification coverage. For teams building agentic systems in other domains, the pattern is transferable: build a high-fidelity simulator of your action space, codify your safety invariants as checkable predicates, and treat simulation CI/CD as a gate on deployment — not a research tool.
Human-robot interaction as a safety requirement
Mikell Taylor's work at GM's Autonomous Robotics Center reframes a common blind spot: the human in the loop isn't a fallback — they're part of the system's operating envelope. Proteus, Amazon's first autonomous mobile robot, had to navigate fulfillment centers where humans don't follow traffic rules. The robot's behavior planner doesn't just avoid collisions; it models human intent, communicates its own intent through motion and lighting, and degrades gracefully when prediction uncertainty exceeds a threshold. This isn't UX polish. It's a safety argument: the system's correctness depends on the human's ability to understand and predict it.
For agentic AI deployed in shared spaces — warehouses, hospitals, construction sites — this means your agent's policy must be interpretable by design, not explained post-hoc. The action space should include explicit signaling behaviors (yielding, pausing, indicating intent) that are verified alongside navigation and manipulation tasks. The test matrix must include adversarial human behavior: distraction, non-compliance, deliberate probing. If your safety case assumes cooperative humans, it's not a safety case — it's a hope.
Building the evidence chain
Across all three domains, the deployment decision rests on an evidence chain that connects component-level verification to system-level safety claims. This chain typically includes:
- Formal methods on the controller layer: model checking or theorem proving that the low-level control laws satisfy stability and constraint satisfaction for all states in a verified region.
- Scenario-based testing in simulation: millions of parameterized scenarios covering nominal, edge, and adversarial conditions, with automated pass/fail criteria tied to safety requirements.
- Hardware-in-the-loop (HIL) campaigns: the exact compute stack, sensors, and actuators running closed-loop in a lab with injected faults (bit flips, latency spikes, sensor drops).
- Structured field testing: operational test plans with defined coverage metrics (miles per operational design domain, encounters per scenario type) and independent safety oversight.
- Runtime assurance: a verified monitor that watches the learning-enabled component and can override or safe-hold when behavior leaves a certified envelope.
None of these layers is sufficient alone. The argument for deployment is the composition of evidence across layers, with traceability from top-level safety requirements down to individual test cases. This is where most agentic AI projects stall: they have model cards but no safety case; they have eval benchmarks but no traceability to operational hazards.
Regulatory reality as a forcing function
The panel emphasized that regulatory engagement isn't a post-development checkbox — it shapes architecture from day one. Shield AI works within military airworthiness processes (MIL-STD-882, DO-178C analogs for autonomy). Waabi engages with NHTSA, FMCSA, and state-level AV frameworks. GM navigates OSHA, ANSI/RIA R15.06, and emerging standards for industrial mobile robots. In each case, the regulatory pathway dictates what evidence must exist, what independence is required in the assessment, and what operational constraints will be imposed.
Teams building agentic systems for regulated domains should treat the target standard as a requirements document, not a compliance hurdle. If DO-178C requires modified condition/decision coverage (MC/DC) for Level A software, your neural network verification strategy must produce equivalent evidence — or you must architect the system so the learning-enabled component is never Level A. That architectural decision (e.g., a runtime assurance monitor that is Level A and constrains a Level D neural network) is a design choice, not a certification afterthought.
Culture: the hidden infrastructure
Michael, Urtasun, and Taylor each described organizational practices that sound mundane but are load-bearing: blameless postmortems for every disengagement, weekly safety reviews with authority to stop deployments, dedicated safety engineering teams separate from model development, and explicit "stop work" criteria tied to leading indicators (simulator regression failures, HIL fault injection results, field anomaly rates).
This cultural infrastructure is what prevents the normalization of deviance — the gradual acceptance of "it worked last time" as evidence of safety. In agentic AI, where the system's behavior emerges from learning rather than specification, this discipline is the difference between a demo and a product. The most effective lever is often organizational: give the safety team veto power, fund the simulation infrastructure as a first-class product, and make the evidence chain auditable by external assessors.
What this means for your agentic system
If you're building an LLM-based agent for customer support, the stakes are lower — but the engineering pattern holds. Define your safety invariants (no PII leakage, no unauthorized actions, bounded cost per session). Build a simulator of your tool-use environment with adversarial user policies. Create a runtime monitor that validates each action against policy before execution. Instrument every deployment with structured telemetry that feeds back into your scenario database. Treat the agent as a component in a larger assured system, not the system itself.
The companies on that Disrupt stage didn't solve alignment with better prompts. They built assurance stacks around learning-enabled components — stacks that include deterministic fallbacks, formal verification, high-fidelity simulation, structured field testing, and organizational accountability. That's the blueprint for any agentic AI that has to work when "almost right" isn't good enough.
Frequently Asked Questions
How do you verify a neural network that can't be formally verified?
You don't verify the network in isolation. You wrap it in a runtime assurance architecture: a formally verified monitor that checks the network's output against a safe envelope before the action executes. If the output violates the envelope, the monitor overrides with a deterministic safe action (hold, abort, handover). The verification effort shifts to the monitor and the envelope definition — both of which are tractable for formal methods — while the network operates as a "best effort" optimizer within those bounds.
What's the minimum simulation fidelity needed for meaningful safety evidence?
Fidelity must match your safety-critical failure modes. If your agent fails when sensor latency spikes, your simulator needs accurate latency models. If it fails on adversarial human behavior, you need a human behavior model calibrated from real data. Start with the hazard analysis: identify the top 10 ways the system could cause harm, then build simulation fidelity specifically around those scenarios. Don't chase photorealism — chase causal fidelity for your hazard scenarios.
How do you handle the combinatorial explosion of agentic action spaces?
You don't test the full combinatorial space. You use property-based testing and scenario reduction: define safety properties that must hold regardless of action sequence (e.g., "never command force above threshold," "always maintain communication heartbeat"), then use formal methods or statistical model checking to verify those properties over the abstracted state space. For the concrete scenarios, prioritize coverage of your operational design domain's critical regions using importance sampling guided by your hazard analysis.
Related reading: Runtime Assurance Patterns for LLM Agents | Simulation CI/CD for Autonomous Systems | Building Safety Cases for Production AI


