Why This Conversation Matters Now
Agentic AI is moving from chat interfaces into aircraft, trucks, and factory floors. The error modes change fundamentally: a language model hallucinates text, but an autonomous stack controlling physical systems can ground a fleet, crash a vehicle, or injure a worker. Three domains — defense aviation, long-haul trucking, industrial robotics — face the same readiness question with different constraints. Founders building vertical agentic applications need cross-domain patterns, not just autonomous vehicle best practices.
The TechCrunch Disrupt 2026 session "Building AI Systems When Failure Is Not an Option" brings together Shield AI CTO Nathan Michael, Waabi founder Raquel Urtasun, and GM Director of Robotics Strategy Mikell Taylor. The event runs October 13-15 at Moscone West. This preview extracts transferable principles from each speaker's public track record, since the session itself hasn't happened yet.
Shield AI: Platform-Agnostic Autonomy in Contested Environments
Nathan Michael leads Hivemind, Shield AI's platform-agnostic mission autonomy software. His background spans AI, control, perception, and multi-robot systems at Carnegie Mellon's Robotics Institute, where he directed the Resilient Intelligent Systems Lab. That academic rigor meeting production timelines is the first transferable lesson: assurance-driven development doesn't happen in a research vacuum.
Hivemind's selection for the USAF Collaborative Combat Aircraft program implies specific verification gates. The program demands autonomy that operates in GPS-denied, communication-contested environments — conditions where the agent must plan, adapt, and coordinate without reliable external input. This requires a stack architecture where the autonomy layer is decoupled from the airframe, sensors, and communications hardware. Platform-agnostic design isn't a luxury; it's a verification necessity. When you certify the autonomy logic separately from the platform integration, you can reuse assurance evidence across aircraft types.
The $1.5B Series G at $12.7B post-money valuation (February 2026) reflects investor confidence in this approach. For teams building agentic AI systems development stacks, the takeaway: design your autonomy core as a portable, verifiable module from day one. Contested environments expose the difference between "works in simulation" and "provably safe under bounded uncertainty."
Waabi: Simulation-First Validation at Scale
Raquel Urtasun spent 25 years in AI and autonomous vehicles, including chief scientist at Uber ATG, before founding Waabi. The company's $1B raise (January 2026) and Uber partnership targeting 25,000+ robotaxis create a forcing function: production-grade safety cases at fleet scale. Waabi World, their simulator, reframes testing from "miles driven" to "scenario coverage completeness."
The key architectural decision: Waabi World is differentiable and reusable as a validation layer, not just a training environment. Traditional AV stacks treat simulation as a pre-deployment checkpoint. Waabi treats it as a continuous, differentiable component of the development loop — scenario generation, policy evaluation, and regression testing all feed back into the same world model. This means when a corner case appears in the real world, it can be instantiated, varied, and stress-tested in simulation before the next software release.
Urtasun has stated Waabi's autonomous trucks still need full validation before driverless deployment. That transparency about the validation gap is itself a lesson. Most teams over-claim readiness; Waabi's public stance treats validation as an explicit, measurable milestone. For your stack, ask: can you generate a coverage report showing which operational design domain scenarios are validated, which are tested only in simulation, and which remain untested? If not, your validation infrastructure is the bottleneck.
General Motors: Human-Robot Teaming as a Design Constraint
Mikell Taylor leads robotics strategy for GM's Autonomous Robotics Center. She previously led the Amazon Robotics team that developed Proteus, Amazon's first autonomous mobile robot designed to work alongside people in fulfillment centers. That experience — building robots that operate in human spaces from day one — shapes a different readiness criterion: adoption and trust as engineering requirements, not afterthoughts.
Proteus had to navigate dynamic aisles, yield to human workers, communicate intent through motion and lighting, and degrade gracefully when sensors failed. These aren't UX polish; they're safety layers. A robot that behaves unpredictably forces humans into defensive behaviors that create new failure modes. Legible intent — the ability for a nearby worker to predict what the robot will do next — reduces cognitive load and collision risk.
GM's mandate emphasizes practical reliability over demo impressiveness. Factory floors don't need robots that dance; they need robots that show up every shift, handle the edge cases that broke last month's run, and fail in ways the line supervisor can diagnose in minutes. Taylor's perspective: the human-robot interface is part of the safety case. If operators don't trust the system, they'll work around it, creating shadow workflows that invalidate your assumptions.
Cross-Domain Patterns: What Transfers to Your Stack
Three domains, one validation pyramid. The pattern that emerges across Shield AI, Waabi, and GM:
- Safety culture: Independent assurance teams separate from development. Incident taxonomy with severity classification. Blameless postmortems that feed back into scenario libraries.
- Validation pyramid: Unit tests → simulation regression → shadow mode (agent runs parallel to human, no control authority) → supervised autonomy (human on loop) → unsupervised (human on call). Each layer requires explicit exit criteria.
- Regulatory strategy: Evidence packages assembled continuously, not at submission. Regulator as stakeholder from early design reviews. Continuous certification frameworks where software updates trigger targeted re-verification, not full recertification.
- Trust metrics: Mean time to recovery after autonomy disengagement. Scenario coverage percentage against operational design domain. Human takeover rate per 1,000 hours. Predictability scores from human factors studies.
The validation pyramid is where many teams stall. They invest heavily in simulation but lack shadow mode infrastructure. Or they run shadow mode but don't instrument the comparison logs for systematic discrepancy analysis. RAG pipelines and AI automation workflows can feed simulation scenario generation from real-world incident databases, regulatory guidance documents, and domain expert knowledge bases — turning unstructured lessons into structured test cases.
From Disrupt Stage to Your Roadmap
Before the session airs, run a self-assessment on your own stack:
- Bottleneck identification: Which validation layer is your current constraint? Unit test coverage? Simulation scenario diversity? Shadow mode logging fidelity? Supervised deployment telemetry?
- Team composition: Do you have independent safety/assurance capacity? Not a QA engineer — someone whose career incentive is finding flaws, not shipping features.
- Partner ecosystem: Simulation infrastructure that supports differentiable scenario variation. Regulatory advisors who understand your domain's evidence standards. Pilot customers willing to run supervised deployments with instrumented feedback loops.
Armor Tech partners with teams building agentic systems that need to earn trust in production. We design validation frameworks, simulation infrastructure, and safety-aware architectures that close the gap between lab demos and deployable systems. The patterns above aren't theoretical — they're what separates systems that stay in pilot from systems that ship.
Frequently Asked Questions
How does agentic AI validation differ from traditional ML model validation?
Traditional ML validation measures static accuracy on held-out datasets. Agentic AI validation must measure dynamic behavior across state spaces: does the agent recover from perception failures? Does it maintain safety invariants under distribution shift? Does it degrade gracefully when sensors fail? The validation target is a policy, not a prediction.
What's the minimum viable validation infrastructure for a seed-stage team?
Start with a scenario definition language that captures your operational design domain, a simulation environment that can replay and vary those scenarios deterministically, and a logging pipeline that records every autonomy decision with sufficient context for post-hoc analysis. Shadow mode against human operators comes next. Regulatory evidence packages come when you have a pilot customer.
How do you handle the simulation-to-reality gap without massive real-world fleets?
Domain randomization in simulation (varying lighting, sensor noise, physics parameters) combined with targeted real-world data collection on edge cases identified in simulation. The loop: simulation finds failure modes → targeted real-world collection validates or refutes them → updated simulation parameters. Waabi's differentiable simulator accelerates this by making the gradient from scenario parameters to policy performance computable.


