Runtime Assurance for Autonomous Systems: Beyond Vehicle Health
Runtime assurance for autonomous systems means checking each course of action against mission contracts, with auditable evidence for DoD test and evaluation.
A fleet of autonomous platforms can fail its mission while every vehicle in it reports green. Batteries are within limits, links are up, and navigation is nominal, yet the plan the fleet is executing no longer covers the search area, relies on a sensor feed that went stale ten minutes ago, or has two assets assigned to the same task. Design-time verification and per-vehicle health monitoring were never built to catch that.
This is the gap runtime assurance is meant to close, and for mission autonomy it has to be closed at the level of the mission, not only the machine.
Why design-time verification and vehicle health monitoring fall short
Design-time verification is essential. Formal methods, simulation campaigns, and developmental testing establish that a system behaves correctly across the conditions its designers anticipated. The difficulty is that mission autonomy increasingly relies on planners and learned components whose behavior in novel situations cannot be exhaustively characterized in advance. The operational environment will produce combinations nobody tested.
Per-vehicle health monitoring answers a different question. It tells an operator whether each platform is functioning. It does not tell them whether the fleet’s course of action still satisfies the mission’s intent, given what has changed since the plan was generated.
The question program offices and operators actually need answered at runtime is: does the current plan still satisfy the mission, and if not, which requirement is violated and why?
The runtime assurance concept and simplex architecture
Runtime assurance (RTA) is a well-established pattern for this kind of problem. The classic form is the simplex architecture:
- a complex function (often high performing, hard to verify, or treated as untrusted) does the primary work
- a monitor checks the complex function’s outputs or the system state against a defined safety envelope
- a recovery path (a simpler, verifiable fallback) takes over or constrains behavior when the monitor detects a violation
The value of the pattern is that assurance no longer depends on fully verifying the complex function. It depends on verifying the monitor and the recovery path, which are deliberately kept simpler.
The aviation community has formalized this idea. ASTM F3269, published by ASTM International, is a standard practice for methods to safely bound the flight behavior of unmanned aircraft systems containing complex functions. It reflects the same insight: bound the untrusted component with a monitor and a recovery mechanism, rather than trying to prove the complex component correct in every case.
Most RTA work to date focuses on vehicle-level safety envelopes such as geofences, flight dynamics, and separation. Mission autonomy needs the same structure applied one level up.
Mission-level contracts for mission autonomy assurance
A mission contract is a runtime-checkable claim about a course of action (COA). Rather than asking whether a vehicle is safe, it asks whether the plan honors a specific mission requirement. Five contract classes cover much of what goes wrong in multi-platform missions such as ISR, logistics, and search and rescue:
| Contract | The question it checks | Example violation |
|---|---|---|
| Resource feasibility | Can assigned assets actually complete their tasks with available endurance, payload, and capacity? | A platform is tasked beyond its remaining endurance |
| Coverage persistence | Are required regions covered at the required rate for the required duration? | A search sector goes uncovered after an asset is reassigned |
| Task deconfliction | Are tasks and assets assigned without conflicting overlaps? | Two assets are assigned the same task while another goes unassigned |
| Contingency response | Does the plan account for active contingencies such as a lost link or a pop-up hazard? | A replan ignores a newly declared no-go area |
| Information freshness | Is the plan based on a current view of the world? | A COA keyed to an outdated world state is still being executed |
Each COA the planner emits is adjudicated against these contracts. The planner itself stays a black box. That is a feature: the monitor does not need access to planner internals, which makes it applicable across vendors and algorithms.
Auditable fault cards and evidence for test and evaluation
Detection is only half the requirement. For assurance to support decisions, every alert has to be explainable after the fact.
A fault card is a structured record of a contract violation. At minimum it should carry:
- the contract that was violated
- the COA and world state version it was evaluated against
- the evidence behind the determination (which assets, tasks, regions, and data timestamps)
- a recommended mitigation hook, such as forcing a replan with the active contingency injected as a hard constraint, or rejecting a stale COA
- timestamps for when the violation became true and when it was detected
Exported as structured logs, fault cards become test and evaluation artifacts. Testers can replay scenarios, compare detections against injected faults, and trace every alert back to its evidence.
Metrics that make runtime assurance measurable
A runtime monitor should be scored like any other detection system. Four metrics are a reasonable baseline for DoD autonomy test and evaluation:
- Precision: the fraction of raised faults that correspond to real contract violations
- Recall: the fraction of real violations that were detected
- Time-to-detect: the delay between a violation becoming true and the fault being raised
- Traceability coverage: the fraction of alerts whose evidence chain can be followed end to end
Randomized scenario sweeps with fault injection allow these metrics to be measured across many conditions rather than a handful of curated demonstrations.
The policy context for trustworthy autonomy
DoD policy points in the same direction. DoD Directive 3000.09, “Autonomy in Weapon Systems,” updated in January 2023 and available through the DoD Issuances site, emphasizes rigorous verification, validation, and test and evaluation, and requires that systems be designed to allow commanders and operators to exercise appropriate levels of human judgment. Although that directive addresses weapon systems specifically, its expectations about V&V and T&E rigor are a useful benchmark for autonomy more broadly.
The Department’s AI Ethical Principles, adopted in 2020, call for AI capabilities that are responsible, equitable, traceable, reliable, and governable. Runtime assurance maps directly onto several of these. Traceability requires that decisions can be audited. Reliability requires that behavior is tested and monitored across the life cycle. Governability requires the ability to detect unintended behavior and disengage or correct it.
A practical sequence for program teams considering mission-level runtime assurance:
- Write down mission intent as explicit, checkable requirements, not only as planner objectives.
- Map those requirements to contract classes and define what evidence each check needs.
- Identify the recovery paths available when a contract is violated, and who approves them.
- Build fault-injection scenarios for each contract class.
- Measure precision, recall, time-to-detect, and traceability coverage across randomized sweeps.
- Preserve fault logs as T&E evidence and review them with operators.
How QuantumWorks approaches mission-level runtime assurance
MissionGuard applies this approach: it checks black-box mission-autonomy outputs at runtime against five declared mission contracts and raises fault cards with evidence and mitigation hooks, exported as JSON and CSV for audit. It is an assurance monitor. It does not control vehicles, and it contains no targeting or weapons logic; its focus is assurance of mission autonomy for work such as ISR, logistics, and search and rescue. It is part of our autonomy and robotics work, and more on how we approach public-sector problems is on our government page. If your program is working through how to assure mission autonomy at runtime, we would welcome the discussion.