How Safe Is Safe Enough — and Who Decides?

Evidence, not footage — post 3 in the Physical AI Deployment Assurance series.

Post 2 parked a question in the agent pillar.

The vendor is asked for a safety number and handed no threshold.

Safe compared to what?
At what confidence?
Decided by whom?

This post picks up that question.

Post 2 ended with the shape of the problem: risk numbers are easy to produce; thresholds are hard to own. Here is why, and who is supposed to do the owning.

Imagine a hospital considering whether to move a delivery robot from pilot to routine operation. The vendor can report a collision-risk estimate. The safety team asks whether the estimate is low enough. The hospital board asks who is supposed to sign off. The insurer asks how the remaining risk should be priced. The engineer can compute a number, but the organization still has to decide what number it is willing to own.

That question looks technical.

It is not.

It is a question about signatures, which means it lives in the operations pillar while wearing the agent pillar’s clothes.

The mathematics starts too late

Let me begin with a confession about my own field.

Nearly every tool in quantitative safety takes the threshold as an input. Call the threshold α: the tolerated probability, tail mass, or risk level below which the system is considered acceptable for the purpose at hand.

Risk-constrained control bounds the conditional value-at-risk at level α. Conformal methods guarantee coverage 1−α. Chance-constrained planning caps the violation probability at α.

Different machinery, same signature:

Give me α, and I will give you guarantees.

I have spent much of my career on the mathematics that begins after α has been chosen.

This post is about what happens before.

In many projects, what happens before is almost nothing. Ask where α came from and the room goes quiet. The number arrives from a previous project, from a reviewer’s intuition, from an internal convention, or from whatever made the simulation pass.

The most consequential parameter in the safety argument is often the one with the thinnest paper trail: the threshold everyone computes against, but no one remembers choosing.

Three answers from older industries

Physical AI is not the first field to face this problem.

Aviation, rail, automotive, and process industries confronted versions of the same blank long before robots began sharing corridors with people. Their answers sort into three broad archetypes.

Comparative

The new system must be at least as safe as what it replaces.

European rail practice has used principles in this family — the French GAMAB principle, roughly “globally at least as good”: the new system should be at least as safe as the existing one. The baseline is history. The threshold is not chosen from nowhere; it is chosen relative to a system society already accepts.

Absolute

Anchor the threshold to something outside the immediate system.

Aviation offers the most famous engineering example. For catastrophic failure conditions in large transport aircraft, certification guidance has long used “extremely improbable” objectives, commonly associated with probabilities on the order of 10⁻⁹ per flight hour.

That number is often treated as bedrock. But the usual reconstruction is more interesting. The existing fleet’s accident record implied a tolerable rate of catastrophic accidents. A share of that budget was attributed to system failures. The budget was then divided across the many potential catastrophic failure conditions of an aircraft. Out came the number.

The apparent constant began life as a comparison:

No worse than the fleet we already accept.

Through decades of use, that comparison hardened into an absolute.

Thresholds are born as comparisons and retire as constants.

Procedural

Refuse to name a universal number; name a duty instead.

The UK’s ALARP doctrine — risk must be as low as reasonably practicable — does not begin by declaring one global threshold. It asks what additional safety measures are available, what risk reduction they would buy, and whether the sacrifice required would be grossly disproportionate to the benefit.

Regulators may publish tolerability bands around such reasoning. But the doctrine’s core is procedural. It does not say, once and for all, “this number is safe.” It says that the responsible party must keep reducing risk until the case for stopping can be defended.

One more device deserves naming.

Automotive and industrial functional-safety standards often route the question through tables. Parameters like severity, exposure, and controllability go in; an integrity level comes out. Notice what the table really does: it distributes the signature.

No individual engineer chooses α. A committee chose the table, and the engineer applies it. For mature hazard classes with decades of data, that is not necessarily a flaw. It is the point. The value judgment gets made collectively once, then retired from daily engineering.

Physical AI does not yet have such a table.

It lacks the two ingredients that make tables legitimate: accumulated failure data and a community willing to turn value judgments into shared thresholds.

Why Physical AI bends all three answers

Physical AI inherits these archetypes, but bends each of them.

The comparative answer

The comparative answer arrives in one familiar phrase:

Safer than a human.

The phrase is doing far less work than it appears to.

Which human?

The median operator?
The trained professional?
The tired worker near the end of a shift?
The distracted visitor?
The impaired tail that drives much of the baseline statistics?

Safer on average, or safer everywhere?

A system can beat the human aggregate while being worse in specific, describable situations. Describable matters, because describable scenarios are the ones that become procurement objections, safety concerns, insurance exclusions, and litigable narratives.

Then comes an asymmetry no threshold can ignore: machines can fail in ways humans rarely do, and publics may weigh machine failure differently from human failure.

I am not endorsing that asymmetry or complaining about it. I am observing that any threshold which pretends it does not exist may not survive contact with its first incident.

I will not tell you whether “safer than the average human” is enough.

I am telling you the sentence is underspecified.

The absolute answer

The absolute answer needs an anchor, and the anchor needs exposure history.

Aviation could build thresholds on enormous accumulated exposure and comparatively standardized operating regimes. Physical AI has neither the exposure base nor the site homogeneity.

Post 2 used the word specific deliberately. Deployment assurance is site-specific. The same robot can be low-risk in a quiet logistics mezzanine at 3 a.m. and much higher-risk in a hospital corridor during shift change.

Aviation could average over enormous flight hours.

A hospital corridor does not average the same way.

The procedural answer

The procedural answer presumes that the next increment of safety can at least be argued about: what measure is available, what risk reduction it buys, and whether the sacrifice is grossly disproportionate to the benefit.

For learned systems, even that can be unstable.

Is the next increment of safety more data?
A different architecture?
A restricted operating envelope?
A lower speed limit?
A human in the loop?
A better map?
A runtime monitor?
A different sensor suite?
A narrower deployment site?

The options are not enumerable in advance. The marginal cost of risk reduction is often unknown until after the engineering work begins. That makes “reasonably practicable” a moving target.

And underneath all three answers sits the deepest bend.

A learned controller’s risk number is conditional.

The honest object is not α.

It is:

α | A

A risk threshold conditional on an assumption set.

A is not fine print. It is part of the claim.

The honest statement is never simply:

10⁻⁷ per hour.

It is:

10⁻⁷ per hour, under assumptions A.

Lighting within tested bounds.
Human traffic within the training distribution.
The map current.
Network latency within the validated range.
No unmodeled obstruction in the designated path.
No adversarial interference.
No post-deployment update that changes the controller’s behavior outside the evidence base.

Signing α without signing A is signing nothing.

So the question doubles.

Who signs the number?
Who signs the assumptions?
And who is watching A on the Tuesday after the renovation, when the assumption set has quietly stopped being true?

A signable safety claim therefore needs more than a threshold. It needs a threshold, an assumption set, and a monitoring plan for detecting when those assumptions have stopped being true.

The circle of deferral

Now watch the actors handle the blank.

The operator asks the vendor: just tell me it is safe. A yes-or-no question about a conditional distribution.

The vendor asks for a spec: give us the number and we will meet it. This is rational, not evasive. Authoring the number is authoring liability.

The regulator writes process, not numbers. A number creates exposure in both directions: too strict and it may stall an industry; too loose and it may be blamed after the next accident. Process requirements are more survivable than constants.

The certifier audits what the standard specifies. Where the standard contains no number, it audits process. “You followed your procedure” may be true, useful, and still incomplete.

The insurer is the actor structurally forced to produce a number, because the premium is one. With little actuarial history, it prices conservatively, or declines to price at all. In practice, that price can quietly become one of the industry’s working thresholds.

The court chooses last, after the accident, with the outcome visible and the counterfactual invisible — an epistemically difficult seat, and often the one holding the final signature.

Trace the circle.

Everyone touches the threshold.

No one signs it.

Six actors — operator, vendor, regulator, certifier, insurer, and court — arranged in a circle around the safety threshold α, each deferring the choice to the next.
Everyone touches the threshold. No one signs it.

And when the circle completes without a signature:

If no one chooses α, the accident chooses it.

That is not a moral judgment about any single actor. It is a system failure.

When the threshold is not owned before deployment, it is often reconstructed after an incident, under grief, scrutiny, and litigation.

An accident is a terrible chooser.

It sets the threshold retroactively, one case at a time — and often for an entire industry, not just for the party that failed.

What α actually is

I promised no answer, and I will keep the promise.

What I can offer is a decomposition, because much of the dysfunction above comes from treating α as a number.

α is not a number.

It is a five-part sentence:

This level of risk — relative to this baseline — under these assumptions — valid until these monitored conditions change — on this party’s authority.

That sentence has five parts.

PartQuestion
NumberWhat level of risk is being accepted?
BaselineSafe compared to what?
Assumption setUnder what conditions does the claim hold?
Validity windowUntil what changes?
SignatoryOn whose authority?

Remove any one part and the remaining four become difficult to sign.

Most deployment conversations I have seen negotiate the first part and improvise the other four after the contract is signed, which is another way of saying: after it is too late.

Here is the honest division of labor.

Engineering can make the sentence legible. It can compute the number, expose the assumptions, monitor their decay, quantify uncertainty, and clarify the baseline comparison.

Choosing the sentence is a governance act.

Governance here does not mean bureaucracy. It means deciding which residual risk an organization is willing to own, under what assumptions, and on whose authority.

The recurring mistake runs in both directions: asking engineers to smuggle governance inside the mathematics, or asking boards to sign mathematics they cannot read.

The job of deployment assurance, in the end, is to hand each party a sentence it can actually sign.

Where this goes

If the threshold is a five-part sentence, the next question is what a signable version of that sentence looks like in practice.

What has to be in a safety claim for a skeptic with something at stake to accept it?

What assumptions must be stated?

What evidence must be attached?

What monitoring plan keeps the claim alive after deployment?

And what is usually missing?

Risk numbers are easy to produce. Thresholds are hard to own. A defensible safety claim is the document that makes ownership possible.

That is the next post:

The anatomy of a defensible safety claim.


Views are my own and do not represent my employer or any organization I advise.