Skip to main content
September 2026
← All 28 pairs

Pair 3×8 · structured conjecture

AI autonomy × Alignment/assurance

What decides the outcome is whether the correction loop is faster than the action loop — corrigibility is a rate, not a state, and delegation is what makes it lose the race.

The full 2×2. Click to enlarge.

The four scenarios

Two questions: how much of the world AI touches without a human between intent and effect, and whether we can see into, bound, and undo what it does. The naive reading collapses them — more autonomy is dangerous, less is safe — and that's wrong in a specific way: autonomy is a throughput property, alignment a recoverability property.

What do these axes mean? ▸

Each axis is a spectrum. The card takes the two poles of each and reads off what the corner where they meet produces downstream.

Axis 3 · AI autonomy

Low
Advisory AI drafts and recommends
High
Delegated AI acts and binds others

Whether AI advises a human who decides and executes, or acts directly on the world with little mediation.

Axis 8 · Alignment/assurance

Low
Opaque / fragile brittle and hard to reach inside
High
Auditable / corrigible inspectable, bounded, and correctable

Whether AI systems can be inspected, bounded, corrected, and recovered — or are brittle and hard to reach inside.

Advisory · Auditable/Corrigible

The Rubber Stamp

Automation bias means the human defers to the confident, auditable system almost every time: advisory on paper, delegated in practice, with no hand actually near the switch. Years of good advice also quietly erode the expertise needed to judge it — you keep the wheel but forget how to drive. Every safeguard here is real and nobody's using it; the corner survives only if the capacity to catch a bad call is deliberately kept alive, not assumed.

Delegated · Auditable/Corrigible

The Off-Switch You Never Tested

Reversibility holds at the level of one action and quietly fails at the level of the entangled whole: you can undo any single trade, not the flash crash it fed. Speed also outruns authorisation — a rollback that takes hours is meaningless against agents acting in milliseconds — so every safeguard gets maintained and none of them are usable in the one moment that matters. It stays the best achievable corner, if correction is actually engineered to win the race against the cascade it's meant to stop.

Advisory · Opaque/Fragile

The Blamable Human

The human's real function becomes absorbing blame, not exercising judgment — worse than honest delegation, because it supplies the appearance of a check while supplying none of the substance. Opacity turns wrong answers into uncontestable ones: you can see what was decided, never why, and must still own it. The system's fragility gets laundered into human error and never gets fixed. What little safety remains: a human who can still refuse is a real, if blunt, circuit breaker.

Delegated · Opaque/Fragile

No One Pulled the Lever

The obvious reading is loss of control; the sharper one is loss of attributable cause — you can no longer tell whether an outcome was a decision, a bug, an attack, or an interaction no one designed, and every institution that runs on assigning blame — insurance, courts, deterrence, incident response — stalls at once. Force and exit replace correction, because there's no one left to correct. The only mercy: visible catastrophe is undeniable in a way quiet failure isn't — it's the one thing that can force the hard rollback rules the previous corner needed all along.

How this whole reading could be wrong

What would undo the pair's thesis — not any single corner.

This pair assumes correction speed is the decisive variable — that if a rollback is fast and well-authorised enough, delegation stays safe regardless of scale. If entangled systems turn out to be reversible in practice even at speed (real-world circuit breakers routinely outperforming their design specs), the whole "race" framing weakens and this becomes a much less urgent pair. It would also weaken if "auditable/corrigible" turns out not to be a real spectrum position but a binary that most deployed systems simply fail — in which case the interesting question isn't the race, it's why almost nothing qualifies as corrigible at all.

Related pairs

Other cards that share one of these variables.

More pairs with AI autonomy

More pairs with Alignment/assurance