Omnafy

July 22, 2026

Autonomy is earned: how we tier agent risk

"Human in the loop" is the most agreed-upon and least specified idea in enterprise AI. Everyone says it. Almost nobody says which loop, which human, for how long, or what would have to be true to stop.

Left unspecified, it collapses in one of two directions. Either every action needs approval forever, in which case the agent has moved the work rather than removed it. Your staff now read proposals instead of doing tasks, and the economics never close. Or supervision quietly erodes, approvals become reflexive, and you end up with full autonomy that nobody ever decided to grant.

We avoid both by making supervision a property of the capability rather than of the agent, and by defining in advance what graduation requires.

Four tiers

Every tool in the catalog is classified into one of four risk tiers when it is built. The tier decides how much supervision the operation requires and whether it can ever require less.

T0, Observe. Read-only monitoring, reports, alerts. Nothing changes state. Autonomous from day one; there is nothing to approve.

T1, Internal actions. Reversible writes inside systems you control. Proposed and approved to start with. As evaluations accumulate, individual tasks graduate to autonomous, one task at a time and never in a batch.

T2, External actions. Anything that leaves the building: records written to an external system, a fax, a message to a patient or a partner. Reversibility is weaker or absent, so these stay approval-first for considerably longer and graduate selectively.

T3, Judgment calls. Decisions that belong to a licensed professional stay with that professional. The agent assembles the case and recommends with evidence; a person decides. T3 never graduates. That is not a maturity milestone we have yet to reach. It is a permanent property of the tier.

What graduation actually requires

Autonomy is granted against measured accuracy on the specific task, not against a general sense that the agent is doing well. Every approval decision a reviewer makes, whether approved, corrected, or declined, is an evaluation label. Those labels accumulate into a record for that task, and the record is what a graduation decision is made against.

Two properties matter as much as the threshold:

  1. Graduation is per task, not per agent. An agent that has earned autonomy on visit scheduling has earned nothing on billing adjustments.
  2. Graduation is reversible. If accuracy degrades because a process changed, an upstream system changed, or the caseload changed, the task drops back to propose-and-approve. That is a routine operational action, not an incident.

Approvals have to show the evidence

A tiering scheme is worthless if the approval step is a bare yes or no. If the reviewer has to open three other systems to check the agent's work, you have added cost rather than control, and within a month they will stop checking.

So every proposed action arrives with the evidence attached: what the agent read, what it inferred, what it proposes to do, and what happens if it is wrong. The reviewer's job is to exercise judgment on a prepared case. That is work worth a person's time, it is fast enough to keep up with, and it produces the labels that let the task graduate.

Autonomy earned this way is defensible. You can say exactly which capabilities run unsupervised, on what evidence, and what would send them back.