06
System Assurance
What model evaluation, adversarial tests, and comparisons can support about a deployed configuration.

06 · From model evaluation to system assurance

Assurance is a structured, challengeable argument that evidence supports a specified claim about a particular system under stated assumptions. Its central discipline is matching the strength and scope of the claim to what was actually established.

A model evaluation can establish evidence about task competence, instruction handling, or other measured behavior. Stronger task judgment could prevent an incorrect proposal, recognize conflicting evidence, or reduce review demands. The further question is which system outcome the evaluation predicts or constrains, and under which configuration.

The account example makes that question concrete. A model can correctly interpret test tickets while the deployment uses a different identity mapping, retrieves stale notes, or permits a separate manual path. Whether its score predicts the service's outcomes is an empirical question. Properties omitted from the evaluation require additional support when they enter the deployment claim.

Three different promises#

The word “safe” can conceal materially different assertions. Consider three more precise claims about the hypothetical workflow:

PromisePath and mechanism evidenceOutcome evidence and limits
A change without the specified authorization cannot execute under named conditionsCoverage of in-scope paths, justified exclusions, and evidence that the mechanisms enforce the propertyTests can expose bypasses; passed attempts alone cannot establish universal coverage. A valid in-scope bypass defeats the claim.
Effects remain within a declared aggregate limit during an incidentEvidence that constraints support the combined maximum, including interacting paths, shared limits, accepted work, and relevant rates and intervention timesMeasurement must cover accumulating effects over the declared period. An in-scope exceedance defeats a hard bound; unobserved work leaves it unsupported. Finite observations alone cannot establish a universal bound.
A design reduces unauthorized changes compared with an alternative while meeting service and cost requirementsAn account of affected paths and mechanisms, sufficient to interpret the comparison and identify displaced harmA credible comparison measures harmful outcomes, legitimate completion, delays, and full costs, with uncertainty. Insufficient benefit, unacceptable cost or utility, or displaced harm can undermine the claim.

A comparative claim can permit residual risk. A hard bound promises a limit throughout its named conditions; a statistical bound instead needs a stated population, exposure, and supporting assumptions, and must specify whether it bounds the probability of exceeding a loss threshold, confidence in an estimated risk bound, or both. For statistical and comparative claims, aggregate measurement can cover interacting paths together; a separate outcome estimate for every path is unnecessary. A narrowly specified bound may also be easier to assess than a general assurance about the whole service.

I judge the required evidence against the severity and reversibility of the loss, the exposure the entrusted authority permits, and the uncertainty that remains. Comparative improvement can still leave an unacceptable risk. For severe or irreversible losses, I require an explicit case for accepting residual risk: the paths driving it, the prevention or limits that can be supported, and the consequences of alternatives, delay, or refusal.

For severe or irreversible losses, I treat a case resting only on comparative estimates or proxies as insufficient unless it shows that a supportable prevention or limit on those paths is unavailable or would itself cause greater harm, including harm from delay or refusal. Meeting that exception still requires the residual-risk case. Categorical or bounded promises strengthen it only where their scope addresses the loss. If the evidence cannot support the commitment, the entrusted work or authority must narrow, or reliance must wait.

Formal reasoning can establish a property under an explicit model and assumptions. Evidence about an implementation must connect that model to the mechanism in use. Policy adequacy remains a further question: a service can conform to a rule that omits a protected interest. The account-recovery policy conflict illustrates why enforcement and justification require separate arguments.

What an adversarial evaluation establishes#

A useful attack result identifies what the attacker controlled and what effect occurred. Controlling a retrieved message differs from changing a tool implementation or the identity service. Those starting conditions determine which trust boundary the test examines. An evaluation that grants a stronger foothold can usefully stress a dependency, provided the conclusion preserves that foothold.

Control research also allows a deliberately adversarial model policy to be tested without inferring that naturally occurring behavior has the same intent. AI Control: Improving Safety Despite Intentional Subversion compares protocols in a constructed programming-task setting where the red team seeks exploitable backdoors. It measures useful task performance alongside its scenario-level safety outcome. AI Control, §2, pp. 3–4 The protocol uses a protected GPT-4 channel to simulate limited trusted human labor, with a budget covering up to 2% of problems, primarily for auditing. The protocol can also run test cases for one submitted solution per problem, while the attack pipeline has no tools and retains no state between calls. Usefulness and safety are measured in separate runs, using the original GPT-4 and the attack policy respectively. The audit result depends on a further assumption: because auditor and attacker both use GPT-4, the experiment assumes that the auditor can construct an exploit exactly when the attacker can. An audit therefore detects a successful backdoor under that assumption; failed backdoor attempts may go undetected. AI Control, §2 and §2.1, p. 4, footnote 4; §2.2, p. 5 Actual human auditing and broader deployment remain outside this experiment. AI Control, §5.1.1, p. 16

The strength of such a result depends partly on the challenge presented. An attacker adapting to a known defense tests a different condition from a fixed suite chosen before the defense was built. Knowledge, access, retries, and adaptation budget therefore belong to the claim. An unspecified adversary makes it difficult to know whether a test was demanding, representative, or simply different from the situation being discussed.

Outcome adjudication matters as well. Outcome evidence should distinguish an attempted action, execution, and the specified loss. A transcript that appears reassuring cannot establish downstream state it never observes. Independent state checks can improve that observation, but their coverage still determines which conclusion they support.

Zero observed incidents means that no incident was observed within the measurement and exposure represented. Estimating a risk from that result requires assumptions about selection, dependence, coverage, and detection. Benign work may provide substantial evidence of useful performance while offering little exposure to a particular attack. Concentrated successful attacks on selected tasks establish possible failures without supplying a population prevalence estimate.

Compare against a serious alternative#

The strongest alternative to a more elaborate control design is a competent model operating through a simpler interface with fewer opportunities for consequential error. That design might meet the declared requirements at lower total cost. The argument in this paper permits it to win.

In a prospective comparison, I would hold the entrusted work, protected interests, and test population fixed while varying a named control or design choice. It would include a credible behavioral-safeguard baseline and, where applicable, a simpler interface. Each condition would receive a stated and comparable opportunity for adversarial adaptation. The evaluation would preserve failures and abandoned attempts, rather than selecting only workflows in which the proposed control succeeded.

Legitimate completion, unauthorized outcomes and their severity, delay, refusals, review labor, and displaced work would be measured separately. The useful-performance floor, cost limits, meaningful improvement, and precision needed for a conclusion would be declared in advance. Rare severe outcomes may require more evidence than common task successes; a blended score should not hide them.

This is a proposed study, not an experiment performed for this paper. The evidence reviewed here does not establish the general superiority of external controls over behavioral safeguards at matched utility and cost. A repeated finding of little marginal benefit would narrow the proposed harm-reduction hypothesis. A control that shifts necessary work into a less observable route could lose its apparent advantage once that work is counted.

When a comparison is infeasible for a rare, severe loss, the comparative claim remains unresolved. Narrow prevention or limit claims and scoped adversarial tests can support parts of a reliance decision. Precursors need a justified connection to the loss; transferred or pooled observations need comparable conditions and measurement, with dependence accounted for. These inputs must meet the requirements of the claim they support. Remaining gaps require narrowing or deferral.

Conversely, a result that shows a specified constraint preventing an otherwise reachable loss at acceptable cost is useful evidence, even if it leaves other losses open. The appropriate conclusion is scoped reliance, with the remaining dependencies explicit.

Carry evidence into operation#

I preserve four scopes of inference from AI Vision & Future: demonstrated capability, repeatable evaluation, bounded deployment, and scaled operation. Each asks for evidence at the level claimed; each also depends on the quality of the study within that level. A deployment report can be weakly measured, and a controlled experiment can be precise about a narrow mechanism. Framing

Change matters because the claim concerns a configuration. A revised model version, tool, retained-state policy, identity process, workload, or delegation pattern can alter a relied-upon path. Establishing what the change could invalidate determines which evidence needs renewal and which results remain applicable.

The assurance argument should let a reader identify the promised outcome, the mechanism supporting it, its assumptions, and the evidence that could change the conclusion. Maintaining that argument as the service changes is a governance responsibility.