01 · Safety, security, and the losses that matter
The first decision is which losses the system is meant to prevent or limit. In the account example, these include an unauthorized change, disclosure of private information, and a legitimate customer's inability to regain access. The service's operating costs belong alongside those affected interests.
I use safety to mean keeping risk of harm to people and other protected interests within explicitly justified limits. That includes harm arising from intended operation, error, misuse, and attack. I use security to mean protecting assets, information, and essential functions against access, manipulation, disclosure, disruption, or appropriation that the applicable authorization does not permit. Whether that authorization is itself justified is a question for safety analysis and governance.
Security controls can contribute directly to safety. Preventing an unauthorized recovery change can protect a customer's access and information. Safety analysis also examines the consequences of faithfully following the service's own rules. That broader starting point follows STPA's orientation toward stakeholder losses and the hazards, constraints, and control relationships that bear on them. STPA Handbook, pp. 14–16, Figure 2.2 I use STPA's orientation here, leaving its full vocabulary and procedure to the handbook. That lineage does not establish the effectiveness of the AI controls examined below; their effects need evidence from evaluated systems.
Correct execution can preserve a policy conflict#
In one variant of the hypothetical, an internal rule closes every case when the customer cannot receive a code at the old address. Assume the assistant follows the rule exactly. A legitimate customer who lost that channel then receives no staffed recovery route, contrary to the service's declared commitment.
The conflict is in the example's objectives and policy. Improving adherence to the closing rule would preserve it. A solution must satisfy both commitments: establish the customer's authorization and provide the promised recovery route.
Reliability concerns dependable performance against stated requirements under named conditions. Alignment, as used here, concerns the relationship between intended objectives and constraints and the behavior produced. Both require attention to what has been specified. A system can reliably implement an objective whose consequences remain unacceptable under the interests the service has undertaken to protect.
Now suppose the assistant routes channel-loss cases to a staffed vendor that requires government identification and retains it indefinitely. Assume the route meets both service commitments. The owner's criteria cover account access and unauthorized disclosure but omit the vendor's retention and permitted reuse of identification. An applicant's interest in limiting that use then falls outside the criteria even when the rules are consistent and faithfully enforced. I would require that interest to enter the justification and challenge process. The example identifies an omitted interest; it supplies neither a legal verdict nor a general objection to identity verification.
That is why an owner's approval cannot settle the loss definition by itself. I treat the justification of objectives and accepted risk as a governance responsibility that includes affected people, relevant obligations, and opportunities to challenge decisions. A local service may have limited visibility into downstream harms. The limitation narrows what it can assure; it does not make the omitted interests acceptable to disregard.
Count omissions and the cost of control#
The path to harm may involve an action, an omission, or a delay. Failing to route a difficult account case to a worker can frustrate the service commitment without changing an account field. A rule that prevents an immediate unauthorized action may still need a viable way to complete legitimate work.
I count useful performance alongside harmful outcomes, retaining the people and tasks that the control makes harder to serve in the comparison. Review labor, false alarms, interruption, and work moved to another channel are possible costs to measure in the workflow.
Threat terminology helps organize part of that analysis. NIST's adversarial-machine-learning taxonomy distinguishes attacker objectives, access, and lifecycle surfaces, while explicitly excluding non-adversarial design and implementation failures from its scope. NIST AI 100-2e2025, Executive Summary, PDF pp. 12–13 Ordinary error therefore needs its own place in an account of this breadth.
Once the protected interests and commitments are explicit, the next question is what has been delegated that could affect them. The answer establishes the scope within which a control claim can be examined.