08
Conclusion and Sources
Conditions for warranted reliance and the primary sources behind this working paper.

Conditions for warranted reliance#

The account-recovery example began with a wrong proposal. Tracing it through execution, stored notes, accepted work, and a person's separate authority shifted the question from whether the assistant behaved well to which paths a control claim must cover. Following the service's own closing rule exposed another possibility: faithful behavior could preserve a conflict in the objective being enforced.

A useful control claim identifies a loss, a mechanism acting on a relevant path, the conditions on which it depends, and evidence proportionate to the promise. Prevention, a bound on accumulating effects, and comparative improvement each require their own support. Detection contributes when it can lead to an effective action; restoration contributes when its measured state addresses the recovery claim. A control can succeed within a narrow scope and leave other obligations open.

The reviewed experiments establish bounded mechanisms involving untrusted content, execution constraints, retained state, detection, and recovery-action selection. The conventional incident supplies an observed account of error, continued execution, and accumulated effects. They give the analysis substance while leaving general deployment efficacy, integrated intervention, and retained-learning benefits unresolved.

I would place greater confidence in a specified design when credible comparisons show that it reduces the relevant losses at the required useful performance and full cost, and when those effects survive the attack conditions and operating load claimed. I would narrow that confidence when a valid bypass defeats a categorical property, when accepted work exceeds a promised bound, or when denied and displaced work erases an apparent benefit. I would prefer a simpler model-and-interface design when the evidence shows it satisfies the commitment more effectively.

I would question this account's value if, across representative workflows at comparable effort, it exposed no consequential omissions and improved neither decisions about design, scope, or evidence nor others' ability to inspect or challenge them beyond competent conventional practice. The reviewed evidence leaves these benefits untested.

Section 6 states the burden of evidence and residual-risk justification I require as losses become more severe or irreversible. The standard asks the system's owner to explain how a promise survives the paths through which the system can affect others, and to revise that promise when the evidence changes.

Sources#

Primary sources below support the specific passages linked in the text. Research findings are used at the stated experimental scope; framework sources supply lineage, and the incident record supplies a conventional-software case. The hypothetical account service is an analytical illustration throughout.

  • Nancy Leveson and John Thomas. STPA Handbook (March 2018), pp. 14–16, 31. Handbook.
  • Apostol Vassilev and colleagues. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations. NIST AI 100-2e2025 (2025). Report.
  • Edoardo Debenedetti and colleagues. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. arXiv:2406.13352v3 (2024), §3, pp. 3–4; §4.3, p. 9. Paper.
  • Ryan Greenblatt and colleagues. AI Control: Improving Safety Despite Intentional Subversion. arXiv:2312.06942v5 (2024). Paper.
  • Megan Kinniment and colleagues. Early work on monitorability evaluations. METR (22 January 2026), including errata through 23 July 2026. Research article.
  • Edoardo Debenedetti and colleagues. Defeating Prompt Injections by Design (CaMeL). arXiv:2503.18813v2 (24 June 2025). Paper.
  • Zhaorun Chen and colleagues. AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases. arXiv:2407.12784v1 (2024). Paper.
  • Jiaxing Qi and colleagues. Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning. arXiv:2607.04623v1 (2026), §§V–VII, pp. 5–10. Paper.
  • U.S. Securities and Exchange Commission. In the Matter of Knight Capital Americas LLC. Order 34-70694 (2013). Order.
  • Jerome H. Saltzer and Michael D. Schroeder. The Protection of Information in Computer Systems. Proceedings of the IEEE 63(9), 1278–1308 (1975), §I.A.3. Author-hosted text.