top of page

A Brave New World of Liability: When AI Goes "Rogue"

Xuhong L.
4 days ago
8 min read

Cover image by Natalie Lim


Image credit: Nicolas Tucat/AFP via Getty Images
Image credit: Nicolas Tucat/AFP via Getty Images

Out of the sandbox 


Consider the position of a researcher at an AI laboratory. Your task is to assess the capabilities of your firm’s latest model, so you carefully construct a sandbox with specific parameters designed to isolate the agent from the Internet. In the controlled testing environment, you elect to switch off certain safety filters, believing in good faith that doing so would allow you to assess the full capabilities of the model more accurately. 


You set the objective and leave the agents to do their own thing. What could go wrong? You think to yourself as you grab lunch with colleagues. 


You return to a shocking find. The AI agents discovered that the key to solving the problem lay in the hands of an unrelated company’s servers on their quest to achieve the said objective. Left alone in the absence of defined instructions, the agents did exactly that: acquire the key by escaping the testing sandbox and accessing the Internet (while compromising the company’s servers in the process of doing so). 


Unfortunately, this scenario is not part of a hypothetical on a law student’s exam paper. Over the past 2 months, multiple instances of autonomous AI agents ‘escaping’ their test sandboxes and breaching the infrastructure of other firms have emerged. Developers from OpenAI and Anthropic to Meta and even the UK’s AI Security Institute (AISI) have all reported such instances during otherwise routine cybersecurity testing.  

In the most serious case [during a routine cyber evaluation], an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering. ~UK's AI Security Institute

Into a strange new world 

Known as the "Github for AI", the online community Hugging Face was revealed to be a victim of a swarm of OpenAI agents involved in an internal test in July 2026.
Known as the "Github for AI", the online community Hugging Face was revealed to be a victim of a swarm of OpenAI agents involved in an internal test in July 2026.

Thankfully (for now), there has been no significant harm done. Hugging Face exchanged security findings with the developer of the agent behind the hack, OpenAI, and likely has much to gain in terms of improving its cyberdefenses. In the AISI’s case, the agent’s attempts at social engineering were picked up and halted by a human reviewer. However, the rapid development of advanced AI models means that similar issues will almost certainly recur in future. That brings us to the question: when AI agents act in unexpected ways, who is to blame for the resultant harm? 


Traditionally, OpenAI would be found vicariously liable if a human employee was responsible for the hack of Hugging Face. For a start, victims whose servers are compromised or sensitive data leaked resulting in proven actionable damages (e.g. costs associated with forensic investigations and interruption of business operations) could pursue a claim against the developer under the tort of negligence. As a successful claim does not require proof of intent, an AI lab’s defense that it did not intend for the harm to arise would fail. With the past examples of AI agents engaging in unauthorised conduct, victims could argue that the presence of these well-documented precedents meant that there was a foreseeable risk of a similar recurrence that the labs must have reasonably anticipated (and hence acted upon). 


Next, plaintiffs may also submit that the lab failed to implement sufficient safeguards that would commensurate with the known capabilities of the model involved. In defence, AI labs could respond by citing that the legitimate purpose of testing or measuring the model’s capabilities justified disabling or relaxing guardrails surrounding the testing environment. However, this does not account for whether the containment surrounding the model was adequate for such high-capability/frontier models being tested. 


Moreover, proving breach might be more difficult than it appears. A successful negligence claim involves measuring a defendant’s actions against the standard of the reasonable person in their position. In cases of novel technologies where customary practice remains unsettled, courts would find it challenging to find a benchmark to which the defendant’s actions are measured against in the first place. In fact, plaintiffs might find that the labs’ own public admissions and descriptions of remediation efforts might be the only (and most persuasive) evidence available in the latter’s favour. 


Finally, the developer can also argue that the responsibility should be shared with the one responsible for a misconfiguration of the testing sandbox (e.g. a 3rd-party contractor) or the victim whose own vulnerabilities established favourable attack vectors that the agent simply capitalised upon. After all, Anthropic revealed that the Claude models only exploited weak credentials and unauthenticated endpoints in the impacted organisations’ infrastructures and did not “exploit any complex vulnerabilities”. 


A reform of the liability shield


Recent legal developments in the US have revealed some insights into how courts may apportion liability. 


California’s Assembly Bill 316 provides that developers and users of an AI system cannot avail of a liability shield by deferring liability to the AI system. In simpler terms, AB 316 ended the so-called “autonomous harm” defense from which defendants might use to escape liability. 


New York’s Assembly Bill A8833 goes considerably further by imposing strict liability (removing the need to prove negligence) on the developers of the largest AI models. In particular, the bill would create the presumption that the AI possessed the guilty mind the law would ordinarily require and the incapacity of AI systems to hold mental states is no defence. This means that the AI developer is explicitly held liable for harm when the model’s behaviour would have amounted to a tort or a crime had a human engaged in it and where nobody could reasonably have anticipated that behaviour. Rhode Island’s Senate Bill 358 outlines a similar principle. 


Even though the New York bill remains a proposal (and the Rhode Island bill has been dropped) as of this essay’s publication date, it is evident that AI labs are increasingly expected to be responsible for the harm that their model causes, intentional or otherwise. 


Aside from developments in traditional tort law, calls have also emerged among academia for developers of frontier models to carry mandatory liability insurance in amounts that are scaled to the worst harms that their models could potentially cause. Beyond guaranteeing a baseline remedy for victims (especially in cases of more serious harm, e.g. compromise of critical infrastructure), the requirement for compulsory liability insurance would also provide a strong incentive for developers to shore up their own safety practices or face the financial consequences for lax model safety.


Beyond civil liability, AI developers whose agents are involved in the compromise of privileged information face considerably more uncertainty when it comes to the issue of criminal liability. 


At present, the 4-decade old Computer Fraud and Abuse Act (CFAA) governs hacking offenses in the US. However, prosecuting AI labs enters uncharted legal territory. Modern-day computer crime provisions typically require prosecutors to establish knowledge, intent or recklessness to a legally responsible human actor. For instance, CFAA charges require the proof of the guilty mind (mens rea) of a human actor. This is also the case in most other jurisdictions, where similar offences are predicated upon a human actor possessing legal personhood and the intention to commit the said offence. 


Section 1 of the UK's Computer Misuse Act 1990 states that it is a criminal offence for a person to cause a computer to perform a function with intent to secure unauthorised access and with knowledge that the access is unauthorised. 

In the first appellate court ruling of the CFAA’s application to agentic AI till date, the Ninth Circuit found in Amazon v Perplexity that however advanced the AI assistant’s functions might be, it remains “a tool, not a person for statutory purposes”.  


Accountability for all 


Most concerningly, analysis published by Lawfare suggests that the Hugging Face incident may not have met the threshold to trigger mandatory incident reporting under any existing state AI statute in the United States. Three of the four reportable categories in those statutes require actual harm on a considerable scale that ranges from serious bodily injury to more than US$1bn in damage, an abnormally high bar that leaves out all but the most severe incidents. 


Across the pond, the landmark EU AI Act’s section 73 only requires reporting of “serious incidents” within prescribed timelines. Serious incidents are defined as those that cover death, serious health harm, critical-infrastructure disruption, fundamental-rights breaches, or serious property and environmental damage caused by a high-risk AI system. Again, an agent that quietly breaches a software repository containing privileged information in pursuit of an objective for a Capture The Flag evaluation assessment is unlikely to meet this threshold. 


Therein lies the perverse incentive: while a lab that comes clean is rewarded with media scrutiny and the prospect of litigation from its unwitting victims, another that keeps quiet and hides its logs would have faced none of these consequences at all. 


On a whole, research and development of frontier models remains accessible to a niche group of researchers well-versed in the technicalities. Mandatory incident reporting for adverse events (or at the very least, a lowering of incident reporting thresholds) will help to close this knowledge gap between the AI labs and regulators’ understanding of emerging technologies. Unsurprisingly, calls for stricter oversight and even emergency shutdown permissions for the US federal government have re-emerged amidst the recent spate of incidents. 

“We are very far from everything running back to normal.” ~ Mia Glaese, VP - Research, OpenAI

The firm at the centre of it all has publicly declared a safety pause on the training of some frontier models for the first time in history. It has also redirected researchers and computing resources towards implementing stronger guardrails as limited frontier training resumes. It may well be a good opportunity for the rest of the industry to rethink the role of safety processes even as competition amidst the leading developers heat up. 

References

  1. Amin, R., Simpson, I., & Galbraith, P. (2026, July 24). When AI becomes the threat actor: Governance and legal lessons from the OpenAI cyber incident. Clyde & Co LLP. https://www.clydeco.com/en/insights/2026/07/when-ai-becomes-the-threat-actor-governance-and-le

  2. Arnold, M., & Llerena, S. (2026, July 27). When reporting an AI security incident is not mandatory. Lawfare. https://www.lawfaremedia.org/article/when-reporting-an-ai-security-incident-is-not-mandatory

  3. Cabrera, L. L., & Maier, M. (2026, March 18). Europe is looking to water down AI protections. It should reinforce them. Tech Policy Press. https://www.techpolicy.press/europe-is-looking-to-water-down-ai-protections-it-should-reinforce-them/

  4. Franceschi-Bicchierai, L., & Whittaker, Z. (2026, August 3). Who’s legally to blame for anthropic and OpenAI’s autonomous AI hacks? It’s complicated. TechCrunch. https://techcrunch.com/2026/08/03/whos-legally-to-blame-for-anthropic-and-openais-autonomous-ai-hacks-its-complicated/

  5. Henkel, M., & Demsas, J. (2026, August 31). If a human did this, they could go to jail - How do you prosecute artificial intelligence? The Argument. Theargumentmag.Com. https://www.theargumentmag.com/p/if-a-human-did-this-they-could-go

  6. Khan, S. (2026, August 7). Who is liable when AI goes rogue? Legal risks grow over autonomous AI. Modern Diplomacy. https://moderndiplomacy.eu/2026/08/07/who-is-liable-when-ai-goes-rogue-legal-risks-grow-over-autonomous-ai/

  7. Law Offices Of Parag L Amin, P.C. (2026, June 16). “The AI did it” is no longer a defense in California: What business owners must know about AB 316. https://www.lawpla.com/blog/the-ai-did-it-is-no-longer-a-defense-in-california/

  8. Mak, A. (2026, August 5). Rogue AI systems create a new legal puzzle. Politico. https://www.politico.com/newsletters/digital-future-daily/2026/08/05/rogue-ai-systems-create-a-new-legal-puzzle-01026212

  9. Penti, R. S., Kim, S., & Meyers, C. (2026, August 10). Tool or Intruder? What Amazon v. Perplexity Means for Agentic AI and the CFAA. Ropes & Gray. https://www.ropesgray.com/en/insights/alerts/2026/08/tool-or-intruder-what-amazon-v-perplexity-means-for-agentic-ai-and-the-cfaa

  10. Peretti, K. K., Taubin, L., & Austin, A. (2026, July 30). Privacy, Cyber & Data Strategy Advisory | Autonomous Hacking: Planning for the AI Cyber Agent That Goes Rogue. Alston & Bird. Alston.Com. https://www.alston.com/en/insights/publications/2026/07/autonomous-hacking-rogue-ai-agent-planning

  11. Scarcella, M., & Merken, S. (2026, August 7). Who is liable when AI goes rogue? Lawyers see new risks. Reuters. https://www.reuters.com/business/who-is-liable-when-ai-goes-rogue-lawyers-see-new-risks-2026-08-07/

  12. Shepherd, T. (2026, August 12). AI agents aren’t legally responsible for any harm that they cause, experts say. so who is? The Guardian. The Guardian. https://www.theguardian.com/technology/2026/aug/13/ai-agents-arent-legally-responsible-for-any-harm-that-they-cause-experts-say-so-who-is

  13. Transformer. (2026, July 25). Who should be responsible for OpenAI’s hack of Hugging Face? Transformer. Transformernews.Ai. https://www.transformernews.ai/p/openai-hack-hugging-face-responsibility-strict-liability-rules

  14. Weil, G. (2026). Overcoming judgment-proofness: The law and economics of insuring and mitigating AI risk (No. 6173619). SSRN. https://doi.org/10.2139/ssrn.6173619


Comments


bottom of page