Who decides when AI is too dangerous?
- AI systems are becoming more autonomous and capable, but our understanding of the associated risks and effective controls is still evolving.
- Evidence from researchers and independent evaluations shows that AI can take unexpected actions that were not anticipated by its developers.
As AI becomes more capable and autonomous, who has the evidence, independence, and authority to decide when the risk is no longer acceptable?
I’d like to frame the governance question ‘Who decides when AI is too dangerous?’ as a security question and take you through some of the practicalities of where we are today.
We continue to develop AI systems with greater autonomy, with access to tools and very capable cyber security skills. At the same time, we’re trying to learn how those systems really behave when operating independently, and the controls begin to fail.
As a practitioner, I think security professionals should be forgiven for reaching for trusted playbooks and attempting to reduce privilege, restrict access, isolate systems, and monitor behaviour. All in an attempt to limit the blast radius when it goes wrong.
I don’t believe we really understand the risk or required controls well enough. In my view it will take a sufficiently serious incident before this is treated differently.
The risk predictors of doom
There are now daily AI risk predictions making the headlines. From those risks that can be managed with good engineering and security best practice to the other end of the spectrum, with the frontier AI development community in particular, predicting catastrophic outcomes
As cyber security practitioners, we shouldn’t choose between these two extremes, (aside from the reality that this is a very fast-changing space), and instead we should attempt to find supporting evidence for these claims.
Not all evidence is necessarily dramatic. Models are becoming more capable of acting autonomously, using tools, and interacting with external systems as part of a cyber security task. At the same time, researchers release evidence that they’re seeing behaviours they didn’t expect or explicitly authorise.
Whilst I do have many concerns and worries about AI capability exceeding controls, I do not subscribe to the notion that we’re approaching human extinction. This is probably the innate optimist in my nature but dismissing the whole discussion because some predications are scary and extreme isn’t unhelpful.
We can’t ignore the evidence
The UK’s AI Security Institute (AISI) continues to report numerous instances of AI Agents taking unsanctioned actions during cyber evaluations. OpenAI has separately disclosed examples of models creating additional instructions for themselves, attempting to hide mistakes, searching for credentials, sharing files without permission, and communicating through various novel ways that researchers had not foreseen.
None of this demonstrates that AI has become uncontrollable. It does prove that AI can become malevolent, and very capable systems can find ways of achieving objectives that their designers didn’t anticipate (see my article Managing malevolent AI agents).
With the risk of repeating myself and previous articles, for security professionals, that will sound familiar. Our adversaries have always looked for the path the defender didn’t consider. The difference now, is that we’re deliberately building systems capable of finding paths through complex environments themselves. The cyber security profession rarely waits for an attack technique to become catastrophic before taking action and collaborating to take the underlying weakness seriously.
What does ‘too dangerous’ actually mean?
If we saw dangerous activity, could we pause, reflect, and or stop? For example: the ability to autonomously identify and exploit vulnerabilities, to circumvent the controls within a test bed or sandbox, to acquire credentials, and deceive the operator? What about doing what it takes and if necessary, coordinating with another agent, replicating and resisting intervention?
There is no simple answer, and there may never be. Risk in this context depends on the capability, access, autonomy, environment, and consequences. A capability that’s benign inside an isolated research lab could become much more significant when it’s connected to a production system with credentials, code repositories, or our critical infrastructure.
The frontier developers are beginning to define ‘dangerous AI’ for us—hold back those cynical thoughts of IPO and stock valuations driven by fear! —through capability classifications, preparedness frameworks and, responsible scaling policies. So are governments and independent evaluators, although we’ve already seen when those assessments don’t align and these bodies profoundly disagree. I discussed one of these and their novel index ranking system in this article The cyber weapon index, when AI stops assisting the attacker.
Related articles
David Pool explains how to build a cost-efficient AI enterprise by balancing model, governance, data, FinOps, and workforce capability.
Building the cost-efficient AI enterprise
10 September 2026
Why healthcare organisations must prioritise workforce transformation alongside AI adoption to achieve sustainable and effective change.
Before AI transformation comes workforce transformation
8 September 2026
Anthropic has announced a research preview of its Model Hardware Standard (MHS), a new specification designed to help AI agents interact with physical devices. If Model Context Protocol (MCP) gave AI a common way to discover and use software tools, MHS does something similar for hardware.
Anthropic's Model Hardware Standard: what happens when AI agents can use physical tools?
28 August 2026
Forward deployed engineers are emerging as a key AI role, helping organisations bridge the gap between technical capability and business value.
What the rise of the forward deployed engineer tells us about how organisations are building capability
27 July 2026
London’s plans to prepare workers for AI-driven change reflect what QA’s analysis of AI readiness across UK sectors has found: adoption is accelerating, but workforce skills, confidence, and capability aren’t keeping pace. By examining how AI is being adopted across industries, we’ve identified a widening gap between giving people access to AI and equipping them to use it effectively.
Closing the AI readiness gap starts with skills
21 July 2026
In uncertain markets, organisations look for confidence signals and this is a clear one. The Government is signalling that skills matter.
What the government's AI skills compact signals for boardrooms
14 July 2026
Vibe coding is accelerating software creation, but is security governance keeping up? Richard Beck explores the emerging AI risk gap.
Vibe coding and the AI security governance gap
9 July 2026
The launch of Anthropic’s Fable 5 should have been a landmark moment. Instead, it became a warning shot.
Fable 5 was pulled but should such powerful models be available in the first place?
15 June 2026
Why enterprise businesses must invest in AI training to unlock productivity, manage risk, and stay competitive in a rapidly evolving economy.
From AI adoption to supercharged progress: the skills challenge facing Britain
12 June 2026
Danny Jessee explains how AI is now a core business capability, and how winning organisations can accelerate adoption, reduce friction, and unlock faster returns.
The AI Skills inflection point: From experiment to enterprise advantage
20 April 2026
About the Author
Richard Beck