Who decides when AI is too dangerous?
- AI systems are becoming more autonomous and capable, but our understanding of the associated risks and effective controls is still evolving.
- Evidence from researchers and independent evaluations shows that AI can take unexpected actions that were not anticipated by its developers.
As AI becomes more capable and autonomous, who has the evidence, independence, and authority to decide when the risk is no longer acceptable?
I’d like to frame the governance question ‘Who decides when AI is too dangerous?’ as a security question and take you through some of the practicalities of where we are today.
We continue to develop AI systems with greater autonomy, with access to tools and very capable cyber security skills. At the same time, we’re trying to learn how those systems really behave when operating independently, and the controls begin to fail.
As a practitioner, I think security professionals should be forgiven for reaching for trusted playbooks and attempting to reduce privilege, restrict access, isolate systems, and monitor behavior. All in an attempt to limit the blast radius when it goes wrong.
I don’t believe we really understand the risk or required controls well enough. In my view it will take a sufficiently serious incident before this is treated differently.
The risk predictors of doom
There are now daily AI risk predictions making the headlines. From those risks that can be managed with good engineering and security best practice to the other end of the spectrum, with the frontier AI development community in particular, predicting catastrophic outcomes
As cyber security practitioners, we shouldn’t choose between these two extremes, (aside from the reality that this is a very fast-changing space), and instead we should attempt to find supporting evidence for these claims.
Not all evidence is necessarily dramatic. Models are becoming more capable of acting autonomously, using tools, and interacting with external systems as part of a cyber security task. At the same time, researchers release evidence that they’re seeing behaviors they didn’t expect or explicitly authorize.
Whilst I do have many concerns and worries about AI capability exceeding controls, I do not subscribe to the notion that we’re approaching human extinction. This is probably the innate optimist in my nature but dismissing the whole discussion because some predications are scary and extreme isn’t unhelpful.
We can’t ignore the evidence
The UK’s AI Security Institute (AISI) continues to report numerous instances of AI Agents taking unsanctioned actions during cyber evaluations. OpenAI has separately disclosed examples of models creating additional instructions for themselves, attempting to hide mistakes, searching for credentials, sharing files without permission, and communicating through various novel ways that researchers had not foreseen.
None of this demonstrates that AI has become uncontrollable. It does prove that AI can become malevolent, and very capable systems can find ways of achieving objectives that their designers didn’t anticipate (see my article Managing malevolent AI agents).
With the risk of repeating myself and previous articles, for security professionals, that will sound familiar. Our adversaries have always looked for the path the defender didn’t consider. The difference now, is that we’re deliberately building systems capable of finding paths through complex environments themselves. The cyber security profession rarely waits for an attack technique to become catastrophic before taking action and collaborating to take the underlying weakness seriously.
What does ‘too dangerous’ actually mean?
If we saw dangerous activity, could we pause, reflect, and or stop? For example: the ability to autonomously identify and exploit vulnerabilities, to circumvent the controls within a test bed or sandbox, to acquire credentials, and deceive the operator? What about doing what it takes and if necessary, coordinating with another agent, replicating and resisting intervention?
There is no simple answer, and there may never be. Risk in this context depends on the capability, access, autonomy, environment, and consequences. A capability that’s benign inside an isolated research lab could become much more significant when it’s connected to a production system with credentials, code repositories, or our critical infrastructure.
The frontier developers are beginning to define ‘dangerous AI’ for us—hold back those cynical thoughts of IPO and stock valuations driven by fear! —through capability classifications, preparedness frameworks and, responsible scaling policies. So are governments and independent evaluators, although we’ve already seen when those assessments don’t align and these bodies profoundly disagree. I discussed one of these and their novel index ranking system in this article The cyber weapon index, when AI stops assisting the attacker.
Related articles
David Pool explains how to build a cost-efficient AI enterprise by balancing model choice, governance, data, FinOps, and workforce capability.
Building the cost-efficient AI enterprise
September 10, 2026
Why healthcare organizations must prioritize workforce transformation alongside AI adoption to achieve sustainable and effective change.
Before AI transformation comes workforce transformation
September 8, 2026
Forward deployed engineers are emerging as a key AI role, helping organizations bridge the gap between technical capability and business value.
What the rise of the forward deployed engineer tells us about how organizations are building capability
July 27, 2026
Vibe coding is accelerating software creation, but is security governance keeping up? Richard Beck explores the emerging AI risk gap.
Vibe coding and the AI security governance gap
July 9, 2026
The launch of Anthropic’s Fable 5 should have been a landmark moment. Instead, it became a warning shot.
Fable 5 was pulled but should such powerful models be available in the first place?
June 15, 2026
Danny Jessee explains how AI is now a core business capability, and how winning organizations can accelerate adoption, reduce friction, and unlock faster returns.
The AI Skills inflection point: From experiment to enterprise advantage
April 20, 2026
There’s a lot of talk these days about the speed of AI. How fast it’s evolving, and how fast it can accelerate our work… I’m more interested in how fast people must now learn in order to keep up.
AI acceleration demands skills at high velocity. Here's how learning can help organizations keep up
March 25, 2026
In The World Economic Forum’s latest scenarios, Four Futures for Jobs in the New Economy: AI and Talent in 2030, they reinforces what we know: businesses will race to automate faster than many workers can currently reskill.
World Economic Forum: AI will displace workers ‘faster than reskilling can respond’
March 16, 2026
The team here at QA is excited to announce our new collaboration with the world’s most valuable company, NVIDIA. They’re at the center of the new generation of technologies, powering the 4th industrial revolution and rightly dominated headlines through 2025.
Who is NVIDIA? And why our new collaboration matters to you
February 9, 2026
Dr Vicky Crockett gives her view of the top 10 AI trends to expect in 2026, including agentic AI, AI in apps and AI experimentation.
What’s next in AI? 10 predictions for 2026
December 15, 2025
About the Author
Richard Beck