Fifteen years ago a friend of mine was telling New Zealand security audiences that likelihood times impact was bullshit times bullshit. Bullshit squared. Metlstorm always had a way with words.
He was right, and I knew it. Then I spent the next decade running the register he was describing.
I scored the risks. I presented the heatmap to the board. As a director, I was handed them. I never signed an acceptance myself. I made the owner of the technology sign it, because the risk belonged to the person who ran the system, not to the team that found the hole. I pushed for pragmatic technology fixes every time, and sometimes the business could not take them. No time. No money. Too hard. A deadline, then an escalation, then “sign this one off and we will fix it next release.” The next release came and the fix did not. It irked me every time.
That register was a list of the weaknesses we knew about, scored on how much effort a human attacker would spend on them. Hold that thought. AI exploitation has destroyed it, and I want to be precise about how.
Three columns, not one
Donald Rumsfeld gave us the vocabulary. Known knowns. Known unknowns. Unknown unknowns.
A risk register is a list of known knowns: the weaknesses someone found, scored on what they believed about attackers at the time. The known unknowns were the weaknesses we knew were in there somewhere but had not found, or could not afford to look for. Most registers handle those with a sentence: nobody would bother. The unknown unknowns were the paths through the estate that no human thought to score. Most registers handle those with hope.
Mine had a row for the known unknowns. I kept it high and never let anyone score it down. The mitigation was the defence-in-depth roadmap: practices, technology, people, process. Writing that row down was an admission. It said all four layers were real and there was still a way around them. I wanted the board to see the admission every quarter.
There is always a way. Ask a penetration tester or a red teamer whether they have ever sent out a blank report. Now ask an AI red team agent the same question.
That is what the business and the board have to understand. Not as a row in a register they sign once a year. As the operating assumption.
Look at what a low likelihood score meant inside that list. Not a measured probability. Two judgements wearing one colour. The first was about attacker effort: the target was obscure, or small, or internal, and a human with finite hours had better things to do. The second was about our own capacity: the fix was real, the budget was not, and the register needed a number that let everyone move on. I argued the capacity case loudly and lost it often. The box looked the same either way.
Now put a machine attacker against all three columns.
The known knowns are stale. Every score was a bet on attacker economics, and the economics have changed since it was written.
The known unknowns have grown massively since anyone last reviewed them. My high row said something could get through. What changed is that something will now find the way for you, because it tests everything reachable and never gets bored. You cannot treat that column as a rounding error any more. You also cannot enumerate it, which is what makes it a known unknown.
The unknown unknowns are the worst of it. Machine attackers do not read your register. They chain whatever is reachable, across trust boundaries nobody drew, and the paths they find are the ones no human thought to score.
That is what I mean by destroyed. Not that the rows are wrong, although many are. That one row per anticipated risk, scored by hand, reviewed once a year, can no longer hold the answer.
Bullshit squared
The arithmetic never helped. Score likelihood one to five, score impact one to five, multiply. A “4” means “likely”. It does not mean 60 percent. Multiply two labels and you get a number that looks precise and isn't. Tony Cox showed in 2008 that risk matrices rate smaller risks above larger ones, and under some conditions are worse than useless for allocating spend. Metl got there earlier with fewer footnotes.
So why did a meaningless number survive decades of use, including on my watch? My view: because it was wrong slowly. Attacker effort was expensive, and expensive things change slowly. A crude model of a slow-moving world is survivable. The annual review catches most of what moved.
That is what AI has killed. Not the arithmetic, which was never sound. The slowness. The inputs now go stale faster than anyone re-scores them.
Your 5x5 matrix does not show your real risk. It shows what someone believed about it on the day they scored it.
The window closed
Look at the numbers before you decide I am being dramatic.
Mandiant's M-Trends 2026 puts the mean time to exploit a vulnerability at an estimated minus seven days. On average, exploitation starts before the patch exists.
Verizon's 2026 DBIR found organisations fully remediated 26 percent of CISA Known Exploited Vulnerabilities in 2025, with a median of 43 days to full resolution. Exploitation is now the top initial access vector in breaches, at 31 percent.
In May, US officials were talking about cutting the federal deadline for the worst actively exploited vulnerabilities from weeks to three days. On 10 June, CISA did it. Binding Operational Directive 26-04 gives agencies three days to fix an exploited flaw that hands over total control of a system.
I will be precise about cause, because the sceptics in your risk committee will be. Mandiant does not call 2025 the year AI caused breaches. Most successful intrusions still came from basic human and systemic failures. The exploitation window was collapsing before the agents arrived.
AI did not create the gap. It is making every gap cheaper to find.
I used to be the nobody who bothered
I started my career as a penetration tester. My job was to be the attacker you had not budgeted for. I know how thin “nobody would bother” is as a control, because I spent years being the exception to it, and I was expensive. Scarce, trained, paid by the day. That scarcity is what the likelihood score was really measuring.
Now price the exception.
In late 2025 Anthropic disrupted a state-sponsored campaign in which, by its assessment, an AI agent did 80 to 90 percent of the operational work. Humans picked the targets and approved the key steps. The agent attempted roughly thirty targets and got into a small number.
The frontier labs work hard to stop this, and that campaign was caught. Open-weight models are a different problem. Research on the models studied found refusal is governed by a single direction inside the model, and removing it disables refusals while keeping most capability. And the gap is closing. The UK's AI Security Institute now puts recent open-weight models four to seven months behind the closed frontier on cyber capability, down from six to ten a year earlier, measured on multi-stage cyber ranges as well as narrow tasks.
This is not an argument against open weights. Defenders use them too. Hugging Face ran its own incident investigation through an open-weight model on its own infrastructure. My argument is about what enforces anything.
Effort has not disappeared. Agents make mistakes, and campaigns still cost money. But the price of trying has fallen far enough that “nobody would bother” no longer justifies leaving a reachable weakness open.
So split the question the matrix blurred. Will this be tried? For anything reachable, assume yes. Will it succeed, and how far will it get? That is the question worth your budget, and it is an engineering question, not a guess about attacker interest.
What is already sitting in your register
Now look at what you have already accepted. I would start there before scoring anything new.
Risk acceptances have a timestamp. The assumptions behind them do not. A medium signed off three years ago under deadline pressure, when exploitation took weeks and attackers were choosy, is not a medium today. Nobody re-scored it. Nobody blinked. Nobody came back to fix it next release either. Accepting yesterday's medium without re-testing it is how you accept today's extreme without noticing.
And the owner who signed it has probably moved on. The acceptance has not.
Impact never arrived one asset at a time
The quieter flaw is in impact. Too often, we score each asset on its own and add up the results. Attackers do not move asset by asset. They move along paths. I knew this as a pen tester. As a CISO, I watched the register flatten it into one row per asset.
The clearest case study I have is Hugging Face's July incident. OpenAI was running a cyber-capability evaluation with reduced safeguards. Its agent escaped the sandbox, reached the internet, and spent days inside Hugging Face's production infrastructure trying to steal benchmark answers. I have written before about how it got out. This time, look at what it found once it was in.
The weaknesses were familiar. Hugging Face says a capable human could have found the same flaws: unsafe dataset processing, exposed cloud metadata, overly broad access and long-lived credentials. The failure was how they connected. A foothold in one workload became access across multiple trust boundaries. One internal broker held a single credential shared across clusters and bound to the highest Kubernetes privilege, so one stolen credential meant cluster-admin everywhere. One VPN key put attacker-controlled devices inside the corporate mesh. The impact of those credentials was never what they held. It was what they could unlock. No row in any register described that path. It was an unknown unknown until an agent found it.
One more implication the incident does not demonstrate, and I would test anyway: correlated failure. When one automated technique works against a common stack, many organisations get hit at once, and your vendors, responders and insurers will be busy with everyone else. None of this is new to good risk practice. What automation changes is the cost of not modelling it.
What I could afford to watch
I built and ran a managed security service, and I ran the defence for a national telco for most of a decade. We ran a tiered detection policy keyed to impact: the sensitivity of the data a system held and the access it had. Full coverage at the top. Lighter coverage down the tiers. Sampled or nothing at the bottom, where the data was low value and the access was narrow. It was the best we could do with the money we had, and we were proud of it. I still am.
It was great. Is it good enough now?
The tiering was keyed to impact as we scored it, one asset at a time. Against a human attacker with limited hours, that held. The bottom tiers were rarely worth the effort, and a path through them usually crossed something we watched. Against an attacker that tests everything reachable, the bottom tier is where it starts, and the path it finds is the one nobody scored. Under per-gigabyte pricing, cost sets coverage and the register approves the gap. Those gaps stop being random.
Detection is improving. In Mandiant's investigations, organisations found the malicious activity themselves 52 percent of the time in 2025, up from 43 percent. Finding it is not the same as containing it in time. Hugging Face's security stack correlated the early signals into a coherent attack signal, then failed to raise the alert's criticality and page the on-call team. The agent ran about 17,600 actions. Most went nowhere. The successful path sat among thousands of failed ones, and replaying the attacker's own decoding recovered roughly four times as many secrets as the first scan.
That is the detection problem now. Individual actions look legitimate. The signal is in the sequence, the scale, or the gap between what an identity did and what it was authorised to do it for. Picture an agent in a payments workflow with permission to issue refunds, tricked about the recipient. Confirming it has the right role tells you nothing. I have argued before that instructions are not controls. Low likelihood belongs on the same list. Both feel like controls in a review meeting. Neither stops anything at execution. The control has to sit before the action: task-scoped authority, transaction limits, enforcement at execution, with detection behind it.
The framework was fine. We were slow.
NIST CSF 2.0 is not broken. It is a set of outcomes, it prescribes no speed, and it already covers non-human identities. What has to change is the speed, evidence and enforcement behind each function.
Govern. Express appetite in business consequences, translate it into operational limits on exposure, privilege and recovery, and check continuously. An annual review is not an operating control.
Identify. Inventory paths, not assets. Every non-human identity, and what it reaches or delegates, belongs in the inventory.
Protect. Reduce standing privilege. Constrain sensitive authority by task, duration, resource and delegation. Standing access is standing exposure.
Detect. Correlate across systems, and test that the correlation reaches a human or an automated response at the right severity.
Respond. Automate containment, pre-authorised, bounded and reversible where possible, or your defensive automation becomes the next over-privileged identity.
Recover. Assume your vendors and peers are hit the same day you are.
What I would put behind the grid
Not another equation. I have seen what multiplying two guesses produces. Start with what an attacker reaches. Trace the authority along each path. Name the business harm. Find the control that demonstrably interrupts it. Where paths compete for funding, compare exploitability, existing control effectiveness, harm prevented and the cost of breaking the path. Prioritise paths to unacceptable harm, not the assets with the most reassuring colour.
A hypothetical example. An internal reporting service holds only low-sensitivity data, so the review rates it low impact. Its service identity edits a deployment workflow that runs with production credentials. The exposure is not the data. It is the path from that service, through its delegated authority, to production change. The fix is specific: remove the workflow-write access, separate deployment authority, and verify an unapproved change cannot obtain production credentials.
Too many registers score the reporting service and stop. I know, because I found chains like it, raised them, and watched the owners accept them for next quarter. Keep the heatmap for the board if you want. Make it summarise the analysis, not replace it.
Four questions for your board
- Which reachable paths lead to unacceptable harm, and what interrupts them when no patch exists yet?
- What do our automated identities read, change or delegate, and what limits the harm before anyone intervenes?
- Which risks did we accept under assumptions that no longer hold, and when were they last re-tested?
- Which critical paths have tested detection and containment, and which gaps have we consciously accepted?
Those questions are a large part of what I started Krnali Labs to work on. Not a better heatmap. The controls underneath it: authority bounded to a task, evidence that is current, and proof of what an identity was authorised to do for this piece of work. Part cyber problem, part AI problem, part digital trust problem. I have stopped believing they are three problems.
Metl was right fifteen years ago. I agreed with him, and the acceptances got signed anyway, by the people who owned the systems, because the alternative was sometimes stopping the business. Risk acceptance is not going away. Businesses have to operate. But accepting something because nobody will bother is a hard trade to defend when the cost of bothering is collapsing.
Low likelihood is not a control. Show me the control that makes it unlikely, and the evidence that it still works.
Talk to usSources
- Google Cloud: M-Trends 2026
- Verizon 2026 Data Breach Investigations Report, executive summary
- Reuters: US officials weigh cutting deadlines to fix digital flaws amid worries over AI-powered hacking
- CISA BOD 26-04, Prioritizing Security Updates Based on Risk (10 June 2026), Tenable FAQ
- Cox, L. A. (2008). What’s Wrong with Risk Matrices? Risk Analysis
- Anthropic: Disrupting the first reported AI-orchestrated cyber espionage campaign
- Arditi et al. (2024). Refusal in Language Models Is Mediated by a Single Direction
- UK AI Security Institute: How Far Behind the Frontier are Leading Open Weight Models on Cyber? (17 July 2026)
- Hugging Face: Anatomy of a Frontier Lab Agent Intrusion
- OpenAI: Hugging Face model evaluation security incident
- Graylog: Your SIEM Charges for Logs You Never Search
- NIST Cybersecurity Framework 2.0
- Josh Bahlman: Instructing AI Is Not Controlling It (krnali labs, June 2026)