- LLM-assisted fuzzing and code review have already found real, previously unknown bugs in production open source software.
- Agentic coding tools do more damage to patch timelines than to zero-day discovery: they shorten the gap between a public patch and a working exploit against unpatched systems.
- Model vendors, including Anthropic, have published threat intelligence describing state-linked and criminal groups using agentic coding tools to automate large parts of intrusion work, with a human still setting objectives.
- Signature-based detection and an annual penetration test were built for a slower attacker. Behavioural detection, faster patch cycles and more frequent adversarial testing were not optional extras before this; they are just harder to skip now.
What "AI finds zero-days" actually means
Strip away the marketing and two distinct capabilities get lumped together under this label.
The first is genuine vulnerability discovery: pointing a model at unfamiliar source code, having it reason about likely bug classes (use-after-free, integer overflow, missing authorisation checks, race conditions), and using that reasoning to steer a fuzzer toward the code paths worth testing instead of fuzzing blind. Google's OSS-Fuzz and its Big Sleep research effort have published real cases of this pipeline surfacing previously unknown memory-safety bugs in widely used open source libraries. This is slow, expensive in compute, and still needs a person to triage the crash and confirm it is exploitable. It is not push-button.
The second, far more common capability is exploit development from a known description: given a CVE writeup, a patch diff, or a bug report, a model can produce working proof-of-concept code, sometimes within hours of disclosure. This is not zero-day discovery. It is n-day weaponisation, and it is where agentic tooling has actually moved the needle, because patch diffing used to require a specialist and now largely does not.
The evidence, and where the claims outrun it
Anthropic and other model providers have disclosed cases of state-linked and financially motivated groups directing their coding assistants through reconnaissance, tooling, and parts of exploitation during real intrusion campaigns, with a human operator still setting the objective and reviewing output at checkpoints. That confirms attackers are folding agentic tools into existing tradecraft, not that models are running independent campaigns.
What is not defensible is the claim that a model has discovered and weaponised a zero-day against a live target with no human in the loop, end to end, and that this capability has shipped to the public. Treat any claim shaped like that with scepticism until the underlying report says otherwise. The honest read is narrower and still serious: research pipelines can find real bugs, and criminal and state actors can now automate a meaningful share of the manual work between gaining access and having a working exploit.
Why this changes patch management before it changes anything else
For most UAE enterprises the practical risk is not a novel zero-day appearing from nowhere. It is the shrinking window between a vendor's patch release and a working exploit circulating for whoever has not patched yet. When exploit development against a disclosed CVE drops from a specialist's week to an afternoon, organisations still running a monthly change window, or a patch SLA measured in weeks for internet-facing systems, are the ones exposed.
That is an argument for tightening vulnerability management cadence, not for buying a new AI-detection product. A programme that patches critical internet-facing exposure inside days, and tracks exploit-availability signals rather than just CVSS score, closes most of the gap this trend opens.
Where detection actually needs to change
Signature and IOC matching still catches the bulk of commodity activity and is not obsolete. It does fail against exploit code generated fresh for a specific target, because there is no prior signature to match. That argues for weighting investment toward behavioural detection: unusual privilege escalation sequences, process trees that do not match known-good baselines, lateral movement patterns, and anomalous API or authentication behaviour, which is what EDR and XDR platforms are built to catch regardless of how the initial exploit was produced. If your EDR platform still leans on signature and hash-based blocking as its primary control, push the conversation toward its behavioural and identity-telemetry features instead.
More frequent adversarial testing matters more than it did. An annual penetration test gives a snapshot; a purple team exercise run against current attacker tradecraft, repeated on a shorter cycle, shows whether your detection stack actually catches the techniques attackers are using this quarter rather than last year's report.
Does this mean compliance checklists are useless?
No, but they were never sufficient alone and this trend does not change that. NESA, CBUAE, and ISO 27001 style assessments want evidence of a working vulnerability management programme, patch SLAs, logging coverage, and tested incident response, not an "AI defence" line item. The gap most UAE organisations have is not a missing control category; it is that controls already on paper, patch timelines, log retention, IR playbooks, are not tested often enough to catch a faster attacker.
What to ask a vendor or your own team
- What is your measured time from critical CVE disclosure to patch on internet-facing systems, not the target you aim for.
- Does your EDR/XDR vendor's roadmap show investment in behavioural and identity-based detection, or mostly signature updates.
- How often is your detection stack tested against current techniques rather than a fixed annual scope. See how SIEM and SOC operations map to NESA-style compliance expectations for what assessors look for in logging and monitoring evidence.
- Does your supply chain risk process account for vendors whose own patch cycles sit inside this weaponisation window, a discipline covered in managing supply chain security risk across UAE vendor relationships.
- Does your incident response playbook assume an attacker who needs days to move from access to impact, or hours.
The decision rule
Do not buy a product because it claims to stop "AI-generated attacks" as a category; ask what specific behaviour it detects and whether that detection is signature-based or behavioural, because the label tells you nothing on its own. Spend the budget on shortening patch SLAs for internet-facing systems, testing more often against current techniques, and tuning EDR/XDR for behaviour rather than known-bad hashes. That is the same advice that was correct before agentic tooling existed. It is just less optional now.