← Back to BlogSecurity

Google: AI-Found Flaws Are RCE Half the Time

GTIG: AI-found bugs cause remote code execution nearly 2x more than human-found ones, and attackers weaponize patch diffs within days.

AI-found vulnerabilities turn into remote code execution bugs about 50 percent of the time, nearly double the 26 percent rate for vulnerabilities found by other means, according to a Google Threat Intelligence Group (GTIG) report published October 1, 2026. The same report found that exploitation of vulnerabilities GTIG rates High Risk more than doubled, from 28 incidents across all of 2025 to 75 in the first eight months of 2026, as attackers increasingly use large language models to turn a patch diff into a working exploit within days of disclosure.

Sources: Google Cloud: Vulnerability Discovery and Exploitation Trends in the AI Era, SecurityWeek: Google: AI Is Changing the Pace and Profile of Vulnerability Discovery, Help Net Security: AI-discovered vulnerabilities more likely to lead to remote code execution

That severity shift is layered on top of a volume shift GTIG measured over the same stretch: monthly CVE disclosures roughly doubled, from 5,045 in January 2026 to 10,740 in August. The raw count is not the interesting part. What changed underneath it is the mix. A growing share of what gets disclosed is no longer a low-risk configuration quirk a team can schedule for a quarterly review. It is the category of flaw that hands an attacker code execution, and it is reaching working-exploit stage faster than most remediation pipelines were built to handle.

Sources: Google Cloud: Vulnerability Discovery and Exploitation Trends in the AI Era, SecurityWeek: Google: AI Is Changing the Pace and Profile of Vulnerability Discovery

AI-found bugs skew toward remote code execution

GTIG compared vulnerabilities it could attribute to AI-assisted discovery against everything else disclosed in the same window. Half of the AI-attributed vulnerabilities led to remote code execution. Vulnerabilities found through other methods, manual research, fuzzing, conventional static analysis, reached that outcome only 26 percent of the time. GTIG's own report states the finding in those exact terms: "Exactly 50% of all AI-discovered vulnerabilities result in Remote Code Execution (RCE), compared to just 26% across the broader CVE ecosystem." That is a direct figure from the primary document, not a rounded estimate relayed secondhand through press coverage.

Sources: Google Cloud: Vulnerability Discovery and Exploitation Trends in the AI Era, SecurityWeek: Google: AI Is Changing the Pace and Profile of Vulnerability Discovery, Help Net Security: AI-discovered vulnerabilities more likely to lead to remote code execution

The "likely" qualifier is worth taking seriously rather than skipping past. GTIG is not claiming certainty on every individual attribution; it describes vulnerabilities it identified as likely discovered using AI, a deliberately conservative framing from a report whose author also sells AI-focused security tooling. That caveat does not undercut the finding so much as bound it: the 50-versus-26 split describes a pattern GTIG's own researchers were willing to publish under their name, not a marketing estimate rounded up for effect.

The gap matters because RCE is the category that turns a disclosure into an incident rather than a line item. A cross-site scripting bug or an information-disclosure flaw is a problem to triage and schedule. A remote code execution bug is a problem that ends with someone else running their own code on your infrastructure, and GTIG's data suggests AI-assisted discovery is disproportionately good at finding exactly that category rather than vulnerabilities in general. The tools are not just surfacing more issues. They are surfacing a meaner distribution of issues.

Exploitation is outpacing disclosure

Disclosure volume doubling is one curve. A steeper one sits underneath it: how many of those disclosed vulnerabilities actually get weaponized, and how quickly. GTIG counted 141 distinct vulnerabilities exploited in the wild across the first eight months of 2026, already ahead of the 127 exploited across all of 2025. Within that set, exploitation of vulnerabilities GTIG rates as High Risk went from 28 incidents in 2025 to 75 in roughly two-thirds the time, a jump the researchers tie directly to AI-assisted weaponization rather than a simple increase in attacker headcount.

Sources: Google Cloud: Vulnerability Discovery and Exploitation Trends in the AI Era, SecurityWeek: Google: AI Is Changing the Pace and Profile of Vulnerability Discovery, Help Net Security: AI-discovered vulnerabilities more likely to lead to remote code execution

Only a small fraction of disclosed vulnerabilities, a figure the report puts at roughly 0.23 percent, or about 1 in 431, ever get exploited at all. That detail keeps this from being a panic headline about every CVE in the feed. Attackers are not exploiting more vulnerabilities indiscriminately. They are getting faster and more selective at finding the narrow, high-value subset worth weaponizing, and AI-assisted discovery and triage are what let them cover that much ground without a matching increase in analyst time.

Sources: Google Cloud: Vulnerability Discovery and Exploitation Trends in the AI Era, SecurityWeek: Google: AI Is Changing the Pace and Profile of Vulnerability Discovery

Patch diffs have become exploit blueprints

GTIG describes a specific mechanism behind the speed increase. Threat actors are finding it more efficient to point an LLM at the difference between a patched and an unpatched software version, read that diff alongside the vendor's disclosure announcement and any public proof-of-concept code, and use the combination to rapidly weaponize the underlying n-day vulnerability. The patch, the thing that is supposed to close the hole, doubles as a precise map of exactly where the hole was.

Sources: Google Cloud: Vulnerability Discovery and Exploitation Trends in the AI Era, Help Net Security: AI-discovered vulnerabilities more likely to lead to remote code execution

  • Diff the patched release against the previous version to isolate exactly what code changed
  • Cross-reference the vendor advisory and any public proof-of-concept code for the same CVE
  • Prompt an LLM to reconstruct a working exploit from the delta between those two inputs
  • Fire the exploit at every unpatched instance reachable before defenders finish rolling the patch out

GTIG's case study is CVE-2026-1731, an unauthenticated OS command-injection flaw in BeyondTrust's Privileged Remote Access and Remote Support products that was itself first identified by a third-party AI research agent, before human attackers turned the patch into a weapon. GTIG observed one threat cluster exploiting it within four days of public disclosure, and five additional threat clusters exploiting the same flaw within seven days. A vulnerability that might once have given defenders a week or two of runway between "patch available" and "mass exploitation" now gives them days, and the clusters moving fastest are the ones automating the patch-to-exploit pipeline with the same kind of model their targets use to write code in the first place.

Sources: Google Cloud: Vulnerability Discovery and Exploitation Trends in the AI Era, Help Net Security: AI-discovered vulnerabilities more likely to lead to remote code execution

Insight

Four days from public disclosure to first exploitation is not an outlier anymore. For any vulnerability an LLM can turn into a working proof-of-concept straight from the patch diff, it is close to becoming the baseline.

The patch cadence mismatch

Most organizations still run remediation on a cadence built around a monthly patch cycle: a vendor ships fixes on a fixed schedule, internal teams batch testing and rollout, and a 30-day SLA for high-severity findings gets treated as reasonably aggressive. GTIG's CVE-2026-1731 timeline does not fit inside that cadence. Four days from disclosure to first exploitation is roughly an eighth of a typical 30-day SLA window, and seven days to six separate threat clusters means the vulnerability was already being exploited at scale before most monthly patch cycles would have even scheduled the fix for testing. The cadence was designed for a world where weaponization took about as long as patch rollout. GTIG's numbers describe a world where weaponization finishes first.

The numbers, before and after

Put the volume, severity, and speed findings next to each other and the shape of the problem gets easier to see: more disclosures overall, a disproportionate share of the new ones severe enough for code execution, and a shorter runway between a patch shipping and that patch getting reverse-engineered into an exploit.

MetricBaselineAI era (2026)
Monthly CVE disclosures5,045 (Jan 2026)10,740 (Aug 2026)
RCE rate: human-found vs AI-discoveredHuman-found: 26%AI-discovered: 50%
High Risk vulnerabilities exploited28 (all of 2025)75 (Jan-Aug 2026)
Distinct CVEs exploited in the wild127 (all of 2025)141 (Jan-Aug 2026)
Time to first exploitation (CVE-2026-1731)Historically: weeks4 days to first exploit, 7 days to 6 clusters

Why "we'll patch it next sprint" is now a worse bet

None of this changes what remediation actually requires: triage the finding, build the fix, test it, ship it. What it changes is the price of the gap between those steps. When vulnerability discovery was mostly manual and slower-paced, carrying a two-week remediation backlog was a reasonable risk, because the realistic window before a given flaw got weaponized was usually measured in weeks too. The two clocks moved at roughly the same speed, so a backlog did not have to outrun much.

GTIG's numbers describe a different environment. A disproportionate share of what gets found is RCE-class, and the weaponization clock for at least some of those flaws is now running in days, not weeks, because attackers are pointing the same kind of automation at the patch diff that defenders use to understand it. A backlog sized for the old clock does not get safer because the surrounding threat got louder. A critical finding sitting in a ticket queue for a sprint, waiting for the team that owns that service to get to it, is the exact gap GTIG's data describes as newly expensive.

The operational fix is not more dashboards or more alerts about how bad things are. It is a process that refuses to let a known-critical issue age past the point an automated attacker needs to turn it into working code, which, going by GTIG's example, can be as short as four days. That means remediation gates that block a deploy or a merge while a critical finding is open, not a priority label that gets reviewed at the next planning meeting.

Operationally, zero backlog does not mean fixing everything instantly forever. Medium and low-severity findings can still queue normally. This is specifically about the critical and high-severity tier, the one AI-assisted discovery is disproportionately surfacing in GTIG's data. For that tier it means continuous scanning instead of a periodic sweep, a fix-forward default where a patch ships the same day a critical finding is confirmed rather than waiting for a release train, and a gate, automated where possible, that physically blocks a deploy while the finding is open instead of relying on someone remembering to circle back.

Pro Tip

If a critical or high-severity finding cannot be fixed the same day it is found, treat the delay itself as a tracked risk with an owner and a deadline measured in hours, not a backlog item that waits for capacity. The GTIG numbers argue that the delay is now the bigger exposure, not the original bug.

The uncomfortable part of GTIG's framing is the symmetry. The same automation that lets a security team summarize a patch diff and triage a backlog in minutes is available to whoever is trying to exploit that diff before the triage finishes. Neither side has an inherent speed advantage anymore. The side that wins is whichever one closes the loop from "finding exists" to "finding resolved" faster, and for the attacker, that loop increasingly runs through the same kind of model that could have been used to fix the bug instead of weaponize it.

GTIG's report is about vulnerabilities in deployed software and the patches vendors ship for them, not about bugs an AI coding agent introduces while writing a feature. The underlying math holds either way: a critical-severity finding that waits for the next sprint is a bet that nothing moves fast enough to exploit it in the meantime, and that bet keeps getting worse. TLM Forge applies the same zero-backlog logic to code an AI agent writes before it ever merges, pairing an independent diff review with an adversarial red-teaming pass and a convergence gate that will not let a merge through while any critical or high-severity finding is still open. It does not patch third-party CVEs and it will not shrink GTIG's numbers. What it removes is the backlog teams create for themselves: shipping now and promising to fix the finding later. See how it works for the full review sequence.

The report's loudest number is 10,740 disclosures in a single month, but the number that should move a remediation calendar is smaller: 75, the count of High Risk vulnerabilities GTIG watched get exploited in eight months, up from 28 in all of the year before. Severity went up, the clock got shorter, and most remediation processes were built for neither.

Frequently asked questions

01What did Google's GTIG report find about AI and vulnerability discovery?

GTIG found that vulnerabilities likely discovered using AI lead to remote code execution about 50 percent of the time, versus 26 percent for vulnerabilities found by other methods, while monthly CVE disclosures roughly doubled in 2026, from 5,045 in January to 10,740 in August.

02How much more likely are AI-discovered vulnerabilities to cause remote code execution?

Close to twice as likely. GTIG's October 2026 report found AI-attributed vulnerabilities resulted in remote code execution about 50 percent of the time, compared with 26 percent for vulnerabilities found through manual research, fuzzing, or conventional static analysis.

03How fast are attackers exploiting vulnerabilities after disclosure in 2026?

GTIG documented CVE-2026-1731 being exploited by one threat cluster within four days of public disclosure, with five more clusters exploiting it within seven days, as attackers use LLMs to analyze patch diffs and automate n-day weaponization.

04Did CVE disclosures actually double in 2026?

Yes. GTIG recorded monthly CVE disclosures rising from 5,045 in January 2026 to 10,740 in August, roughly doubling over eight months, alongside a jump in High Risk exploited vulnerabilities from 28 in 2025 to 75 in the same period.

05What should remediation teams change in response to the GTIG findings?

Treat open critical and high-severity findings as a live risk rather than backlog. With AI-found bugs skewing toward RCE and exploitation starting within days of disclosure, a gate that blocks merges or deploys until issues hit zero matters more than it did when discovery was slower and manual.

Ship AI-written code you can trust

TLM Forge is the missing process layer for Claude Code: a spec audit, independent multi-agent review, enforced TDD, and an adversarial red-team gate.

Get TLM Forge