Per-PR checklists break down across many repos and agents. At scale, security review needs automated gates enforced on every merge.
Pull requests are merging with zero review, human or automated, 31.3% more often than before AI-assisted coding became routine. That is not a handful of teams cutting corners. It is what happens once an engineering org runs dozens of repositories, hundreds of AI-assisted pull requests a week, and several coding agents opening branches in parallel, and reviewer capacity does not grow to match. The constraint stops being how carefully any one person reviews a diff. It becomes whether review happens at all, and the 2026 data says it increasingly does not.
Where the per-PR checklist assumes the wrong scale
Security review of AI code covers what a reviewer should look for once a diff is open: injection, broken access control, hardcoded secrets, unsafe deserialization, missing input validation, and authorization ordering. That checklist is correct, and running it on every pull request that touches sensitive logic is still the right instinct. It also has a hidden assumption built into it: one reviewer, one diff, and enough time to actually work through the list. That assumption holds for a five-person team shipping one product out of one repository. It stops holding once an org is running AI coding agents across many teams, many repositories, and many pull requests at once, because nothing about the checklist itself changes what any single reviewer can personally get through in a day. This post covers the layer above the checklist: what has to change operationally once review has to happen at that scale, not what to look for once someone finally opens the diff.
What breaks first: the data
Four numbers from 2026 industry data describe the same failure from different angles: code volume is up, review capacity has not followed at the same rate, and teams are absorbing the gap by skipping review rather than by reviewing faster.
- Pull requests merging with no review at all, human or automated, are up 31.3% since AI-assisted coding became routine, a shift researchers attribute to reviewers unable to keep pace with the volume arriving for their attention rather than a deliberate decision to skip oversight.
- Agentic AI pull requests wait 17.6 hours at the 75th percentile before anyone even picks them up for review, against 3.4 hours for unassisted work, roughly a 5.25x gap, based on an analysis of 8.1 million pull requests across 4,800 engineering teams.
- AI-generated pull requests merge within 30 days only 32.7% of the time, against 84.4% for unassisted work, meaning a large share of AI-authored changes sit unresolved well past the point a manual review process was ever designed to handle.
- 62% of security teams say keeping pace with the volume of AI-generated code is getting harder, and 66% spend more than half their working time validating whether findings are real rather than fixing the vulnerabilities underneath them, in a survey of 200 security practitioners at mid-to-large enterprises.
- Only 3.7% of engineering leaders say their existing review process is sufficient to maintain quality and governance as coding agents take on more of the work, and 89% report at least one AI-related production incident already.
Read together, these numbers describe an industry-wide capacity mismatch, not a handful of under-resourced teams. Code volume is rising because AI-assisted engineers and agents produce more pull requests per week than a human-only team ever did, and reviewer headcount is not rising at the same rate in any of these datasets. The result is not slower but still-thorough review. It is review that silently stops happening on a rising share of changes, which is a materially worse outcome for exactly the vulnerability classes a checklist exists to catch, because none of those classes announce themselves in the diff.
What 'at scale' actually means operationally
At scale does not mean the same review, done faster. It means three things changing at once that a single-PR checklist has no mechanism for handling. The number of surfaces to cover multiplies, because it is many repositories instead of one. The number of actors writing code multiplies, because it is many engineers and several autonomous agents, each opening branches independently and on its own schedule. Review capacity does not multiply alongside either of them, because a security or senior-review team grows on a hiring timeline, not a pull-request timeline. A team that adds agents to write code without adding equivalent capacity to review it has not removed a bottleneck. It has moved the bottleneck downstream, to the point where a vulnerable change actually ships, and made that bottleneck invisible until then.
Four ways manual review fails once volume rises
- Consistency drops per repository. A reviewer who runs the checklist carefully on their own team's repo has no way to confirm whether the same checklist ran at all on a sibling team's repo, and nothing forces it to.
- Coverage becomes selective under load. When the review queue backs up, reviewers triage by skimming diffs that look risky and waving through ones that look routine, and 'looks routine' is exactly the camouflage an AI-generated vulnerability wears.
- Bypass becomes the default, not the exception. A pull request that waits long enough gets merged on a stale approval or with no review at all, which lines up with pull requests merging without review rising 31.3% as volume grew, not an isolated lapse.
- Validation time exceeds fix time once findings do surface. Security teams end up spending more time confirming which findings are real than fixing the ones that are, the same ratio ProjectDiscovery's survey found industry-wide.
A gate that can be skipped under deadline pressure is not a gate. It is a suggestion with a compliance checkbox attached, and the data above is what happens once 'skip it this once' quietly becomes the default under volume, not the rare exception a team assumes it still is.
One reviewer versus an enforced gate
| Dimension | Single-PR manual review | Org-wide automated gate |
|---|---|---|
| Consistency across repos | Depends on which engineer reviews and how careful they are that day | Same policy runs the same way on every repository, every time |
| Coverage as PR volume grows | Falls as volume rises; reviewers triage by skimming | Holds constant; the gate runs on every PR regardless of volume |
| Behavior under backlog | PRs wait, then merge on a stale approval or no review | PR stays blocked until the check passes; nothing merges silently |
| Ownership of the bar | Implicit, whoever happened to be assigned as reviewer | Explicit, one policy defined and owned centrally |
| Cost as team or repo count grows | Scales roughly with headcount and repo count | Scales with compute, not with reviewer headcount |
| Typical failure mode | Silent: a vulnerable pattern ships because nobody had time to look | Loud: the merge is blocked and the finding stays visible until fixed |
What has to change: policy owned once, enforced everywhere
Fixing this is structural, not behavioral, and three changes matter more than any individual reviewer's effort. One owner defines the security bar once, instead of leaving it to be re-derived by whichever engineer happens to review a given pull request. The check that enforces that bar runs as a merge gate on every repository the org owns, not as a recommended step a busy team can skip under deadline pressure. And the gate blocks rather than warns, because a warning a human can dismiss under time pressure reduces back to the same manual-diligence problem it was supposed to replace.
Who owns the policy
- A central security or platform team owns the definition of what gets checked: the vulnerability classes, the severity thresholds, and the exceptions process, not each repository's own maintainers deciding it independently.
- Individual teams inherit the bar rather than redefining it locally. An attacker moving laterally only needs the weakest repo's interpretation of 'good enough' to find a path in, so local opt-outs defeat the point of an org-wide gate.
- The gate, not a person, is accountable for whether a pull request merges. That is what makes the policy enforceable across a hundred repositories at once, instead of dependent on a hundred individual reviewers agreeing to apply it the same way.
Rolling out an org-wide gate without stalling every team at once
Turning a checklist into an enforced, org-wide gate is itself a scale problem, and rolling it out badly stalls every team on day one. Three practices keep that rollout from becoming its own outage. Start with the highest-risk repositories, the ones touching authentication, payments, or customer data, rather than flipping the gate on everywhere at once, so a small surface absorbs the first wave of false positives and process friction before it reaches every team simultaneously. Run the gate in warn-only mode for a fixed, short window per repository, not indefinitely, and convert it to a hard block on a published date rather than leaving it as an opt-in that quietly never gets enabled. And track the org-wide numbers with the same rigor as a single diff: the percentage of pull requests merging without the gate running, the time between a pull request opening and the gate's first result, and how many exceptions are currently active and why. A rollout with no expiry date on its exceptions is not a rollout. It is a permanent carve-out with a start date.
Where TLM Forge fits, and where a single scanner doesn't reach
This is the layer TLM Forge is built for. A spec audit catches an underspecified goal before an agent writes against it, which matters more at scale, because a vague spec handed to a dozen parallel agents produces a dozen inconsistent implementations of the same wrong idea instead of one reviewer's mistake on one PR. Independent diff review checks each pull request against its plan without depending on any one engineer having spare bandwidth that week. A red-team pass runs the same adversarial questions security review of AI code describes for a single diff, on every diff, not only the ones that happen to look risky to whoever is triaging the queue that day. The convergence gate is what turns all three into an enforced stop rather than a recommendation: nothing merges while a critical or high-severity finding is open, on any repository running it, regardless of how backed up the review queue gets.
That is a different layer than GitHub's own Code Quality merge gate, which runs CodeQL and AI-authorship detection as a pre-merge check for repositories that turn it on. It catches a real class of issues, and most engineering orgs should enable it. It is scoped to what that one scanner detects, and only on whichever repos have it enabled, not to a spec audit before code exists, an independent review of the actual diff against its plan, or an adversarial pass that reasons about the specific change rather than matching it against a fixed rule set. An org-wide security posture needs both kinds of coverage, not one instead of the other.
Frequently asked questions
01How do you review AI generated code for security at scale?
Not by asking every reviewer to run a longer checklist by hand. Define the vulnerability classes and severity thresholds once, centrally, then enforce that policy as an automated merge gate on every repository, so coverage does not depend on which engineer is free that week or how backed up the review queue is.
02Why does a per-PR security checklist stop working once an org has many repos?
The checklist assumes one reviewer with time for one diff. Once dozens of repositories and several coding agents produce pull requests in parallel, review capacity does not grow with that volume, so the same checklist gets skimmed, skipped, or applied differently from team to team.
03Who should own AI code security policy across an engineering org?
One central team, usually security or platform engineering, not each repository's own maintainers deciding independently. A policy re-derived per team produces a different bar in every repo, and an attacker moving laterally only needs the weakest one, which is why centralizing it matters more as repo count grows.
04Does GitHub's Code Quality gate cover security review at scale?
Partially. It runs CodeQL and AI-authorship detection as a merge gate, which is genuinely useful, but it only covers what that one scanner detects, and only on repositories where a team has turned it on. It does not replace a spec audit, an independent diff review, or an adversarial pass.
05What actually changes operationally when review moves from one PR to org-wide?
The unit of ownership shifts from an individual reviewer to a written policy, and the unit of enforcement shifts from a suggestion to a merge block. Consistency, coverage, and the power to actually stop a merge all have to come from something that does not depend on any one person's bandwidth that day.