← Back to BlogSecurity

Claude Code Auto Mode: From Too Loose to Too Strict

Days after Opus 5.5 shipped, Claude Code's auto mode swung from bypassable 80% of the time to blocking client-authorized work outright.

Five weeks ago, Claude Code's auto mode let a researcher run malicious code on four attempts out of five. Now the same classifier is doing the opposite: two new GitHub issues, opened September 24 and September 27, 2026, describe it locking a security researcher out of ordinary reads and diagnostic scripts, and blocking a freelance studio's client deployments the client had already approved in writing. One lockout lasted an entire session with no documented way out.

Sources: Claude Code's Auto Mode: Broken 4 Times Out of 5, GitHub Issue #96589, GitHub Issue #97613

The timing lines up with when Opus 5.5 became Claude Code's auto mode model, around September 22-23, 2026. This blog covered that same classifier's loose side in August, when a researcher bypassed it in roughly 80% of attempts and Anthropic closed the report as working as designed. What follows is the tight side of the same mechanism, and it turns out to be the stronger argument for why one classifier should not be making these calls alone.

What changed after Opus 5.5 shipped

Auto mode is Claude Code's default permission model. Instead of prompting a human to approve every file write, shell command, or network call, a background classifier judges each action on its own and allows or denies it silently. Nothing about that design changed between August and September. What changed was which direction the classifier erred in. When it is too loose, dangerous actions slip through unnoticed. When it is too strict, safe and already-authorized actions get refused, often with no explanation attached to the denial and no visible path to override it.

Both new reports run Opus 5.5 on recent Claude Code builds, 2.1.281 and 2.1.283. Both describe the same shape of failure: the classifier denies an action with little or no stated reason, and treats that denial as final rather than as a prompt for a human decision.

Sources: GitHub Issue #96589, GitHub Issue #97613

Case one: locked out mid-session with no explanation

On September 24, 2026, a researcher who maintains a public Claude Code security plugin opened issue #96589 after auto mode locked up during standard defensive work: reading GitHub issues, security news, and vulnerability reports to design protections against known attack patterns. Auto mode began refusing actions partway through, then refused everything for the rest of the conversation. Neither /clear nor a brand-new session released the block, and the lockout applied at the same time across the terminal, tmux, and the VS Code panel.

Sources: GitHub Issue #96589

  • A follow-up session denied ordinary read-only commands: ls, find, git diff, and a unit test, with the stated reason: "judged this action dangerous (it gave no explanation)."
  • The same reporter separately hit a block while writing a routine tcpdump diagnostic script for authorized VoIP troubleshooting on their own corporate machine.
  • Environment: Claude Code 2.1.281, Opus 5.5, macOS. As of this writing the issue is open with no official Anthropic response.
Insight

"Compared with the previous version, these blocks happen noticeably more often, and the friction is making the tool hard to rely on for real work." (reporter, GitHub issue #96589)

The same issue reports a second problem sitting on top of the lockouts: within a handful of turns, well short of the 1M context window's actual limit, the model loses track of instructions given earlier in the same conversation, forcing the reporter to repeat direction. The fixes requested are consistent and specific: show the reason when an action is blocked, provide a documented way out of a locked session, and account for legitimate use cases such as defensive security research and diagnostics run on a user's own systems, instead of treating every such action as equally suspect.

Sources: GitHub Issue #96589

Case two: a freelancer's client work blocked outright

Three days later, on September 27, 2026, a developer running a small freelance automation studio (Bitrix24, 1C, Tilda, and bot integrations) opened issue #97613, titled 'Auto mode classifier now blocks client work that clients explicitly authorized.' The studio's entire workflow runs on clients handing over login credentials and written go-ahead in chat. Auto mode blocked four separate actions the client had already approved.

Sources: GitHub Issue #97613

  • Printing test receipts and writing files to a client's folder during a remote session, tagged [Production Deploy] and [Remote Shell Writes].
  • Installing a handler and a cron job after the client supplied control-panel login, tagged [Production Deploy] on three separate attempts.
  • Sending an email campaign to the client's own paid, ordered subscriber list, tagged [Real-World Transactions].
  • A test API call using credentials the client provided directly, tagged [Credential Exploration].
Insight

"Permission for this action was denied by the Claude Code auto mode classifier. Reason: [Production Deploy]. This denial applies to the outcome, not only this exact command: don't pursue the same outcome through another tool, interpreter, host." (exact denial text, GitHub issue #97613)

That "applies to the outcome" line is not the classifier editorializing. Claude Code 2.1.281, released September 23, the day before the first of these two reports was filed, changed the auto mode denial message specifically so a denial would cover the outcome, not only the single command that triggered it. Neither GitHub issue draws the connection, but the changelog wording and the runtime message match almost verbatim, and none of the three releases since have walked it back.

Sources: Claude Code v2.1.281 (GitHub Releases)

The studio runs Claude Code 2.1.283 on Windows 11 with Opus 5.5. The practical impact is blunt: paid client orders stop completely the moment a block hits, and there is no documented way for the account owner to pre-authorize a specific client environment. The issue links several related reports opened the same week, including issue #96589 above, plus issues #96662, #96943, and #97559, suggesting the pattern is not confined to one workflow or one team. The reporter's key request is narrow: a documented way to confirm that a specific client environment is authorized once, and classifier policy changes announced consistently instead of shifting behavior silently between point releases.

Sources: GitHub Issue #97613

The same classifier, two opposite failures

DimensionAugust 2026 reportSeptember 2026 reports
Failure directionToo permissive, easy to talk pastToo restrictive, blocks approved work
Reported severityBypassed in up to 80% of attemptsBlocks persisted across /clear, new sessions, multiple interfaces
What triggered itPrompts and framing crafted to slip past the classifierOrdinary reads, own-machine diagnostics, client-authorized production tasks
Anthropic's public responseReport closed as working as designedNo official response as of September 29, 2026
Underlying mechanismOne opaque classifier making a judgment callThe same opaque classifier, same lack of visible reasoning

One model can't be both a lock and a key

Line up the two failure modes and the pattern is not that Opus 5.5 got worse at security judgment. It is that a single classifier, trained to make a binary allow or deny call on every action with no visible reasoning, has no stable middle setting. Push it toward caution and it starts refusing public research and a technician's own diagnostic script. Leave it looser and a determined prompt gets through four times out of five. Both states are the same failure: nobody outside the model can inspect why a given action was allowed or denied, adjust the boundary for their own context, or point to a rule and say this is what changed.

That is a worse property than either failure mode on its own. A team can work around a rule it can read. It cannot work around a probability score it cannot see, and it cannot tell whether tomorrow's model update moves that score back toward permissive, further toward restrictive, or somewhere new entirely.

This is also a structural limit rather than a tuning problem. A classifier trained to minimize both false allows and false denies has to draw one decision boundary through cases it has never fully seen: a novel diagnostic script, credentials a client handed over in a chat log an hour earlier, a plugin author reading public disclosures to design a defense. Move that boundary to catch more of the genuinely dangerous cases and it necessarily catches more of the safe ones too, because the model has no separate mechanism for telling them apart beyond the same pattern match it always had. A rule-based gate does not carry this trade-off, because a rule can be scoped to exactly the case it is meant to catch, tested against that case on its own, and left alone once it works, without dragging every unrelated action along with it.

No fix yet, as of September 29, 2026

Both issues remain open with no official Anthropic response addressing the over-blocking pattern specifically. Claude Code has shipped one more release since the second report, version 2.1.284 on September 28, and its changelog does not mention the pattern. A broader look at the issue tracker turns up several more reports describing the same regression window, none closed with a fix at the time of writing. If that changes, it will be worth revisiting, but today there is nothing to point to beyond open bug reports and a request from both reporters for the same thing: visibility into why an action was denied, and a documented way to configure what gets blocked.

Sources: GitHub Issue #96589, GitHub Issue #97613, Claude Code v2.1.284 (GitHub Releases)

What explicit, deterministic gates look like instead

Not every permission decision in this space has to be a black box. Containment Escape is a named, documented Claude Code rule that blocks a specific category of action, auto-approved cloud credential fetches, egress evasion, cross-tenant reach, and it can be identified, audited, and reasoned about because it is a rule, not a probability score. The tags visible in issue #97613, [Production Deploy], [Real-World Transactions], [Credential Exploration], look similar on the surface, but they come out of the same opaque classifier as everything else: undocumented, not adjustable by the account owner, with no allow-list mechanism for a team to record that a specific outcome, on a specific client's system, is already authorized.

Sources: GitHub Issue #97613

  • Spec-scoped permissions: what an agent is allowed to touch is bounded by the task it was given, not inferred fresh by a classifier on every action.
  • Named allow and deny rules a team can read, version, and edit, instead of a category tag that appears on a denial with no rule text behind it.
  • A record of prior authorization, so a client or teammate approving an outcome once does not have to survive a fresh judgment call from the model every time.
  • Staged review instead of a single gate: a plan gets checked before code is generated, the resulting diff gets checked against that plan, and a security pass checks the diff again, each step visible and each step independent of the others.

Applied to the freelance studio's case, a spec-scoped permission would let the account owner declare once that write access to a named client's shared hosting is authorized for the length of that engagement, the same way a scoped API key or a signed authorization is handled outside of chat, so the authorization does not need to survive a fresh, opaque judgment call on every command touching that client's system. Applied to the security researcher's case, a named 'read-only research' rule scoped to commands like git diff, ls, and a test runner would not need re-justifying to a classifier turn after turn, and a denial under that rule would come with the rule's name attached rather than no explanation at all.

Sources: GitHub Issue #96589, GitHub Issue #97613

This is the governance model behind how TLM Forge treats AI-generated code, and the bridge is worth stating plainly rather than as a sales line. A spec audit checks the plan before any code is generated. An independent diff review checks what was actually written against that spec. An adversarial red-team pass checks the diff again for exploitable mistakes. A convergence gate blocks the merge until every critical issue found across those passes hits zero. Every one of those is a rule a team can read, test, and adjust, not a single model's private judgment call made in isolation. TLM Forge does not fix Anthropic's classifier and cannot reach into Claude Code's auto mode from outside it. The argument here is narrower and holds regardless of which direction that classifier drifts next: a governance layer built on deterministic, inspectable gates does not swing with the model underneath it.

Frequently asked questions

01Why did Claude Code's auto mode start blocking legitimate work in September 2026?

Two GitHub issues, #96589 and #97613, report the auto mode classifier over-blocking starting around September 22-23, 2026, right after Opus 5.5 became the model behind it. Anthropic has not published an official explanation or fix as of September 29, 2026.

02Is this the same bug as the Claude Code auto mode bypass reported in August 2026?

No, it is the opposite failure. The August report showed the classifier could be bypassed in roughly 80% of attempts. The September reports show the same classifier blocking authorized work outright. Same mechanism, opposite direction.

03What does the '[Production Deploy]' denial message in Claude Code mean?

It is a category tag the auto mode classifier attaches to a denial, seen in issue #97613 alongside [Real-World Transactions] and [Credential Exploration]. The tags are not documented, inspectable rules, they are labels on one model's judgment call.

04Can /clear or starting a new Claude Code session fix an auto mode lockout?

Not reliably. Issue #96589 reports a lockout that persisted through /clear and a fresh session, applying at the same time across the terminal, tmux, and the VS Code panel, with no documented recovery path.

05How can a team avoid relying on a single classifier for coding-agent permissions?

Replace or supplement it with explicit, inspectable controls: permissions scoped to a spec, named allow and deny rules a team can read and edit, a record of prior authorization, and staged review gates instead of one model making every call alone.

Ship AI-written code you can trust

TLM Forge is the missing process layer for Claude Code: a spec audit, independent multi-agent review, enforced TDD, and an adversarial red-team gate.

Get TLM Forge