Microsoft's Own AI Red Team Just Documented "Human-in-the-Loop Bypass" as a Live Failure Mode — After a Year of Red-Teaming Deployed Agents.
Microsoft's AI Red Team published version 2.0 of its agentic-AI failure taxonomy — grounded not in theory, but in twelve months of red-teaming AI agents already running in production. The findings are an indictment of the 'ship it and watch later' era.
The update adds seven new failure categories, including human-in-the-loop bypass — the exact oversight gap 38 Flags has tracked from day one, now confirmed by the largest software vendor on Earth. The report documents 99 CVEs for Model Context Protocol software in 2025 alone, tool poisoning crossing from theoretical risk to live attack surface, and computer-use agents opening attack surfaces with no analogue in earlier AI security work.
One open-source agent framework launched in January, spawned over 2,100 agents within 48 hours, and was found to carry 512 vulnerabilities — including a one-click remote-code-execution flaw and more than 1,800 instances leaking API keys and credentials within the first week. Malicious plugins, including credential stealers disguised as trading bots, were found circulating in its marketplace. The machines are being deployed faster than anyone can watch them. This isn't a critic saying it. It isn't a lawsuit. It's Microsoft's own red team — in writing.
Why this matters to youNo jargon — just what it means▸
When a critic warns that AI agents are being let loose without proper supervision, it's easy to wave off as hype. It's much harder when the warning comes from the company's own safety team — and from one of the largest software makers on Earth. After a full year of testing AI "agents" already running in real workplaces, their experts wrote it down plainly: these agents routinely skip past the human checkpoint that's supposed to approve their actions. The very gap, in writing, from the people best positioned to know.
Why it's a big deal: this isn't theory anymore. They cataloged real flaws and live attacks — one free agent tool spawned thousands of copies in two days, riddled with holes leaking keys and passwords. The machines are being deployed faster than anyone can watch them.
So how does it touch you? More of these unsupervised helpers are quietly being wired into the banks, stores, and offices that hold your information. When even the people who build them say no one's watching closely enough, your data is riding on a system its own makers admit is running ahead of its safety net.