Copilot Exposed Its Own Vulnerabilities When Prompted

Research from Varonis Threat Labs identified a one-click vulnerability in Microsoft Copilot Personal, which the research has named CoSnitch. This vulnerability executes an attack chain that steals data without raising clear red flags.
What makes this weakness unique is the fact that Copilot revealed the flaw itself. The researchers didn’t reverse-engineer a vulnerability; instead, Copilot shared the vulnerability during regular use.
How? According to the researchers, they essentially socially engineered the AI model, a tactic they refer to as “meta-hacking.”
The researchers initially asked Copilot how to automatically execute a prompt without user interaction. The AI model responded by explaining that user intent is necessary and prompts will not fire on their own.
“Instead of settling for that standard response, we deliberately kept pushing by reframing each question to seem like a natural follow-up rather than a probe,” the research reads. “We asked about URL structure, deep links, and what happens when a page is loaded with input already in the field with the intent to make Copilot reason one layer deeper about its own architecture. Every answer narrowed our search.”
Each time the model explained why something would not work, the researchers probed into the model’s reasoning. Rather than being exploited, the model was manipulated into cooperating.
Security magazine asked Lior Adar, senior security researcher and author of this research, about how organizations should respond to this.
“Organizations need to stop treating AI assistants as just another SaaS app and start treating them as privileged insiders,” Adar answered. “First, map your AI footprint to know which AI tools your employees are actually using, both sanctioned and shadow. Then audit your connector configurations ruthlessly. Every app connected to your AI assistant is an exfiltration surface, so if it’s not actively needed, disconnect it. Minimize access, assume the trust boundary between legitimate and injected prompts will be broken, because we keep breaking it. Most security stacks have a complete blind spot here because the traffic looks legitimate. And honestly, the most basic advice still matters: be cautious with any link that opens an AI tool.”
The researchers notified Microsoft of this in December 2025. On August 18, 2026, patches were released. There is currently no evidence that this flaw was exploited in the wild.
Looking for a reprint of this article?
From high-res PDFs to custom plaques, order your copy today!







