The Safeguards We Trust And the Ones That Don't Exist Yet
- Shavvon Cintron
- 2 days ago
- 4 min read
Siren Warns

AI moved into the enterprise faster than most security teams could write a policy for it. Chatbots draft emails, copilots touch production code, and "AI-powered" is now a checkbox on every vendor's feature list. The tooling is real. The safeguards around it are uneven some solid, some theater, and some simply missing.
Here's an honest look at where we actually stand, across three fronts.
1. AI Inside the Enterprise
Most organizations adopted generative AI tools before they finished writing an acceptable-use policy for them. That gap is where the risk lives.
What's working: Data Loss Prevention rules are increasingly AI-aware, flagging when source code, PHI, or credentials get pasted into a public chatbot. Enterprise-tier AI products generally don't train on customer input, and that distinction matters but only if IT actually procures the enterprise tier instead of letting shadow AI spread through personal accounts.
What's missing: Very few companies have real visibility into what's happening. Recent workforce surveys put unsanctioned AI use at roughly two-thirds to three-quarters of employees, with only a small minority of organizations able to say they have full visibility into that usage or a formal AI security policy in place at all [1][2]. The pattern shows up across industries including inside healthcare, where a large share of clinicians report using unauthorized AI tools despite the compliance stakes [3]. The Samsung engineering team's well-documented 2023 incident, where proprietary source code and internal meeting notes ended up inside a public chatbot, is still the reference case for why "just ban it" doesn't work on its own usage moves further out of sight instead of stopping [1][4].

2. AI in the Hands of Attackers
This is where the safeguard conversation gets uncomfortable, because the defenses were built for human-paced attacks.
Voice cloning now needs a few seconds of audio, not minutes, researchers have demonstrated a usable clone from as little as three seconds of sample audio, often pulled from something as public as a voicemail greeting or a social media clip [5][6]. Deepfake video has cleared a similar bar. In the case that put this on every CISO's radar, a finance employee at engineering firm Arup joined a video call with people who looked and sounded exactly like his CFO and colleagues all of them AI-generated and authorized fifteen transfers totaling roughly $25 million before anyone realized the call itself wasn't real [7][8].
What's working: Email security platforms are training detection models on AI-generated text patterns, and out-of-band verification (calling a known number back, not the one in the email or the meeting invite) is finally getting written into real incident response playbooks instead of just security-awareness slides.
What's missing: Voice and video verification protocols specifically. The Arup case is instructive here standard controls like MFA, endpoint detection, and email filtering barely touched this attack, because it targeted human trust in a live call rather than a system [9]. Most organizations still don't have a defined process for "the CFO called and the voice sounds right now what?"
3. AI Building AI: The Governance Gap
The models themselves are being shipped with safety layers: content filters, refusal training, red-teaming before release. That's genuinely more rigorous than it was a few years ago.
What's working: Major model providers have published frontier safety frameworks and increasingly commission outside firms and academics to red-team their models before release [10][11].
What's missing: Independent verification of those claims. A recent academic scoring of frontier AI providers' safety frameworks found that even when external red teams are involved, the frameworks typically don't establish that those testers are actually independent of the company being evaluated and formal board-level risk oversight was largely absent [11]. Researchers studying the field describe red-teaming as still lacking agreed-upon standards or best practices industry-wide [12]. There's no equivalent yet of a SOC 2 report or a PCI audit that an independent party performs on model behavior. Regulation is catching up the EU AI Act began active enforcement of its transparency and general-purpose-model rules on August 2, 2026, though its toughest high-risk system requirements were pushed back to December 2027 under a revision passed just before that date [13][14].

The Bottom Line
AI safeguards aren't absent they're asymmetric. The technical controls inside well-resourced enterprises are improving quickly. The human-facing defenses against AI-enabled social engineering are lagging behind the attacks. And the audits that would let anyone outside a model provider actually verify a safety claim barely exist yet.
Until that last gap closes, the safest assumption is the oldest one in this field: verify out of band, trust nothing on reputation alone, and treat every "AI-powered" claim defensive or offensive as something to test, not something to take on faith.
Part of the Siren Warns series.
References
Wrivio — "Shadow AI in 2026: What the Numbers Actually Say" — https://www.wrivio.com/blog/shadow-ai-statistics-2026
Salesforce 2026 Workforce AI Survey, via RedTeamPartner — "Shadow AI: 67% of Employees Use AI Tools at Work, Only 18% of Companies Have AI Security Policies" — https://redteampartner.com/blog/shadow-ai-enterprise-risk/
Healthcare Brew survey (Feb. 2026), via Vectra AI — "Shadow AI explained: risks, costs, and enterprise governance" — https://www.vectra.ai/topics/shadow-ai
Technology Radius — "20 Shadow AI Statistics 2024–2026: Enterprise AI Risk" — https://technologyradius.com/statistic/shadow-ai-statistics-2024-2026
McAfee research, via Vectra AI — "AI scams in 2026: how they work and how to detect them" — https://www.vectra.ai/topics/ai-scams
Adaptive Security — "AI Voice Cloning Scams: Detection and Prevention Guide" — https://www.adaptivesecurity.com/blog/the-ultimate-guide-to-ai-voice-cloning-scams-how-to-detect-prevent-and-protect-against-them
CNN Business — "Arup revealed as victim of $25 million deepfake scam involving Hong Kong employee" — https://www.cnn.com/2024/05/16/tech/arup-deepfake-scam-loss-hong-kong-intl-hnk
CFO Dive — "Scammers siphon $25M from engineering firm Arup via AI deepfake 'CFO'" — https://www.cfodive.com/news/scammers-siphon-25m-engineering-firm-arup-deepfake-cfo-ai/716501/
Cyber Helmets — "$25M Deepfake CFO Scam on Video Call" — https://cyberhelmets.com/deepfake-cfo-scam-25m/
OpenAI — "OpenAI's Approach to External Red Teaming for AI Models and Systems" — https://cdn.openai.com/papers/openais-approach-to-external-red-teaming.pdf
Arxiv — "Evaluating AI Providers' Frontier Safety Frameworks" — https://arxiv.org/pdf/2512.01166
Center for Security and Emerging Technology (Georgetown) — "AI Red-Teaming Design: Threat Models and Tools" — https://cset.georgetown.edu/article/ai-red-teaming-design-threat-models-and-tools/
European Commission — "Commission starts enforcing AI Act rules and new transparency requirements on 2 August" — https://digital-strategy.ec.europa.eu/en/news/commission-starts-enforcing-ai-act-rules-and-new-transparency-requirements-2-august
Fello AI — "EU AI Act 2026: Enforcement, Deadlines and Fines" — https://felloai.com/eu-ai-act/


Comments