The $25.6 million call that looked completely normal
In early 2024, a finance employee at a multinational engineering firm joined what appeared to be a routine video call with the company's CFO and several other senior colleagues. Every participant on that call, except the employee himself, was an AI-generated deepfake — recreated from publicly available video and audio of the real executives. Over the course of the call, he was instructed to make 15 transfers totaling HK$200 million (roughly $25.6 million). He complied because the people on screen looked, sounded, and behaved exactly like the executives he worked with every day.
That case is no longer an outlier. Enterprise security teams now report that CEO-and-CFO-style deepfake fraud attempts target roughly 400 companies daily, and the average deepfake fraud incident now exceeds $500,000 in losses, with large-enterprise incidents averaging $680,000. A high-fidelity voice clone can be built from as little as three seconds of source audio — easily pulled from an earnings call, a conference talk, or a public LinkedIn video.
Why video calls are the new attack surface
Real-time face-swap tools overlay a synthetic face onto a live camera feed with enough fidelity to fool a casual glance, and voice-cloning models can now speak in real time with a cloned voice reacting to a live conversation. Combined, an attacker can run an entire fabricated meeting — impersonating one executive, or several at once — without ever needing to pre-record anything. The tell is rarely obvious on a first watch; it shows up in specific, testable details.
The tells that still work in 2026
- Face-edge shimmer. Face-swap models render a synthetic face as a mask laid over the real one. Where that mask meets hair, glasses, or a moving jawline, you'll often see faint shimmering, warping, or a slightly wrong edge — especially during quick head turns.
- Blink timing. Real-time deepfakes still struggle to blink at a natural, irregular human rate. Watch for blinking that's suspiciously rare, or too perfectly metronomic.
- Lip-sync on hard consonants. Because audio and video are frequently generated by separate models, sounds like "p," "b," and "m" — which require a very specific mouth shape — are where lip-sync drift shows up first.
- The unscripted-request test. This is the single most effective live test available to you. Ask the person to do something they couldn't have prepared for: touch their nose, hold up a specific number of fingers, or repeat a random word you just chose. Live deepfake pipelines lag or visibly glitch reacting to genuinely unplanned input.
The control that actually stops it: out-of-band verification
No visual tell is fully reliable against a well-resourced attacker, which is why every serious 2026 guidance on this converges on the same control: verify high-stakes requests through a channel the call itself didn't provide. If a "CFO" on a video call asks for a wire transfer, call that CFO back on a phone number you already had on file — not one given to you during the call — before acting. A pre-agreed code word that's never spoken in normal conversation, changed periodically, remains one of the few things a real-time deepfake genuinely cannot fake, because it isn't derivable from public audio or video of the real person.
Manufactured urgency is the real red flag
Almost every documented enterprise deepfake fraud case, including the $25.6 million incident, centers on urgency: a request to move money or grant access "before end of day" or "before the board finds out," explicitly discouraging you from taking the time to double-check. Treat urgency paired with a financial or access request as the loudest signal in the room — louder than any single visual glitch.
Building this into how your team actually works
Individually memorizing tells isn't a durable defense — teams that hold up well have a written callback-verification policy for any financial request that arrives over video or voice, regardless of how convincing the caller sounds. Score the call against a consistent checklist in real time, and treat "I can't fully rule this out" as reason enough to hang up and verify by another channel before acting.
Frequently Asked Questions
Voice cloning now requires as little as three seconds of source audio, and real-time face-swap tools are widely available, some free. The barrier to entry has dropped enough that CEO/CFO deepfake fraud attempts now target roughly 400 companies daily rather than being a rare, expensive attack.
Ask the person to do something unscripted — touch their nose, hold up a specific number of fingers, or repeat a word you just picked at random. Real-time deepfake pipelines are built to handle normal conversational movement, not genuinely unplanned requests, and tend to lag or glitch when asked to react to something outside their training pattern.
Visual checks reduce risk but never eliminate it against a well-resourced attacker. The control every 2026 security guidance converges on is out-of-band verification — calling the person back on a phone number you already had on file, not one provided during the call itself — before acting on any financial or access request.
Urgency exists to stop you from taking the time to verify. In the widely reported 2024 case where a finance worker transferred $25.6 million after joining a call with deepfake versions of his CFO and colleagues, the requests were framed as time-sensitive. Treat urgency attached to a money or access request as a bigger red flag than any single visual glitch.
Yes — the Deepfake Video Meeting Detector is a 14-point weighted checklist covering face-edge shimmer, blink timing, lip-sync accuracy, and whether you've verified the meeting through a separate channel. It computes a live 0-100 score as you check items off during or right after the call.