Command Palette

Search for tools and commands...

Back to Articles
AISecurityDeepfakesForensics

How to Detect AI Deepfakes: A Practical Forensic Guide

Mr. Robot Team
March 12, 2026

Why "it looks real" stopped being good enough

Generative video models stopped being a curiosity roughly two years ago. Face swaps, voice clones and fully synthetic scenes are now produced by open-source pipelines that run on a single consumer graphics card. The practical consequence is that human judgment at normal playback speed is no longer a defense: a well-made 15-second clip of a familiar person saying something they never said is enough to start a rumor that takes a week to kill.

That does not mean detection is hopeless. It means detection moved from the eye to the signal. Synthetic media is produced by a different pipeline than a camera, and that pipeline leaves measurable statistical traces even when the result looks flawless at 1080p.

What still gives synthetic video away

Professional forensic reviewers look for four families of artifacts, roughly in this order of reliability:

  • Boundary blending. When a face mask is composited onto a real body, high-frequency detail does not survive the cut cleanly. The jawline, hairline and collar are where the seam shows first, usually as a faint shimmer during fast motion.
  • Lighting incoherence. Real scenes have one dominant light direction. A synthetic element re-lit to match approximately — rather than exactly — produces specular highlights on skin that disagree with the reflections in the eyes.
  • Temporal instability. Sensor noise is not perfectly smooth, but it is consistently imperfect. Generative models produce frames whose grain level drifts frame to frame, because every frame is generated independently and then denoised.
  • Structural drift. Hands, teeth, earrings, and text on a shirt change shape subtly between frames. Pause on any close-up and step through it frame by frame.

Measuring it instead of guessing it

The Deepfake Detector on this site automates the third item, because it is the only one a machine can measure cheaply across a whole clip. It works like this:

  1. The video is decoded into a hidden <canvas> downscaled to roughly 320px wide. Full-resolution pixel reads on a 4K clip would stall the main thread, and the heuristic does not need that detail.
  2. Frames are sampled at about 2 frames per second and the decoder is seeked silently (muted, so there is no audio burst at every jump).
  3. For each sampled frame it computes the average horizontal neighbour delta — the mean absolute difference between neighbouring pixels across the red, green and blue channels. This is a cheap proxy for spatial detail: real camera footage has a characteristic grain floor; heavily denoised or generated frames sit noticeably smoother.
  4. It then compares that frame's spatial score with the frame-to-frame delta from the previous sample. In genuine footage the two stay correlated. In synthetic footage the relationship is unstable, and the mismatch is accumulated into an anomaly score.
  5. The average score is mapped to a 0-100 confidence figure, with a threshold at 55.

Everything runs in the tab. The file is read with the File API into memory, and no frame, thumbnail, or score is ever sent to a server — which matters when the clip in question is a leak you are trying to verify.

Honest limits — read this before you rely on it

A heuristic like this is a triage tool, not a verdict. It compares a clip against the statistics of ordinary camera footage, so it will flag heavy-handed generative work and stay quiet on a meticulously tuned one. It also reacts to legitimate heavy processing: aggressive denoising, low-bitrate re-encoding, and stabilization all smooth grain and can push a genuine clip up the scale. Files above 100 MB or longer than two minutes are rejected outright rather than silently taking minutes on your machine.

Use it as a first pass on a suspicious clip, then escalate: pull the original file if you can, compare against a known-good reference, and treat a high score as a reason to keep digging — not as proof on its own.

Manual checks worth doing first

Before touching any tool, spend ten seconds on the clip at quarter speed:

  • Look at the hands during a gesture, and at teeth during a laugh.
  • Watch the specular highlight in each eye — they should point in the same direction.
  • Listen for lip-sync drift: synthetic mouths routinely lead or lag the phoneme by a few tens of milliseconds.
  • Check the frame edges for resampling softness where the clip was cropped.

If a clip fails one of these and scores high on the detector, you have two independent signals pointing the same way. That is a much stronger position than either one alone.