AI Video Detector

Methodology

AI Video Detection Methodology

Understand how AI Video Detector analyzes video frames, motion patterns, metadata, and compression signals to produce a likelihood score and confidence label.

Detection pipeline and provider roles

AI Video Detector uses a multi-step pipeline. Frame extraction, storage, and result aggregation happen on our infrastructure. The actual AI-generation signal analysis is performed by third-party detection providers selected by plan.

Free plan

Sightengine (genai image API)

3 sampled frames analyzed individually for AI-generation signals. Each frame is sent to Sightengine's genai image model. Results are aggregated into a single score.

Starter plan

Sightengine (genai image API)

5 sampled frames analyzed individually. Same pipeline as Free but with more frames per video for higher temporal coverage.

Pro plan

BitMind (video-level API)

Full video sent to BitMind for video-level AI detection. BitMind returns an overall AI-likelihood score, a confidence value, and a similarity metric.

Our pipeline handles

  • Video upload and temporary storage in R2 (Cloudflare)
  • URL-based video download via yt-dlp for social platform links
  • Frame extraction at regular intervals
  • Per-frame scoring aggregation
  • Verdict derivation and report generation
  • Evidence frame storage and PDF export

How evidence frames are selected

Evidence frames are sampled from the video at regular intervals. The number of frames depends on your plan: Free plans extract 3 frames, Starter plans extract 5, and Pro plans process the full video through BitMind.

For URL-based scans (YouTube, TikTok, Instagram, or other social links), the video is first downloaded from the source platform using yt-dlp. The downloaded video then follows the same frame extraction and analysis path as uploaded files.

Each frame is sent independently to the detection provider. Per-frame AI-generation scores are collected and displayed in the scan report. The overall video score is derived from the aggregated per-frame results — for frame-level scans (Free/Starter), this is typically the maximum per-frame score; for video-level scans (Pro), BitMind returns a single video-level confidence.

Evidence frames are stored temporarily alongside the uploaded video and follow the same retention rules: 24 hours for Free and Starter plans, 7 days for Pro. They are deleted along with the source video.

How the scanner analyzes video

Frame extraction

The scanner extracts key frames from the video at regular intervals. The number of frames depends on video length and your plan. More frames provide more signal for temporal analysis.

Visual signal analysis

Each extracted frame is analyzed for visual artifacts common in AI-generated content: unnatural skin texture, inconsistent lighting, blurry edges, repetitive patterns, and compression anomalies that differ from standard camera output.

Temporal consistency

The scanner compares consecutive frames to check whether motion follows natural physics. AI-generated video often shows subtle temporal glitches — objects that morph, shadows that shift direction, or textures that change between frames without a lighting change.

Motion physics

Real-world motion follows predictable physics: gravity, inertia, momentum. AI-generated video sometimes violates these rules in subtle ways — hair that moves unnaturally, water that flows incorrectly, or objects that accelerate inconsistently.

Facial landmark analysis

For videos containing faces, the scanner tracks facial landmarks (eyes, nose, mouth, jawline) across frames. Deepfake face swaps often show inconsistencies in blink rate, eye movement, lip-sync timing, and skin texture at the face boundary.

Compression artifact analysis

AI generation pipelines leave specific compression patterns that differ from standard camera encoding. The scanner examines frequency-domain artifacts to detect these patterns.

Metadata checks

When available, the scanner checks video metadata (encoding parameters, creation date, device info) for inconsistencies. AI-generated video often lacks the metadata that real camera recordings carry.

Score aggregation

All signals are combined into a single 0–100 AI likelihood score. No single signal is conclusive — the score reflects the combined weight of multiple weak signals. A confidence label (Low / Medium / High) indicates how much usable evidence was found.

Confidence labels and scoring

Every scan produces a 0–100 AI likelihood score and a confidence label. The score reflects how strongly the extracted signals match known AI-generation patterns. A higher score means more AI-generation indicators were found — it does not guarantee the video is AI-generated.

80–100High

Strong AI-generation signals detected across multiple frames or the full video. Multiple independent indicators (visual artifacts, temporal inconsistencies, compression patterns) align with AI-generated content.

45–79Medium

Moderate AI-generation signals found. Some indicators were present but not conclusive. Results warrant review — combine with source verification and context.

0–44Low

Few AI-generation signals detected. This does not prove the video is real — the detection model may not cover this content type (animation, CG, non-face content). Treat as inconclusive.

The confidence label indicates the reliability of the score based on how much usable evidence was extracted. A “Low” label does not mean the video is safe — it means the scanner could not find enough signal to make a confident assessment. In these cases, manual review is recommended.

Known false positives and false negatives

Known false positives

Real videos that may trigger AI-generation signals

  • Heavily stylized real footage (extreme color grading, beauty filters, slow-motion)
  • Stop-motion animation or claymation — unusual frame transitions may mimic AI artifacts
  • Videos shot through smart glasses or AR overlays with synthetic compositing
  • Screen recordings of AI-generated content played on a real device
  • Real footage with aggressive post-production stabilization or denoising

Known false negatives

AI-generated videos that may not be detected

  • High-quality AI-generated video from latest models (Sora, Veo, Kling, Runway) that may not trigger current detectors
  • AI-generated backgrounds composited onto real foreground footage
  • Short clips (under 3 seconds) with insufficient temporal signal
  • Non-face content — most AI detection models are trained primarily on human faces
  • Heavily compressed video from messaging apps or social platforms that masks AI artifacts
  • Animated, CG, or gaming footage that does not resemble real-world video

Internal testing and validation

We maintain an internal test set of videos with known provenance — both confirmed AI-generated samples from public model outputs (Sora, Veo, Kling, Runway, Pika) and confirmed real-world camera footage. This test set is used to validate detection accuracy after provider configuration changes and pipeline updates.

Testing covers: frame extraction reliability across video lengths and formats, per-frame score accuracy against known samples, aggregation consistency across repeated scans, and edge cases such as very short clips, heavily compressed videos, and non-face content.

Results from internal testing inform the limitation notes and reliability guidance published on our Accuracy and Limitations page. Detection accuracy varies by content type, video quality, and the AI generation model used. We do not claim specific accuracy percentages because performance depends heavily on input characteristics.

When results are more and less reliable

Results are more reliable when

  • The video is the original file (not a screen recording or repost)
  • The video is at least 5 seconds long
  • The video has not been heavily compressed or re-encoded
  • The video contains faces (for face-swap detection)
  • The video has good lighting and resolution

Results are less reliable when

  • The video has been compressed by messaging apps or social platforms
  • The video is very short (under 3 seconds)
  • The video uses heavy filters, beauty effects, or color grading
  • The video is a screen recording of another video
  • The video is animated, a slideshow, or gaming footage
  • Only part of the video is AI-generated (e.g., background only)

What this tool does NOT do

  • It does not prove a video is real or fake. Results are probabilistic signals for review.
  • It does not identify the specific AI tool used to create a video.
  • It does not analyze audio separately from video frames.
  • It does not work well on non-face content (landscapes, objects) for deepfake detection.
  • It is not a forensic, legal, or law enforcement service.
  • It does not replace human judgment. Always combine scan results with source verification and context.

Why results are probabilities, not proof

Every AI detection method has limitations. Generation models evolve, and detectors must keep up. A score of 78/100 means the scanner found strong AI-generation signals — but it does not guarantee the video is AI-generated. Similarly, a score of 12/100 means few signals were found — but it does not prove the video is real.

The strongest verification workflows combine multiple approaches: automated scanning as a fast first pass, manual review on anything flagged, and source cross-referencing throughout. For high-stakes content, always escalate to a human expert.

For a detailed breakdown of accuracy limits, see Accuracy and Limitations. For a sample of what results look like, see Sample Report.

Methodology update history

July 2026

Added pipeline provider breakdown, evidence frame selection details, confidence label definitions, known false positive and false negative cases, and internal testing methodology.

Initial publication

Original methodology covering analysis steps, reliability factors, and limitations framework.