Frame extraction
The scanner extracts key frames from the video at regular intervals. The number of frames depends on video length and your plan. More frames provide more signal for temporal analysis.
Methodology
Understand how AI Video Detector analyzes video frames, motion patterns, metadata, and compression signals to produce a likelihood score and confidence label.
AI Video Detector uses a multi-step pipeline. Frame extraction, storage, and result aggregation happen on our infrastructure. The actual AI-generation signal analysis is performed by third-party detection providers selected by plan.
Free plan
3 sampled frames analyzed individually for AI-generation signals. Each frame is sent to Sightengine's genai image model. Results are aggregated into a single score.
Starter plan
5 sampled frames analyzed individually. Same pipeline as Free but with more frames per video for higher temporal coverage.
Pro plan
Full video sent to BitMind for video-level AI detection. BitMind returns an overall AI-likelihood score, a confidence value, and a similarity metric.
Evidence frames are sampled from the video at regular intervals. The number of frames depends on your plan: Free plans extract 3 frames, Starter plans extract 5, and Pro plans process the full video through BitMind.
For URL-based scans (YouTube, TikTok, Instagram, or other social links), the video is first downloaded from the source platform using yt-dlp. The downloaded video then follows the same frame extraction and analysis path as uploaded files.
Each frame is sent independently to the detection provider. Per-frame AI-generation scores are collected and displayed in the scan report. The overall video score is derived from the aggregated per-frame results — for frame-level scans (Free/Starter), this is typically the maximum per-frame score; for video-level scans (Pro), BitMind returns a single video-level confidence.
Evidence frames are stored temporarily alongside the uploaded video and follow the same retention rules: 24 hours for Free and Starter plans, 7 days for Pro. They are deleted along with the source video.
The scanner extracts key frames from the video at regular intervals. The number of frames depends on video length and your plan. More frames provide more signal for temporal analysis.
Each extracted frame is analyzed for visual artifacts common in AI-generated content: unnatural skin texture, inconsistent lighting, blurry edges, repetitive patterns, and compression anomalies that differ from standard camera output.
The scanner compares consecutive frames to check whether motion follows natural physics. AI-generated video often shows subtle temporal glitches — objects that morph, shadows that shift direction, or textures that change between frames without a lighting change.
Real-world motion follows predictable physics: gravity, inertia, momentum. AI-generated video sometimes violates these rules in subtle ways — hair that moves unnaturally, water that flows incorrectly, or objects that accelerate inconsistently.
For videos containing faces, the scanner tracks facial landmarks (eyes, nose, mouth, jawline) across frames. Deepfake face swaps often show inconsistencies in blink rate, eye movement, lip-sync timing, and skin texture at the face boundary.
AI generation pipelines leave specific compression patterns that differ from standard camera encoding. The scanner examines frequency-domain artifacts to detect these patterns.
When available, the scanner checks video metadata (encoding parameters, creation date, device info) for inconsistencies. AI-generated video often lacks the metadata that real camera recordings carry.
All signals are combined into a single 0–100 AI likelihood score. No single signal is conclusive — the score reflects the combined weight of multiple weak signals. A confidence label (Low / Medium / High) indicates how much usable evidence was found.
Every scan produces a 0–100 AI likelihood score and a confidence label. The score reflects how strongly the extracted signals match known AI-generation patterns. A higher score means more AI-generation indicators were found — it does not guarantee the video is AI-generated.
Strong AI-generation signals detected across multiple frames or the full video. Multiple independent indicators (visual artifacts, temporal inconsistencies, compression patterns) align with AI-generated content.
Moderate AI-generation signals found. Some indicators were present but not conclusive. Results warrant review — combine with source verification and context.
Few AI-generation signals detected. This does not prove the video is real — the detection model may not cover this content type (animation, CG, non-face content). Treat as inconclusive.
The confidence label indicates the reliability of the score based on how much usable evidence was extracted. A “Low” label does not mean the video is safe — it means the scanner could not find enough signal to make a confident assessment. In these cases, manual review is recommended.
Real videos that may trigger AI-generation signals
AI-generated videos that may not be detected
We maintain an internal test set of videos with known provenance — both confirmed AI-generated samples from public model outputs (Sora, Veo, Kling, Runway, Pika) and confirmed real-world camera footage. This test set is used to validate detection accuracy after provider configuration changes and pipeline updates.
Testing covers: frame extraction reliability across video lengths and formats, per-frame score accuracy against known samples, aggregation consistency across repeated scans, and edge cases such as very short clips, heavily compressed videos, and non-face content.
Results from internal testing inform the limitation notes and reliability guidance published on our Accuracy and Limitations page. Detection accuracy varies by content type, video quality, and the AI generation model used. We do not claim specific accuracy percentages because performance depends heavily on input characteristics.
Every AI detection method has limitations. Generation models evolve, and detectors must keep up. A score of 78/100 means the scanner found strong AI-generation signals — but it does not guarantee the video is AI-generated. Similarly, a score of 12/100 means few signals were found — but it does not prove the video is real.
The strongest verification workflows combine multiple approaches: automated scanning as a fast first pass, manual review on anything flagged, and source cross-referencing throughout. For high-stakes content, always escalate to a human expert.
For a detailed breakdown of accuracy limits, see Accuracy and Limitations. For a sample of what results look like, see Sample Report.
July 2026
Added pipeline provider breakdown, evidence frame selection details, confidence label definitions, known false positive and false negative cases, and internal testing methodology.
Initial publication
Original methodology covering analysis steps, reliability factors, and limitations framework.