Skip to content
FaceOff Technologies

Video Intelligence

Trust Analytics Video

Eight models score the same thirty-second clip independently, and the ACE engine weights them into two numbers: how authentic the behaviour looked, and how certain the engine is of that. Both decompose back to the signal that produced them.

How a score is justified

  • Cross-model agreement
  • Temporal consistency across the clip
  • Behavioural entropy against human baselines
  • Adversarial robustness under deepfake attack

OutputA Trust Factor 1–10, and a confidence level on that number

Models producing an independent, modality-specific confidence score on the same clip

8

Models producing an independent, modality-specific confidence score on the same clip

Trust Factor — how human, congruent and authentic the observed behaviour is

1–10

Trust Factor — how human, congruent and authentic the observed behaviour is

Confidence — how certain the engine is of the verdict it just reached

0–1

Confidence — how certain the engine is of the verdict it just reached

Clip length the models are tuned for, so an ordinary capture is enough

30s

Clip length the models are tuned for, so an ordinary capture is enough

How the fusion works

Eight scores in, two numbers out

Each model interprets one modality and returns its own confidence. The ACE engine does not average them — it weights them by how accurate that model is, how clean its signal was on this clip, and whether the other models agree with it.

Eight modality scores

  • Vision

    Face, eye, posture

  • Audio

    Speech and tone

  • Physiology

    Heart rate, SpO₂

  • Frame integrity

    Deepfake artefacts

ACE fusion engine

  1. Weight by model accuracy
  2. Weight by signal clarity
  3. Test inter-modal agreement
  4. Ensemble re-weighting
  5. Normalise

Two outputs, and their working

  • Trust Factor

    1–10, graded not binary

  • Confidence

    0–1, certainty of that

  • Signal breakdown

    Which model said what

  • Anomaly trace

    Where in the clip it sat

The weights are not fixed. Ensemble learning and the context of the clip adjust them per analysis, so a noisy audio track lowers the weight on the audio models rather than dragging the whole score down with it.

Where it differs

Six dimensions, measured against the category

This compares against how single-modality detection is typically built, not against any particular vendor's product.

DimensionTypical AI systemsFaceOff
Modality coverageOne or two — usually face, sometimes voiceEight in parallel: eye, face, voice, biometrics, posture and more
Signal alignmentFrame-by-frame analysisSpatiotemporal, frequency and attention-based patterns across the clip
Real-world robustnessDegrades under noise and occlusionRecovers via GANs, filtering and statistical drift correction
Deepfake resilienceDetects a limited set of frame inconsistenciesDetects audio-video desync, gaze inconsistency, emotion mismatch and heartbeat
ExplainabilityA basic probabilityFull signal breakdown with anomaly traceability
Decision processEnd-to-end black boxACE fusion engine with explainable trust logic

Why graded, not binary

Real or fake is the wrong question

A binary detector tells you a clip is 87% likely to be fake. It cannot tell you what about the person was off, or whether a reviewer should act on it. The Trust Factor Engine grades the behaviour instead, in bands a human can work with.

A number a reviewer can act on
1–10 maps onto a decision — proceed, review, escalate. A real/fake flag maps onto an argument about the flag.
Weights that move with the clip
Ensemble learning and video context adjust the weighting per analysis, so a degraded channel is discounted rather than allowed to poison the result.
Two numbers, not one
The Trust Factor says how authentic the behaviour looked. The confidence says how much the engine would stake on that. Conflating them is how black-box scores mislead.
Traceable back to the signal
Every score decomposes into which models agreed, which dissented, and where in the clip the anomaly sat — which is what makes it survive being questioned.
Behaviour, not just artefacts
Artefact detectors lose to the next generator. Congruence across face, voice, posture and pulse is a far harder thing to fabricate than a clean frame.
Thirty seconds is enough
The models are tuned for short clips, so the analysis works on what a real case actually produces rather than requiring a forensic capture.

Where it runs

Industries deploying Trust Analytics Video

Each sector page covers the threat model, the controls, and the regulators that apply.

Quantify trust, not just real versus fake.

See an eight-model fusion turn a thirty-second clip into a score you can put in front of a reviewer, with the reasoning still attached to it.