Video Intelligence
Trust Analytics Video
Eight models score the same thirty-second clip independently, and the ACE engine weights them into two numbers: how authentic the behaviour looked, and how certain the engine is of that. Both decompose back to the signal that produced them.
How a score is justified
- Cross-model agreement
- Temporal consistency across the clip
- Behavioural entropy against human baselines
- Adversarial robustness under deepfake attack
OutputA Trust Factor 1–10, and a confidence level on that number
- Models producing an independent, modality-specific confidence score on the same clip
8
Models producing an independent, modality-specific confidence score on the same clip
- Trust Factor — how human, congruent and authentic the observed behaviour is
1–10
Trust Factor — how human, congruent and authentic the observed behaviour is
- Confidence — how certain the engine is of the verdict it just reached
0–1
Confidence — how certain the engine is of the verdict it just reached
- Clip length the models are tuned for, so an ordinary capture is enough
30s
Clip length the models are tuned for, so an ordinary capture is enough
How the fusion works
Eight scores in, two numbers out
Each model interprets one modality and returns its own confidence. The ACE engine does not average them — it weights them by how accurate that model is, how clean its signal was on this clip, and whether the other models agree with it.
Eight modality scores
Vision
Face, eye, posture
Audio
Speech and tone
Physiology
Heart rate, SpO₂
Frame integrity
Deepfake artefacts
ACE fusion engine
- Weight by model accuracy
- Weight by signal clarity
- Test inter-modal agreement
- Ensemble re-weighting
- Normalise
Two outputs, and their working
Trust Factor
1–10, graded not binary
Confidence
0–1, certainty of that
Signal breakdown
Which model said what
Anomaly trace
Where in the clip it sat
The weights are not fixed. Ensemble learning and the context of the clip adjust them per analysis, so a noisy audio track lowers the weight on the audio models rather than dragging the whole score down with it.
Where it differs
Six dimensions, measured against the category
This compares against how single-modality detection is typically built, not against any particular vendor's product.
| Dimension | Typical AI systems | FaceOff |
|---|---|---|
| Modality coverage | One or two — usually face, sometimes voice | Eight in parallel: eye, face, voice, biometrics, posture and more |
| Signal alignment | Frame-by-frame analysis | Spatiotemporal, frequency and attention-based patterns across the clip |
| Real-world robustness | Degrades under noise and occlusion | Recovers via GANs, filtering and statistical drift correction |
| Deepfake resilience | Detects a limited set of frame inconsistencies | Detects audio-video desync, gaze inconsistency, emotion mismatch and heartbeat |
| Explainability | A basic probability | Full signal breakdown with anomaly traceability |
| Decision process | End-to-end black box | ACE fusion engine with explainable trust logic |
Why graded, not binary
Real or fake is the wrong question
A binary detector tells you a clip is 87% likely to be fake. It cannot tell you what about the person was off, or whether a reviewer should act on it. The Trust Factor Engine grades the behaviour instead, in bands a human can work with.
- A number a reviewer can act on
- 1–10 maps onto a decision — proceed, review, escalate. A real/fake flag maps onto an argument about the flag.
- Weights that move with the clip
- Ensemble learning and video context adjust the weighting per analysis, so a degraded channel is discounted rather than allowed to poison the result.
- Two numbers, not one
- The Trust Factor says how authentic the behaviour looked. The confidence says how much the engine would stake on that. Conflating them is how black-box scores mislead.
- Traceable back to the signal
- Every score decomposes into which models agreed, which dissented, and where in the clip the anomaly sat — which is what makes it survive being questioned.
- Behaviour, not just artefacts
- Artefact detectors lose to the next generator. Congruence across face, voice, posture and pulse is a far harder thing to fabricate than a clean frame.
- Thirty seconds is enough
- The models are tuned for short clips, so the analysis works on what a real case actually produces rather than requiring a forensic capture.
The rest of the line
More in Video Intelligence
Where it runs
Industries deploying Trust Analytics Video
Each sector page covers the threat model, the controls, and the regulators that apply.
Quantify trust, not just real versus fake.
See an eight-model fusion turn a thirty-second clip into a score you can put in front of a reviewer, with the reasoning still attached to it.
