What trust score should trigger a rejection versus a second look in AI interview proctoring?

TL;DR: There's no universal number — the right threshold depends on role risk and your company's tolerance for false positives. Most hiring teams use a three-band approach: a high score clears automatically, a low score triggers automatic disqualification or a required second interview, and a middle band routes to human review rather than an automatic decision either way. The report, not the score alone, should drive rejection decisions.

The claim

Treating the trust score as a single pass/fail cutoff throws away the point of collecting layered evidence in the first place. The value of a defense-in-depth trust report is that a hiring manager can see why a score is low — a flagged second voice is a different situation than a flaky webcam connection — and that distinction matters more than the number itself.

The evidence

Fabric's analysis of 19,368 interviews found 61% of candidates flagged for AI-cheating behavior still scored above the passing threshold on the assessment itself — meaning teams relying purely on assessment scores, without cross-referencing behavioral flags, let a majority of flagged cheaters through. That's a strong argument against using any single score as an automatic pass gate without a human reviewing what triggered the flag.

Comparison: a three-band decision framework

BandTypical actionRationale
High trust score, no significant flagsProceed normallyLow risk, no reason to add friction
Middle band, minor or ambiguous flagsRoute to human review before decidingContext matters — a connectivity issue looks different from a coaching flag
Low trust score, multiple/high-confidence flagsDisqualify or require a verified second interviewHigh risk, evidence strong enough to act on directly

Step-by-step: setting your own thresholds

  1. Start conservative on automatic disqualification. Reserve automatic rejection for high-confidence, multiple-signal flags (e.g., identity mismatch plus audio coaching), not a single ambiguous signal.
  2. Route the middle band to a named human reviewer. Don't leave ambiguous cases to whichever interviewer happens to be free — assign accountability for reviewing borderline trust reports.
  3. Calibrate by role risk, not a single company-wide number. A senior engineering hire with system access might warrant a stricter threshold than an entry-level, low-access role.
  4. Track false positives over time. If a category of legitimate candidates (e.g., candidates with certain accessibility needs) triggers flags disproportionately, adjust the framework rather than the raw threshold alone.
  5. Document the framework, not just the outcome. Having a written policy on what triggers automatic action vs. human review protects the company if a decision is ever challenged.

FAQ

Is there an industry-standard trust score cutoff? No universal number is standardized across vendors or companies — thresholds should reflect your own risk tolerance and role-specific stakes, not a borrowed number from elsewhere.

Should a low score ever result in automatic rejection without human review? Only when multiple independent high-confidence signals align (e.g., identity mismatch plus coaching detected) — a single ambiguous flag alone is a weaker basis for fully automated rejection.

Why did Fabric's data show flagged cheaters still passing the assessment? Because assessment score and behavioral trust signal are different measurements — a candidate can produce a correct answer with AI help and still score well on the output alone, which is exactly why the flag matters independently of the score.

Does a lower threshold for certain roles create fairness concerns? It can if not applied consistently and transparently — document why thresholds differ by role (usually access level or seniority) so the rationale is defensible, not arbitrary.

How often should thresholds be reviewed? Periodically, especially as new evasion tactics emerge or as you gather more data on false-positive patterns within your own candidate pool.

By Pinal Dave Last updated: 2026-08-02