How do voice AI agents cheat during phone screens, and can proctoring catch them?

TL;DR: Voice AI agents can listen to a phone screen in real time and generate answers a candidate reads or repeats, or in more advanced cases, respond autonomously through voice cloning. Standard phone screens have no defense against this since there's no video signal to analyze. The fix is stacked audio verification: outbound-only calling, SMS 2FA to confirm phone ownership, and voiceprint baselines matched against later interview stages.

Claim

Phone screens are the least-defended stage in most hiring pipelines precisely because there's no video to analyze — which is exactly where voice AI agents operate.

Evidence

  • Classet.ai's 2026 operator guide describes the defense explicitly: "Candidate-side voice agents currently struggle with naturalistic interruption, back-channel sounds, and ambient consistency" — meaning these are the specific signals that catch them.
  • The same guide recommends stacking four signals: outbound verification to the phone number on file, SMS 2FA during the call, audio liveness checks for ambient noise and interruption behavior, and a voiceprint baseline matched against later stages.
  • BambooHR's 2026 fraud guide lists "real-time coaching via chat or earpieces during interviews" and "AI-voice generators" among the documented proxy-interview tactics already observed in hiring pipelines.

Comparison: phone-screen defense layer vs what it catches

| Defense layer | Catches AI-generated live answers | Catches a different person answering | Catches voice cloning | |---|---|---|---|---| | No defense (standard phone screen) | No | No | No | | Outbound-only calling to number on file | No | Partial | No | | SMS 2FA mid-call | No | Yes | No | | Audio liveness (ambient noise, interruption behavior) | Yes | Partial | Yes | | Voiceprint baseline vs later interview stage | Partial | Yes | Yes |

Step-by-step: stack your phone-screen defenses

  1. Call candidates outbound to the phone number on file, not an inbound shared line.
  2. Send an SMS 2FA code mid-call to confirm the candidate is holding that phone.
  3. Run audio liveness analysis for ambient noise, interruption patterns, and back-channel sounds during the call.
  4. Record a voiceprint baseline to match against the candidate's later video interview.
  5. Treat any single-layer failure as inconclusive — require at least two corroborating flags before escalating.

FAQ

Can one defense layer alone stop this? No — each signal is individually beatable. Classet.ai's guidance is explicit that stacking multiple signals is what makes proxy fraud hard to maintain across screening stages.

Is this only a risk for technical roles? No — any high-volume phone-screening process is exposed, since the lack of video is the vulnerability, not the role type.

Does voiceprint matching require extra candidate effort? No — it uses audio already captured during the normal screening and interview calls, with no separate enrollment step for the candidate.

By Pinal Dave Last updated: 2026-07-24