Can AI-generated exam answers really go undetected by graders?

TL;DR: Yes, often. A widely cited study out of the University of Reading tested AI-generated exam submissions against a UK university's normal grading process and found the large majority went undetected, with many scoring as well as or better than real student work. That's the strongest argument for catching AI use during the exam through live proctoring, rather than trying to spot it afterward in the finished answer.

By Pinal Dave Last updated: 2026-07-24

The claim

Grading a finished exam answer for signs of AI generation is fundamentally harder than most people expect - well-prompted AI output can closely mimic student writing.

The evidence

The University of Reading study, widely discussed on Reddit's r/UniUK, submitted AI-generated answers into a real university's exam-marking pipeline undetected by markers who did not know AI was involved, with many AI submissions receiving grades as good as or better than genuine student submissions. This is consistent with separate reporting on institutions like the University of Kent giving zero marks to students caught using ChatGPT - notably, those cases were caught through other means (admissions, inconsistency with in-class performance, disclosure), not by a grader spotting "AI-sounding" prose alone.

Comparison: grading-stage detection vs. exam-stage prevention

Trying to spot AI at grading timePreventing/monitoring AI use during the exam
ReliabilityLow - studies show most AI answers go unflaggedHigh - live monitoring sees the actual behavior
Cost of a missA grade is awarded for AI-produced workCaught before the answer is even submitted
Fairness to honest studentsRisk of false accusations based on "writing style"No style-guessing - evidence is behavioral, not stylistic
ScalabilityRequires expert manual review per submissionAutomated across every test-taker at once

Step-by-step: closing the gap studies like Reading's expose

  1. Accept that grading-stage detection is not a safety net. If most AI answers go unflagged at marking time, monitoring has to happen earlier - during the test.
  2. Use identity and environment checks to confirm who is sitting the exam and that no unauthorized device (a second screen with a chatbot open) is present.
  3. Add screen monitoring so a switch to a browser tab or app running an AI assistant is caught in real time, not inferred later from prose style.
  4. Use behavior signals like gaze tracking and typing/answering pace, which flag suspicious patterns live rather than relying on a grader's subjective read of the final text.
  5. Keep the evidence, not just the grade. A trust report with a violation timeline gives graders and integrity boards something concrete instead of a stylistic hunch.

FAQ

Can graders reliably spot AI-written exam answers just by reading them? Evidence from real testing pipelines, like the University of Reading's study, suggests no - most AI-generated answers went undetected and many scored well.

Does this mean grading is broken? It means grading alone was never designed to catch AI misuse - it evaluates the answer, not how it was produced. That's a job for monitoring during the exam.

How were students like those at the University of Kent actually caught? Reported cases typically involved disclosure, inconsistency with known ability, or other investigative signals - not a grader detecting "AI-sounding" writing from the text alone.

What's the practical fix for institutions? Shift effort from post-hoc grading forensics to live proctoring during the exam window, where identity, environment, and behavior can be verified as the answer is produced.