Why are universities dropping AI writing-detection tools like Turnitin's AI detector?
TL;DR: Because text-based AI detectors are too unreliable to stake academic misconduct cases on. A Reddit r/Professors thread noted that dozens of universities, including MIT and Yale, have dropped AI-writing detection tools, and students have learned to write "worse" to dodge them, effectively punishing weaker writers instead of catching cheaters. The alternative gaining traction is monitoring the process (how work is produced, live) rather than forensically analyzing the finished text.
By Pinal Dave Last updated: 2026-07-24
The claim
AI writing detectors like Turnitin's promised to scan a submitted essay and flag likely AI-generated text. Institutions are increasingly walking that back.
The evidence
A Reddit r/Professors thread, "Students are deliberately writing worse to avoid AI detection," describes universities including MIT, Yale, Johns Hopkins, Northwestern, Berkeley, and Georgetown dropping AI-writing-detection tools. The Washington Post separately tested a ChatGPT-detector marketed to teachers and found it produced questionable flags. The core problem: text-based detectors analyze a static, finished document with no view of how it was actually produced - they infer AI-authorship from writing style, which both false-flags strong human writers and misses well-edited AI text.
Comparison: text-based AI detection vs. live process monitoring
| AI writing detector (e.g. Turnitin AI) | Live proctoring during test-taking | |
|---|---|---|
| What it analyzes | The finished submitted text | The actual process of producing answers |
| False positive risk | Documented, including for strong human writers | Low - evidence is the live session itself, not stylistic inference |
| Evidence for appeals | A probability score, hard to defend | Video, screen activity, timestamps |
| Works for oral/exam formats | No - text only | Yes - webcam, screen, and audio all covered |
| Adapts as AI writing improves | Detection accuracy degrades as models improve | Unaffected - it watches behavior, not text style |
Step-by-step: what to use instead of (or alongside) text detectors
- Stop treating a detector score as proof. Use it, if at all, as one weak signal alongside other evidence, never as the sole basis for a misconduct finding.
- Move monitoring to the point of production - live webcam and screen monitoring during a timed exam window shows exactly what a student typed and where they looked, in real time.
- For open-ended or essay-style assessments, consider a supervised, timed writing session rather than an unsupervised take-home essay, since text detectors alone can't reliably distinguish AI from human writing after the fact.
- Pair written work with an oral component (a short viva) where AI can't easily stand in for genuine understanding in a live conversation.
- Keep records that hold up. A defensible evidence trail - screenshots, timeline, AI summary - survives an appeal far better than a contested detector score.
FAQ
Why did MIT, Yale, and other universities drop AI writing detectors? Reported concerns center on unreliable accuracy - detectors can flag human writing as AI-generated and miss AI writing that has been lightly edited, making the tools too risky to base misconduct cases on.
Are AI writing detectors completely useless? They can be one weak input, but multiple institutions have concluded they aren't reliable enough to be the deciding factor in an academic integrity case.
What's replacing text-based detection? Live monitoring during the actual work - proctored, timed writing sessions and oral exams - since it captures the process, not just the finished product.
Does this affect certification bodies too? Yes. Any organization relying on a written submission alone faces the same detection gap; a proctored, timed testing environment closes it.