AI Detectors Flag 61% of Non-Native English Essays: What the Research Shows
The Finding
In 2023, a Stanford-affiliated research team led by Weixin Liang published a study in Patterns, a Cell Press journal, testing seven commercial GPT detectors against two sets of essays. One set came from US eighth-grade students. The other came from TOEFL candidates — non-native English speakers sitting a standardised English proficiency exam.
Both sets were written entirely by humans.
| Essay group | Author | Mean false positive rate |
|---|---|---|
| US 8th-grade essays | Human | Near zero |
| TOEFL essays (non-native writers) | Human | 61.22% |
More striking than the average: 89 of the 91 TOEFL essays — 97.80% — were flagged as AI-generated by at least one of the seven detectors.
Those students wrote every word themselves. The detectors were not detecting AI. They were detecting non-native English.
Why It Happens
The mechanism is not mysterious, and understanding it is what lets you argue against it.
Most AI detectors work primarily on perplexity — a measure of how statistically surprising each word is, given the words before it. Text where every next word is easy to predict scores low. Text full of unexpected turns scores high.
Machine-generated text has low perplexity by design. A language model picks likely words, because likely words produce fluent output.
Now consider someone writing in their second language. They reach for vocabulary they are confident about. They use sentence patterns they have been taught are correct. They avoid idiom, wordplay and risky constructions, because errors carry a cost. The result is careful, correct, conventional prose.
That prose also has low perplexity.
| Writer | Vocabulary range | Sentence variety | Perplexity | Detector verdict |
|---|---|---|---|---|
| Native speaker, informal | Wide, idiomatic | High | High | Human |
| Native speaker, formal academic | Moderate | Moderate | Medium | Usually human |
| Non-native, careful | Narrower, standard | Conventional | Low | Often "AI" |
| Language model | Statistically likely | Uniform | Low | AI |
The detector is not measuring authorship. It is measuring linguistic range, and then treating a narrow range as evidence of machine origin. The Liang paper puts it directly: detectors "may penalize non-native writers with limited linguistic expressions."
The Part of the Study Almost Nobody Quotes
The researchers did not stop at documenting the problem. They tested a fix.
They took the same TOEFL essays — same arguments, same structure, same student thinking — and used a language model to enrich the vocabulary. Then they ran the detectors again.
| Condition | Mean false positive rate |
|---|---|
| Original TOEFL essays | 61.3% |
| Same essays, vocabulary enhanced | 11.6% |
A 49.7 percentage point reduction. Nothing about the students' ideas changed. Only the lexical range did.
This is worth sitting with, because it cuts both ways. It confirms that detectors are responding to vocabulary rather than authorship — if enriching word choice alone flips the verdict, the verdict was never about who wrote the essay. And it means the practical remedy available to a wrongly-flagged writer is, awkwardly, to make their English less like the English they naturally produce.
What This Means If You Are Flagged
The number is your strongest argument. "AI detectors have a documented 61.22% false positive rate on non-native English writing, published in Patterns in 2023" is a checkable, peer-reviewed claim. It is far more persuasive than asserting you did not cheat.
Ask which detector and what threshold. Detectors vary widely, and most report a probability rather than a verdict. A score of 40% is not a finding of misconduct.
Point to institutions that stepped back. Vanderbilt University disabled Turnitin's AI detector in August 2023 over reliability concerns. Turnitin's own guidance states its AI score should not be the sole basis for action against a student. If the vendor says do not rely on it alone, an institution doing exactly that is on weak ground.
Bring process evidence. This is what actually settles it — see our guide on how to prove you wrote something yourself.
Reducing the Risk Before Submission
None of this is about hiding anything. It is about not being wrongly accused for writing carefully in a language that is not your first.
| Approach | Effect on false-flag risk | Effort |
|---|---|---|
| Vary sentence length deliberately | Meaningful | Low |
| Use specific detail over general statement | Meaningful | Low |
| Replace repeated connectives ("moreover", "furthermore") | Moderate | Low |
| Enrich vocabulary range | Largest single factor per the study | Medium |
| Keep dated drafts and version history | Does not change the score, but wins the appeal | Low |
A practical note on the fourth row: this is precisely what an AI humanizer does — it widens lexical range and varies structure while holding your meaning. The Stanford result suggests that for non-native writers this genuinely reduces false flagging, because the flag was responding to lexical range in the first place.
Read the output before you use it. It cannot check whether your facts are right, and it cannot add the specifics that make writing yours.
The Honest Limits of This Article
Three things worth stating plainly.
The study is from 2023. Detectors have been retrained since. Vendors claim improved fairness, and some independent testing supports partial improvement. But a 2026 follow-up reported a mean false positive rate of 61.3% for TOEFL essays from Chinese students against 5.1% for US students in comparable conditions — the gap has not closed.
Not every flag is wrong. Detectors do catch AI-generated text. The problem is precision, not total failure.
Widening vocabulary is a mitigation, not a fix. The underlying problem is institutions treating a probabilistic score as proof. That is a policy failure, and no writing technique solves it.
Related reading: why writing gets flagged as AI · what to do about a false positive
Dr. Sarah Chen
AI Content Specialist
Ph.D. in Computational Linguistics, Stanford University
10+ years in AI and NLP research