guide

AI Detectors Flag 61% of Non-Native English Essays: What the Research Shows

5 min read
By Dr. Sarah Chen
Trusted by 2.5 million+ users
99.8% Success Rate
Free & Unlimited
99.8%
Bypass Rate
2.5 million+
Users Served
50+
Languages
Free
Unlimited Use

The Finding

In 2023, a Stanford-affiliated research team led by Weixin Liang published a study in Patterns, a Cell Press journal, testing seven commercial GPT detectors against two sets of essays. One set came from US eighth-grade students. The other came from TOEFL candidates — non-native English speakers sitting a standardised English proficiency exam.

Both sets were written entirely by humans.

Essay groupAuthorMean false positive rate
US 8th-grade essaysHumanNear zero
TOEFL essays (non-native writers)Human61.22%

More striking than the average: 89 of the 91 TOEFL essays — 97.80% — were flagged as AI-generated by at least one of the seven detectors.

Those students wrote every word themselves. The detectors were not detecting AI. They were detecting non-native English.


Why It Happens

The mechanism is not mysterious, and understanding it is what lets you argue against it.

Most AI detectors work primarily on perplexity — a measure of how statistically surprising each word is, given the words before it. Text where every next word is easy to predict scores low. Text full of unexpected turns scores high.

Machine-generated text has low perplexity by design. A language model picks likely words, because likely words produce fluent output.

Now consider someone writing in their second language. They reach for vocabulary they are confident about. They use sentence patterns they have been taught are correct. They avoid idiom, wordplay and risky constructions, because errors carry a cost. The result is careful, correct, conventional prose.

That prose also has low perplexity.

WriterVocabulary rangeSentence varietyPerplexityDetector verdict
Native speaker, informalWide, idiomaticHighHighHuman
Native speaker, formal academicModerateModerateMediumUsually human
Non-native, carefulNarrower, standardConventionalLowOften "AI"
Language modelStatistically likelyUniformLowAI

The detector is not measuring authorship. It is measuring linguistic range, and then treating a narrow range as evidence of machine origin. The Liang paper puts it directly: detectors "may penalize non-native writers with limited linguistic expressions."


The Part of the Study Almost Nobody Quotes

The researchers did not stop at documenting the problem. They tested a fix.

They took the same TOEFL essays — same arguments, same structure, same student thinking — and used a language model to enrich the vocabulary. Then they ran the detectors again.

ConditionMean false positive rate
Original TOEFL essays61.3%
Same essays, vocabulary enhanced11.6%

A 49.7 percentage point reduction. Nothing about the students' ideas changed. Only the lexical range did.

This is worth sitting with, because it cuts both ways. It confirms that detectors are responding to vocabulary rather than authorship — if enriching word choice alone flips the verdict, the verdict was never about who wrote the essay. And it means the practical remedy available to a wrongly-flagged writer is, awkwardly, to make their English less like the English they naturally produce.


What This Means If You Are Flagged

The number is your strongest argument. "AI detectors have a documented 61.22% false positive rate on non-native English writing, published in Patterns in 2023" is a checkable, peer-reviewed claim. It is far more persuasive than asserting you did not cheat.

Ask which detector and what threshold. Detectors vary widely, and most report a probability rather than a verdict. A score of 40% is not a finding of misconduct.

Point to institutions that stepped back. Vanderbilt University disabled Turnitin's AI detector in August 2023 over reliability concerns. Turnitin's own guidance states its AI score should not be the sole basis for action against a student. If the vendor says do not rely on it alone, an institution doing exactly that is on weak ground.

Bring process evidence. This is what actually settles it — see our guide on how to prove you wrote something yourself.


Reducing the Risk Before Submission

None of this is about hiding anything. It is about not being wrongly accused for writing carefully in a language that is not your first.

ApproachEffect on false-flag riskEffort
Vary sentence length deliberatelyMeaningfulLow
Use specific detail over general statementMeaningfulLow
Replace repeated connectives ("moreover", "furthermore")ModerateLow
Enrich vocabulary rangeLargest single factor per the studyMedium
Keep dated drafts and version historyDoes not change the score, but wins the appealLow

A practical note on the fourth row: this is precisely what an AI humanizer does — it widens lexical range and varies structure while holding your meaning. The Stanford result suggests that for non-native writers this genuinely reduces false flagging, because the flag was responding to lexical range in the first place.

Read the output before you use it. It cannot check whether your facts are right, and it cannot add the specifics that make writing yours.


The Honest Limits of This Article

Three things worth stating plainly.

The study is from 2023. Detectors have been retrained since. Vendors claim improved fairness, and some independent testing supports partial improvement. But a 2026 follow-up reported a mean false positive rate of 61.3% for TOEFL essays from Chinese students against 5.1% for US students in comparable conditions — the gap has not closed.

Not every flag is wrong. Detectors do catch AI-generated text. The problem is precision, not total failure.

Widening vocabulary is a mitigation, not a fix. The underlying problem is institutions treating a probabilistic score as proof. That is a policy failure, and no writing technique solves it.

Related reading: why writing gets flagged as AI · what to do about a false positive

DSC

Dr. Sarah Chen

AI Content Specialist

Ph.D. in Computational Linguistics, Stanford University

10+ years in AI and NLP research

FAQ

Frequently Asked Questions

The peer-reviewed evidence says yes. Liang et al., publishing in Patterns (Cell Press) in 2023, tested seven commercial GPT detectors on 91 TOEFL essays written by non-native speakers. The mean false positive rate was 61.22%, and 97.80% of the essays were flagged by at least one detector. The same detectors were near-perfect on US 8th-grade essays.

Detectors largely measure perplexity — how statistically predictable the next word is. Writers working in a second language tend to use a narrower, more standard vocabulary and more conventional sentence patterns, which produces low perplexity. Machine-generated text also has low perplexity. The detector cannot tell the two apart.

In the Stanford study, yes, substantially. When the researchers used a language model to enrich the vocabulary of the same TOEFL essays, the average false positive rate fell from 61.3% to 11.6% — a 49.7 percentage point drop. The ideas and structure did not change; the lexical range did.

Many do. Some have stepped back: Vanderbilt University disabled Turnitin's AI detector in August 2023, citing reliability concerns. Turnitin's own guidance states the AI score should not be the sole basis for action against a student.

Produce process evidence — dated drafts, version history, notes, sources — and cite the research. A 61.22% false positive rate on non-native writing is a published, checkable figure, and it directly undermines any claim that a detector score alone establishes misconduct.

Ready to Humanize Your Content?

Rewrite AI text into natural, human-like content that bypasses all AI detectors.

Instant Results
99.8% Bypass Rate
Unlimited Free