
A student writes an essay entirely on her own, in her second language, using the formal register her teachers spent years training into her. She submits it. An AI detector flags it as machine-generated. She has not touched a chatbot, yet the score says otherwise, and she has no real way to prove a negative.
This is not a rare glitch. It is a documented, measurable pattern, and the research behind it points to a specific and uncomfortable conclusion: the same statistical signals that catch AI writing also catch careful, formal writing by people who learned English as a second language.
The Bias Built Into How Detectors Work
AI detectors do not read for meaning. They score two statistical properties in a passage: perplexity, or how predictable each word choice is, and burstiness, or how much sentence length and rhythm vary from one sentence to the next. Text produced by large language models tends to score low on both measures because the models are built to select likely words in a fairly even rhythm. The problem is that formal, rule-following writing, the kind taught explicitly in ESL and EFL classrooms, produces a strikingly similar signature, not because it was written by a machine, but because it follows the same predictable grammar patterns a second language learner is taught to rely on.
The clearest evidence comes from Liang et al. (2023), a Stanford study published in the journal Patterns. Researchers ran a set of TOEFL essays written by Chinese students through several leading GPT detectors and compared the results against essays written by US-born native English speakers on identical prompts. The detectors misclassified 61.3 percent of the non-native essays as AI-generated, compared to 5.1 percent of the native-speaker essays on the same assignment, more than a tenfold gap produced by the exact same tools scoring the exact same task.
The pattern has not faded with newer models. Reporting from 2026 shows the disparity holding steady enough that a growing number of universities have begun disabling AI detection features entirely, and at least one legal case has already challenged whether a detector’s score can stand as evidence of academic dishonesty on its own. None of this means the underlying models have gotten worse at spotting genuine AI output. It means the statistical shortcuts these tools rely on were never a reliable proxy for authorship in the first place, and that gap simply shows up more sharply for anyone writing outside their first language.
What This Means for Students and Professionals Writing in a Second Language
This bias does not stay confined to classrooms. Anyone writing formally in a second language, a graduate student on a dissertation, a professional drafting a grant proposal, an employee writing a report in their non-native language, runs into the same statistical trap. The irony is that the writers most likely to be flagged are often the ones who worked hardest to get the grammar right, since years spent memorizing sentence structures and formal vocabulary produce writing that, by definition, follows predictable patterns.
Native speakers rarely face this tradeoff. Their informal fluency naturally includes contractions, sentence fragments, and irregular rhythm, the very features that lower a perplexity score and keep a detector satisfied. Non-native writers who were taught to avoid exactly those informal habits end up penalized for following the instructions they were given, and the more diligently a student followed their formal writing instruction, the more likely their work is to score as suspicious.
The situations where this comes up most often include:
- Admissions essays and personal statements written in a learned language
- Graduate research papers and theses following strict academic formatting
- Workplace reports and emails written in a company’s primary business language
- Grant and scholarship applications reviewed under tight academic integrity rules
Protecting Your Work Before You Submit It
Because a false flag can carry real academic or professional consequences, the safest habit is checking formal writing before submission rather than after a problem surfaces. Running a draft through an AI detector ahead of time shows where a passage reads as unusually uniform, giving a writer the chance to naturally vary sentence length and structure before anyone else sees a flagged score.
An AI Humanizer can help at that stage, restructuring rigid, rule-following sentence patterns into varied, natural phrasing while keeping the original argument and evidence exactly as the writer intended. It is worth checking your institution’s specific policy on AI-assisted writing tools first, since confirming that your own original work reads clearly is a different action than using AI to generate the writing itself, and most academic integrity policies draw that line explicitly.
The fix on the institutional side is not asking non-native speakers to write less carefully. It is asking the institutions relying on these tools to treat a single detector score as a starting point for a conversation, not a verdict. Researchers studying this problem consistently recommend confirming any flag across more than one detection tool before taking action, and building in a genuine human review step where a writer can explain their process, a safeguard several universities have already added after facing pushback over wrongly flagged student work. Some institutions have gone further and disabled AI detection scoring altogether for admissions and coursework review, relying instead on interviews, drafts, and in-person writing samples to judge authorship, though that shift is far from universal.
The Bias the Data Reveals
The data is not ambiguous. Non-native English writers are flagged at rates that native speakers rarely experience, using the exact same tools scoring the exact same kind of writing. Until detection technology accounts for that gap, the responsibility falls on students, professionals, and the institutions that rely on these scores to treat a single flag as a question worth asking, not an answer already decided.
For more on how individual detection tools perform on non-native writing and how policies are shifting in response, further reading on the Phrasly blog covers the underlying research in more depth for anyone writing formally in a second language.
Disclaimer: The information provided in this article is for general informational and educational purposes only. It does not constitute professional academic, legal, or educational technology advice. AI detection tools and institutional policies vary widely and are subject to change; readers should verify their own institution’s guidelines before acting on this information. The mention of specific studies, tools, or services is illustrative and does not imply endorsement. The author and publisher disclaim all liability for any academic or professional consequences arising from reliance on this content. Always confirm policy details directly with your school or employer. This article does not guarantee specific detection outcomes or policy responses.
Unlock the door to fearless decision-making—our fearless frameworks help you choose with confidence.






