Why AI detectors flag human-written text

Why AI Detectors Sometimes Flag Human-Written Text

Quick Takeaways

  • AI detectors have flagged the Declaration of Independence, the Book of Genesis, and the 1836 Texas Declaration of Independence as AI-generated, sometimes at confidence scores above 98%, illustrating that formal, highly structured writing can trigger the same signals detectors associate with AI text.
  • Real students have faced academic misconduct proceedings, failing grades, and in some cases lawsuits after being falsely flagged, including cases at UC Davis, Stanford, and Johns Hopkins.
  • Turnitin has publicly stated that its document-level false-positive rate is less than 1% for documents with 20% or more AI writing, while its sentence-level false-positive rate is around 4%. These figures come from Turnitin’s own testing and have specific definitions and conditions.
  • Non-native English speakers have been shown in independent research to be more likely to be flagged by some AI detectors. Emerging research has also raised questions about how detectors classify some neurodivergent writing styles, although the evidence is more limited. Short, formulaic, or highly formal writing can also create reliability problems depending on the tool and the text being analyzed.
  • No detector, regardless of its advertised accuracy, has a zero false-positive rate, and the people building these tools have generally acknowledged this limitation directly in their own materials.

The problem, illustrated

In October 2024, a data scientist ran the preamble to the Declaration of Independence through several AI detectors as an experiment. One tool, ZeroGPT, rated it 97.93% AI-generated. A separate test the same month put it at 98.51%. The following year, a similar test on the 1836 Texas Declaration of Independence returned an 86.54% AI-generated score from the same tool. Around the same period, passages from the Book of Genesis were also flagged as AI-generated by detection tools. These cases illustrate the same underlying limitation covered in how AI detectors actually work: a detector estimates pattern resemblance, not authorship.

None of these documents could possibly have been written by AI. They predate the technology by centuries. What they share instead is formal structure, elevated vocabulary, and highly regular sentence patterns, the same broad qualities that happen to overlap with what some detectors have learned to associate with machine-generated writing.

This isn’t a one-off curiosity. It’s a visible, easily reproduced symptom of a deeper problem: detectors estimate probability based on pattern resemblance, and several human writing styles resemble those patterns closely enough to get flagged.

What actually causes a false positive

A few overlapping factors show up repeatedly in both research and documented cases.

Simple, predictable sentence structure. Detectors often rely partly on measures like perplexity, which estimates how predictable word choices are. Writing that uses common words in a grammatically regular way, whether because it’s formal, because English isn’t the writer’s first language, or simply because that’s the writer’s natural style, can score lower on this measure, resembling the same signature some AI-generated text produces.

Highly formal or historical writing style. As the Declaration of Independence tests illustrate, formal, structured prose (18th-century legal and political writing, religious texts, academic formal writing) can trigger false positives specifically because detectors are typically trained on contemporary writing samples, not centuries-old formal English.

Neurodivergent writing patterns. Some autistic writers and others with atypical but entirely legitimate writing patterns have reported being flagged at high rates. Emerging research has begun examining this concern, but the evidence remains limited and should not be treated as a universal finding.

Short documents and thin samples. Some detectors perform worse on shorter text, since there’s less content to establish a reliable pattern. Turnitin currently requires at least 300 words of prose in a long-form writing format before it generates an AI Writing Report, reflecting the limitations the company has identified with shorter samples.

The detector’s own training limitations. A detector can only recognize patterns it was trained to recognize. Writing styles, dialects, or content types underrepresented in a detector’s training data are inherently more likely to be misjudged in either direction.

Real cases, not hypotheticals

This isn’t an abstract statistical curiosity. It has had direct consequences for real people.

At UC Davis, a student named Louise Stivers was referred to the university’s judicial affairs office after Turnitin flagged part of a paper as AI-written. She later learned she wasn’t the first student in her class to face this exact situation; a classmate had already been given a failing grade after GPTZero flagged his exam answers. Some students facing this situation have turned to AI humanizer tools defensively, rewriting their own human-written work to avoid future false flags, though as covered in that guide, this treats a symptom rather than the underlying reliability problem.

A Stanford PhD student’s entirely human-written text was flagged as 100% AI-generated by Turnitin. A Johns Hopkins professor separately reviewed a case where Turnitin had flagged over 90% of a student’s paper as AI-written. A University at Buffalo student had final papers flagged in a similar pattern.

These aren’t isolated incidents circulating anecdotally. Multiple students have filed lawsuits against universities over false AI misconduct accusations, and legal trackers following these cases report a mix of outcomes, generally turning less on whether the detector itself was accurate and more on whether the student received a fair process before being penalized.

The math behind “small” error rates

Detector accuracy claims can sound reassuring in isolation. A 96%, 98%, or 99% accuracy figure sounds like a small, tolerable margin of error. The problem is what that margin means once it’s applied at scale.

Turnitin has stated that its document-level false-positive rate is less than 1% for documents where its system detects 20% or more AI writing, while its sentence-level false-positive rate is around 4%. In Turnitin’s definition, the sentence-level figure means that a sentence highlighted as AI-written may actually be human-written in about 4% of cases.

Those figures are not universal error rates for all AI detectors. They are Turnitin’s own measurements under specific conditions. Still, they demonstrate the practical issue: even a relatively small false-positive rate can affect a substantial number of people when a detection system is used at large scale.

This is the core tension in how these tools get used. An error rate that looks small in a vendor’s marketing material can represent a very large number of real, specific people once it’s deployed across a large enough population, and the people it affects rarely learn about the aggregate statistics, only about their own individual case.

Who tends to be affected most

Patterns across multiple studies and reported cases point toward a few groups being disproportionately flagged:

  • Non-native English speakers, whose writing has been shown in research to be more likely to be classified as AI-generated by some detectors. A Stanford-led study found that seven detectors classified 61.22% of a specific sample of TOEFL essays written by non-native English speakers as AI-generated. The result came from a particular dataset and set of detectors, so it should not be generalized to every detector or every non-native English writer.
  • Neurodivergent writers, including some autistic students, whose writing style may differ from typical patterns in ways some detectors interpret as suspicious. Emerging research has begun examining this concern, but the evidence remains limited.
  • Students using common academic phrasing, such as formulaic transition phrases, which some detectors have been shown to weight heavily despite that phrasing being entirely ordinary in academic writing.
  • Freelance writers and journalists, who have reported job losses and disputed work after their writing was misclassified by detection tools used by clients or publishers.

What this means in practice

None of this means AI detectors are useless, but it does mean their output should be treated as one input to consider, not as a standalone verdict. Even the companies making these tools tend to acknowledge this directly. Turnitin’s own materials, for instance, state there’s a risk of false positives and that its AI writing assessment should not be used as the sole basis for adverse action against a student. (Turnitin’s current AI Writing Report guidance).

For anyone evaluating a detector’s flag, a few things are worth weighing beyond the score itself: whether the flagged writing matches the person’s known style and history, whether there’s supporting evidence like drafts or version history, and whether the flagged content overlaps with any of the patterns described above (formal writing, short length, non-native phrasing, atypical style). A flag is a prompt to look closer, not a conclusion on its own. For a broader look at how detectors and the tools built to evade them actually relate to each other, AI humanizer vs. AI detector breaks down that dynamic in more depth.

Frequently asked questions

Can a historical document really be flagged as AI-generated?

Yes. Formal, highly structured historical writing, including the Declaration of Independence and biblical text, has been flagged by multiple detection tools in independently reported tests, since detectors are generally trained on contemporary writing and can misread older formal prose as resembling AI-generated patterns.

Why do detectors flag non-native English speakers more often? 

Research suggests it relates to how some non-native English writers use vocabulary and grammatical structures that can resemble the statistical patterns some detectors associate with AI-generated text, even though the writing is entirely human. A Stanford-led study found substantial false-positive rates on a specific sample of TOEFL essays written by non-native English speakers.

Has any student successfully challenged a false AI accusation? 

Some students have pursued formal complaints or lawsuits after being falsely flagged, with mixed outcomes reported so far. Legal analysis of these cases suggests outcomes tend to depend more on whether a fair review process was followed than on the detector’s accuracy being litigated directly.

Do detector companies admit their tools produce false positives?

Yes, generally. Several detection companies, including Turnitin, state in their own materials that a risk of false positives exists and that responsibility for determining actual misconduct rests with the institution or evaluator, not the tool itself.

What should someone do if they’re falsely flagged? 

Gathering evidence of the writing process, such as saved drafts, version history, and outlines, is generally considered stronger supporting evidence than arguing about the detector’s methodology, since it demonstrates the writing developed over time rather than relying on a rebuttal of the score itself.

Noah William is an SEO strategist and technology editor covering SaaS solutions, enterprise software, and cloud computing at Open Herald. He focuses on data-driven search marketing and practical technology breakdowns.