AI Content Detectors: How They Work and When They Fail

Kendall Chris Kendall Chris Sep 12 / 2 days ago
dot shape
AI Content Detectors: How They Work and When They Fail

 

Most people arrive here for one of two reasons. Either you want to check whether something was written by AI, or your own writing got flagged and you need to understand what happened.

Both are worth answering properly, and answering them properly means starting with something the detector companies will not tell you.

AI detectors are a useful signal and unreliable evidence. They produce false positives often enough to matter, and the false positives are not random. They fall disproportionately on people writing in a second language, on technical and formulaic writing, and on anyone whose prose happens to be plain.

This covers how they work, how to use one sensibly, exactly where they fail, what to do if you are wrongly flagged, and what Google actually thinks about AI-written content.

How to Detect AI Content for Free

Three steps:

  1. Paste the text into a detector
  2. Run the check
  3. Read the score

That is the whole process. What matters is how you read the result, and there are four rules worth following.

Use at least 300 words. Below that, results are close to noise across every detector on the market. A paragraph tells you nothing reliable.

Run more than one. Agreement between two or three detectors is a considerably stronger signal than any single confident number. Disagreement tells you the result is unstable, which is itself useful information.

Read the score as a probability, not a measurement. "85 percent likely AI" does not mean 85 percent of the text was generated. It means the classifier's confidence sits at 85 percent, which is a different claim.

Never use a score on its own to accuse someone. This is the rule that matters most and the one most often broken. More on why in the section on failures.

[PLACEHOLDER] Whether an AI detector exists on your platform, its input length limits, and what the output actually reports. If one exists, this section describes it directly. If not, the article works unchanged as an explainer.

How AI Detectors Actually Work

 

How AI Detectors Actually Work

Two concepts do most of the work, and understanding them explains every failure in the next section.

Perplexity measures how predictable each word is given what came before. Language models select likely continuations, so their output tends to be less surprising than human writing. Low perplexity reads as machine-generated.

Burstiness measures variation in sentence length and structure. Human writing tends to swing between long and short, complex and simple. Model output is more uniform. Low burstiness also reads as machine-generated.

A detector is a classifier trained on large collections of human and AI text, learning which statistical patterns correlate with which origin. Some use additional signals, and the better ones combine several approaches, but perplexity and burstiness remain the foundation.

Here is the consequence nobody states plainly.

The detector never sees where the text came from. It cannot. It measures how predictable and how uniform the writing is, then infers an origin from those measurements.

Which means writing that is naturally plain, formulaic or structurally consistent will score as AI-like regardless of who wrote it or how. That is not a bug in a particular product. It is what the method does.

Where Detectors Fail

 

Where Detectors Fail

Every page selling a detector claims high accuracy. Here is what the evidence actually shows.

False positives against non-native English speakers. The most serious failure and the best documented. A Stanford study published in Patterns found that GPT detectors misclassified writing by non-native English speakers as AI-generated at high rates, while classifying native-speaker writing correctly.

The reason follows directly from the mechanism. Simpler vocabulary and more predictable sentence structure are exactly what a detector reads as machine-like. So the tool systematically penalises people for writing in a second language. The Stanford study on detector bias against non-native English writers sets out the findings in full.

Formulaic human writing. Technical documentation, legal drafting, academic abstracts, structured reports, standardised business correspondence. Writing that is supposed to be uniform reads as uniform, and uniform reads as AI.

Light paraphrasing defeats them. Running generated text through a rewriting tool, or simply editing it by hand, drops detection rates substantially. The tool fails against anyone deliberately evading it, while flagging people who were not trying to evade anything. That is the worst possible combination of error types.

Short samples are unreliable. Under roughly 300 words the signal is too weak to distinguish.

Mixed authorship produces meaningless scores. Human writing tidied with AI assistance, or a generated draft substantially rewritten, sits in a space detectors were not built to classify. The number that comes out is not measuring anything coherent.

Grammar tools can trigger flags. Grammarly's own detector page acknowledges that platforms using generative models for their features can cause content to be flagged as AI. Which means accepting a grammar suggestion can contribute to a positive result.

OpenAI withdrew its own detector. In July 2023, OpenAI discontinued its AI Text Classifier, citing a low rate of accuracy. The original announcement carries the withdrawal note.

That last one deserves a moment. The organisation with the most direct possible knowledge of how its own models write built a detector, evaluated it, and concluded it did not work well enough to keep running. Everyone still selling one is claiming to have solved a problem OpenAI decided it had not.

The practical rule: treat a detector score as one input among several. Never as evidence, and never on its own.

What to Do If You Are Wrongly Flagged

If you are here because your own writing was flagged, this section is for you.

Keep your drafting history. This is the strongest evidence available and it only works if it already exists. Google Docs version history, your editor's revision log, a git history, or simply a folder of dated drafts. Anyone who writes in a context where they might be accused should be building this trail as a matter of habit.

Run the text through other detectors. If three tools give you three different answers, that disagreement is itself evidence that the result is unstable. Screenshot all of them.

Ask what the score actually means. A percentage is a classifier's confidence estimate, not a measurement of how much text was generated. Plenty of people applying these tools do not know that, and asking the question changes the conversation.

Point at the documented limitations. The non-native speaker bias is peer-reviewed and citable. OpenAI's withdrawal of its own classifier is a matter of public record. Both are specific and verifiable rather than a general complaint about fairness.

Ask for a conversation rather than a verdict. Most institutions with a considered policy treat a detection score as a prompt for discussion, not as a finding. If the process you are in does not work that way, it is worth asking whether the policy says it should.

Worth being honest that this is harder outside education. A university usually has an appeals process. A client who has decided your work is AI-generated may not, and the drafting history matters more there than anywhere.

Does Google Penalise AI Content?

This is the question most people in SEO actually want answered, and no detector page covers it, because their audience is educators rather than publishers.

No, and Google has been explicit about this.

Its stated position is that content quality matters, not how content was produced. Using automation to manipulate search rankings violates spam policy. Using automation to help produce genuinely helpful content does not. Google's guidance on AI-generated content sets this out directly.

Three things follow from that:

Google does not run an AI detector on your pages. There is no detection score in any ranking system. Nobody is checking your blog posts against GPTZero.

Thin, unhelpful content ranks badly, as it always has. AI made producing that kind of content much cheaper, which is why there is a great deal more of it, but the problem is the quality rather than the tool that produced it.

Content produced at scale primarily to manipulate rankings is spam regardless of whether a human or a model wrote it. Google's scaled content abuse policy covers exactly this.

The practical implication for a publisher: running your own content through a detector before publishing is largely wasted effort. Nobody is checking, and a clean score does not make thin content useful.

Where it genuinely matters is elsewhere. Client contracts with originality requirements. Academic submissions. Journalism, where disclosure standards apply. Those are real obligations, and they are about the agreement you made rather than about search rankings.

When Detection Is Worth Doing

 

When Detection Is Worth Doing

None of the above means detectors are useless. It means they need using carefully.

Academic integrity, with process around it. A score that triggers a conversation is reasonable. A score that triggers a penalty on its own is not.

Commissioned writing. If you are paying for original work and the agreement specifies human authorship, a detector gives you a signal worth following up. Follow it up rather than acting on it.

Editorial triage at volume. If you screen hundreds of submissions, a detector can order the queue. It should not decide the outcome.

Checking your own writing. The most underrated use and the least discussed. A high score on something you definitely wrote yourself often means the writing is generic rather than that anything is wrong. Uniform sentence length, predictable structure, no distinctive phrasing. Treated that way, a detector becomes a bluntness check rather than an authenticity test.

Wrapping Up

Three things worth carrying away.

A score is a signal, not proof. It measures how predictable writing is, not where it came from, and those are different questions.

The errors are not evenly distributed. Non-native English speakers, technical writers and anyone with a plain style bear most of the false positives. That is a property of the method rather than a flaw in one product.

Never accuse on one number. Run several detectors, ask what the score means, and treat the result as the beginning of a conversation rather than the end of one.

Used that way, detection is genuinely useful. Used as evidence, it produces confident wrong answers about real people.

Frequently Asked Questions

Frequently Asked Questions (FAQs) is a list of common questions and answers provided to quickly address common concerns or inquiries.

Are AI content detectors accurate?

Less accurate than their marketing suggests. They produce meaningful false positive rates, particularly on non-native English writing and formulaic text, and light paraphrasing defeats them.

Can AI detectors be wrong?

Regularly, in both directions. They flag human writing as AI and miss AI writing that has been edited. Treat scores as signals rather than proof.

How do AI detectors work?

They measure how predictable the word choices are and how varied the sentence structure is, then compare those patterns against training data of human and AI text.

Why does my human writing get flagged as AI?

Usually because it is plain, uniform or formulaic. Detectors measure predictability, not origin, so simple clear writing scores as machine-like.

Does Google penalise AI-generated content?

No. Google judges content quality, not production method. Content created at scale to manipulate rankings is spam regardless of who or what wrote it.

Can AI detectors detect paraphrased content?

Poorly. Light rewriting substantially reduces detection rates, which means the tool fails against deliberate evasion while still flagging honest writers.

Can Turnitin detect ChatGPT?

It offers AI detection, though several universities have disabled the feature over false positive concerns. Check whether your institution actually uses it.

How much text does an AI detector need?

At least 300 words for a result worth anything. Shorter samples produce results close to guesswork across every tool.

What should I do if I am falsely accused?

Produce your drafting history, run other detectors to show disagreement, and cite the documented limitations, including OpenAI withdrawing its own classifier.

Is it against the rules to use AI to write content?

It depends entirely on the agreement you are under. Google has no rule against it. Your university, client or publisher may.
Kendall Chris
Written by Kendall Chris Kendall Chris

Kendal is an SEO specialist with 5+ years of experience helping small businesses and freelancers grow their organic traffic. She writes about on-page SEO, content strategy and website optimization at SEO Site Checker.

Share on Social Media: