§The journal

Do AI Detectors Agree With Each Other? We Scored the Same 18 Texts Twice and They Disagreed on 7

Two mainstream AI detectors scored the same 18 pieces of text. They reached opposite verdicts on 7 of them, and not once did both agree a text was AI. Every AI verdict came from one detector while the other said human. Full method, all 18 rows, failures included.

Published August 19, 202613 min readBy Abd Shanti
Two identical sheets of printed text side by side on a desk, each marked with a different coloured stamp, one reading pass and one reading fail, with a magnifying glass resting between them.

One detector tells you your writing reads as human. Another tells you the same paragraph reads as AI. Both are sold as reliable, both are used by real institutions to make real decisions about real people, and they cannot both be right. Almost nobody has checked how often this actually happens, because the two groups holding the data have every reason not to look. So we looked, and we published every row.

Why almost nobody publishes this

There are two groups who could run this test easily, and neither one wants the answer in public. Detector companies sell certainty. A study showing a competitor reaching the opposite conclusion on the same paragraph raises the obvious question of which one is wrong, and there is no comfortable version of that answer for either party. So they benchmark against known human text and known AI text, where accuracy can be reported against a label they control, and they leave each other alone.

The other group is companies like ours. A humanizer publishing evidence of a detector catching its output is publishing a list of its own failures. The commercial instinct is to quote one friendly number and move on. We already broke that rule once when we published our real GPTZero pass rate including the 16.7 percent that failed, and the sky did not fall, so we are breaking it again.

The result is a strange gap. Millions of students, writers and freelancers are being judged by these tools every week, and the most obvious question about them, do they agree with each other, has no published number attached to it. Search for it and you find opinion pieces and forum threads. You do not find a measurement.

How the measurement works

The method is deliberately boring, because the finding only means something if there is no room for us to have tilted it.

  • The sample18 pieces of real production text totalling 2,379 words, delivered to real users through the paid pipeline. These were not written for the study. They arrived as ordinary traffic and were selected on length alone, so that a spread of short and medium passages was covered. We did not read them before choosing them and we did not screen them by score.
  • The first scoreGPTZero, through its hosted API, recorded at the moment the text was delivered to the user. This is the same service an institution or a reader would use. It was already stored against each run before this study existed, which means it could not have been shaped by what we were looking for.
  • The second scoreWinston AI, through its hosted content detection API, applied afterwards to the exact stored text. Nothing was regenerated, reworded, trimmed or cleaned up in between. Winston returns a 0 to 100 human score, which we convert to the same 0 to 1 AI scale GPTZero reports so the two are directly comparable.
  • The threshold0.5 on both, which is the conventional split and the one both vendors use in their own reporting. Below 0.5 the detector leans human, at or above 0.5 it leans AI. We report the raw scores as well as the verdicts, so you can draw a different line if you disagree with that one.
  • What we did not doNo text was scored twice with the better result kept. Nothing was dropped for being embarrassing. Every pair we generated is in the table below, including the eleven where the detectors agreed and the study is unremarkable.

One thing worth stating plainly before the numbers. This study measures agreement, not accuracy. Neither detector is ground truth. When they disagree we do not know which one is correct, and we are not claiming to. What we can measure is how often two tools that are both treated as authoritative reach opposite conclusions about the same words, and that turns out to be often enough to matter a great deal to anybody on the receiving end of one of those verdicts.

The headline result, all 18 rows

Scores are AI probability from 0 to 1, where 0 means the detector is confident a person wrote it and 1 means the detector is confident a machine did. The gap column is the absolute difference between the two.

WordsGPTZeroWinstonGapVerdicts
1091.0000.0160.984Opposite
1760.0020.0010.001Both human
1680.0021.0000.998Opposite
1630.0070.7380.731Opposite
2220.0010.0000.001Both human
1220.0020.0020.000Both human
1181.0000.0480.952Opposite
1730.0030.0000.003Both human
970.0010.0520.051Both human
1470.0010.0000.001Both human
2190.0000.0000.000Both human
2100.0000.6520.652Opposite
630.0000.0170.016Both human
760.0000.9880.988Opposite
740.0010.0430.042Both human
720.0000.0000.000Both human
880.0000.8910.891Opposite
820.2580.0120.246Both human

Seven of the eighteen rows end in opposite verdicts. That is 38.9 percent, close to four in ten. The average gap across all eighteen is 0.364, but the average hides the shape of it. The median gap is only 0.051, because when these two detectors agree they agree almost perfectly, often to within a thousandth. There is very little middle ground. Either they land on top of each other or they land at opposite ends of the scale.

That split pattern is the part that should worry anybody relying on a single score. It is not that the detectors are gently noisy around one another, the way two thermometers might differ by half a degree. They either concur completely or they contradict completely. A disagreement here is not a rounding error. It is a different answer to the question.

Not once did both detectors agree it was AI

This is the result we did not expect and the one we think matters most. Across all eighteen texts there is not a single row where both detectors said AI. Eleven times they both said human. Seven times exactly one of them said AI while the other said human. Zero times did they agree on an accusation.

OutcomeCountShare
Both said human1161.1 percent
Both said AI00 percent
GPTZero said human, Winston said AI527.8 percent
GPTZero said AI, Winston said human211.1 percent

Put the consequence in plain terms. In this sample, every single time a piece of text was called machine written, that call came from one tool while an equally mainstream tool looking at identical words said the opposite. There was no case where the evidence lined up. If each of those verdicts had arrived on its own, in an email from a professor or an editor, it would have looked like a finding. Seen side by side, all seven look like a coin landing differently in two different hands.

We want to be careful about how far that stretches. Eighteen texts is a pilot, and a larger sample will almost certainly contain rows where both detectors agree something is AI. We would expect that. But a zero across eighteen paired samples is not nothing either, and it points at something real about how little these tools corroborate each other on the cases that carry consequences.

The disagreements go both ways, and one direction is worse

Five of the seven splits went one way, with GPTZero passing the text and Winston flagging it. Two went the other. That asymmetry is not the interesting part at this sample size, since two and five are close enough to be chance. What is interesting is what each direction does to a person.

WordsGPTZeroWinstonWhat happened
760.0000.988Passed clean on one, condemned on the other
1680.0021.000The widest gap in the study, 0.998
880.0000.891Perfect human score on one, near certainty of AI on the other
1630.0070.738A clear pass turned into a clear flag
2100.0000.652The longest text in this group, and the narrowest of the five
1091.0000.016Reverse direction, condemned on one and cleared on the other
1181.0000.048Reverse direction, same pattern

The five where the first detector passed and the second flagged are the dangerous ones, and they are dangerous precisely because they feel safe. You check your work, you get a green light, you submit it, and then the person receiving it runs a different tool. Nothing you did was careless. You verified. The verification was simply against the wrong grader, and you had no way of knowing that, because nobody tells you the graders disagree.

The two in the reverse direction matter for a different reason. Those are texts one detector was completely certain were machine written, scoring a flat 1.000, and the other was completely certain were not, scoring 0.016 and 0.048. If you had been accused on the strength of the first score, the second score is not a technicality. It is direct evidence that the tool being used against you is not reproducible by an equivalent tool on the same input.

What this looks like at larger scale

The paired study above is small on purpose, because paired scoring across two paid APIs costs real money and we would rather publish eighteen honest rows than a hundred estimated ones. But we hold two much larger datasets that show the same instability from different angles, and they belong next to it.

DatasetSizeWhat it shows
GPTZero scores on delivered output1,000 scored runs74.9 percent pass at or below 0.5. 25.1 percent fail. 23.3 percent score 0.9 or higher, so roughly one run in four is called AI with confidence by a single detector.
Public detector submissions1,000 checks43.0 percent of everything people paste in scores 0.5 or above. 29.9 percent lands in the borderline band, where the detector will not commit either way.
Paired cross-detector study18 texts38.9 percent opposite verdicts. Zero unanimous AI verdicts.

The middle row deserves a sentence of its own. Nearly three in ten of the real texts people submit come back borderline, which is the detector openly saying it does not know. That band exists in every one of these tools and it almost never survives into the way results get used. A borderline score gets rounded, in somebody's head, into a yes or a no, and then a decision gets made on the rounded version.

Read the three rows together and a consistent picture appears. A quarter of confidently delivered text gets called AI by one tool. Three in ten submissions land in a zone the tool itself will not call. And when you put two tools on identical words, they contradict each other four times in ten. None of that is the behaviour of an instrument you would want deciding anything important on its own.

If your work has been flagged and you wrote it yourself

This is the situation the study is most useful for, so here is what we would actually do, in order.

  1. 01
    Get the specific score and the specific tool.Not a screenshot of a colour, the number and the product name. A verdict without a score and a source is not something you can respond to, and asking for it is a completely ordinary request.
  2. 02
    Run the identical text through a different detector.Same words, no edits, no tidying up. If the second tool returns a materially different verdict, you now hold the most relevant piece of evidence available to you, which is that the instrument is not reproducible on your own text.
  3. 03
    Gather your drafting history before you write any reply.Version history in your word processor, timestamps, notes, outlines, browser history on your sources, earlier drafts sitting in your email. Process evidence is far harder to dismiss than a competing score, and it takes minutes to collect while it still exists.
  4. 04
    Keep the reply narrow and factual.You are not arguing that detectors never work. You are pointing out that this specific verdict rests on one tool, that an equivalent tool disagrees, and that you have documentation of how the work was produced. That is a much stronger position than a general argument about whether AI detection is valid.
  5. 05
    Ask what the threshold and the appeal process actually are.Many institutions have neither written down. Asking politely and in writing tends to move the conversation from accusation toward procedure, which is where the facts you have collected are worth the most.

None of that is a trick and none of it depends on our product. It is what we would tell a friend, and most of it is available to somebody who has never heard of us.

What we would do differently if we were you

  • Never trust one detectorIf a decision matters, check the text against at least two independent tools before it leaves your hands. Our data says there is roughly a four in ten chance they will tell you different things, and it is better to learn that before somebody else does.
  • Treat borderline as a failThree in ten real submissions land in the band where the detector declines to commit. Whoever receives your work will not see it that way. They will see the colour, or the headline number, and they will round it. Assume the round goes against you.
  • Keep your draftsVersion history is the cheapest insurance available and it costs nothing until the day it is worth everything. Write in something that keeps revisions, and do not clear them.
  • Be suspicious of any single published pass rateOurs included. A pass rate against one detector is a fact about that detector. This study exists because we could not honestly claim our own GPTZero number generalised, and when we checked, it did not.

Limits of this study

Eighteen paired texts is a pilot and we are calling it that. It is enough to show the effect is real and large, and not enough to pin the disagreement rate to a precise figure. A hundred pairs would tighten it considerably, and we intend to run that.

It covers two detectors out of many. Turnitin, Originality, Copyleaks, Sapling and the rest may behave differently, and Turnitin in particular cannot be tested this way by anyone outside an institution, because it sells no public AI detection API. Any claim you read about a Turnitin pass rate, including claims made by our competitors, has no verifiable measurement behind it for that reason.

The sample is production output from one humanizer, which means it is not a neutral corpus of arbitrary writing. A study on unedited human essays, or on raw model output, would likely produce different disagreement rates. We would expect the effect to persist in some form because it comes from the tools rather than the text, but we have not shown that and we are not asserting it.

And the caveat that applies to the whole exercise. We sell a rewriting tool, so we are not a disinterested party. That is exactly why every row is printed above, including the eleven that make the finding smaller, and why the method section says precisely what was done. Check it, disagree with it, or run it yourself. The numbers are the argument, not us.

We will rerun this at larger scale and with more detectors as we get access to them, and we will publish the result whether it supports what we sell or not. That is the same commitment we made with the GPTZero pass rate study, where the honest number came in at 73.7 percent against an industry that advertises 99.

Frequently asked questions

  • 01Do AI detectors agree with each other?

    Often no. In our paired study of 18 identical texts, two mainstream detectors reached opposite verdicts on 7 of them, 38.9 percent. More strikingly, there was not a single text where both agreed it was AI. Every AI verdict came from one detector while the other said human.

  • 02Why did one detector say human and another say AI on the same text?

    Because they are different statistical models, trained on different data, with different thresholds. They measure how predictable your word choices are rather than detecting a signature, so there is nothing objective for them to converge on. When they disagree the gap tends to be enormous rather than small. In our data the average gap was 0.364 on a 0 to 1 scale, and the widest was 0.998.

  • 03Which AI detector is the most accurate?

    We cannot answer that from this study, and we would distrust anybody who answers it confidently. Measuring agreement is not the same as measuring accuracy, because neither detector is ground truth. What our data does show is that treating any single one of them as authoritative is not supported by how they behave against each other.

  • 04Does passing one AI detector mean I will pass all of them?

    No, and this is the most practical finding here. Five of our eighteen texts passed the first detector and were flagged by the second. If the text is going somewhere consequential, check it against more than one tool before you send it.

  • 05My essay was flagged as AI but I wrote it myself. What should I do?

    Ask for the specific score and the specific tool used, then run the identical unedited text through a different detector. If the verdicts differ, that is direct evidence the result is not reproducible. Collect your drafting history, version records and timestamps before replying, because process evidence is harder to dismiss than a competing score. Keep the reply factual and narrow rather than arguing about AI detection in general.

  • 06How many detectors should I check before submitting?

    At least two independent ones, and treat any borderline result as a fail. In our wider dataset of 1,000 real detector submissions, 29.9 percent came back borderline, which is the tool declining to commit. Whoever receives your work will not read it that generously.

  • 07Can Turnitin be tested this way?

    Not by anyone outside an institution. Turnitin sells no public AI detection API and restricts access to contracted institutions, so no independent party can measure it directly. Any published Turnitin pass rate you encounter, from us or from a competitor, has no verifiable measurement behind it. We would rather say that than quote a number we cannot stand behind.

  • 08Is this study biased because you sell a humanizer?

    We sell a rewriting tool, so we are an interested party and you should read it that way. That is why all 18 rows are published including the 11 where the detectors agreed, why the method section states exactly what was done, and why the limits section is longer than most. The finding also makes our own previously published pass rate look less generalisable than it did, which is not a result a marketing department would ask for.