§The journal

Which AI Is Hardest to Detect? We Tested 4 in 2026

We gave ChatGPT, Claude, Gemini and DeepSeek the same prompts and ran 176 texts through GPTZero. It flagged 176 of them. Here is the full ranking.

Published September 27, 20266 min readBy Abd Shanti
Small brass weights of different shapes lined up beside a notebook with teal, aqua and coral tabs on cream paper

We gave four of the most used AI models the same 32 writing jobs, from a university essay to a LinkedIn post, and ran all 128 answers through GPTZero. It caught 128 of them. Then we told every model to write like a real person so it would not sound like AI. It caught 48 of those 48 too.

176 of 176
Texts from four AI models flagged by GPTZero
175 of 176
Scored the maximum AI probability of 1.000
48 of 48
Still flagged after telling the model to write like a real person

What we tested

Each model got the same 32 prompts, eight for each of four everyday kinds of writing, with a word count so the texts came out at similar lengths. Every model got the plain prompt with no system instructions and no tricks, the way most people use these tools. Every answer was sent once to the GPTZero API as a whole document in September 2026, and we recorded the score it returned. Nothing was retried or dropped.

Kind of writingPromptsExampleTarget length
University essay paragraph8Whether social media has made teenagers lonelierAbout 200 words
Work email8Asking your manager to move a deadline by one weekAbout 150 words
Blog post opening8How to start running after 40About 200 words
LinkedIn post8A lesson from a failed product launchAbout 150 words

The four models were ChatGPT (GPT-5.5), Claude (Sonnet 5), Gemini (3.1 Pro) and DeepSeek (V4 Pro). Grok is not included because we had no access to its API, and Qwen is not included because its run was not complete when we published.

The ranking: from least to most detectable

ModelFlaggedAverage AI scoreLowest scoreSentences marked AI
Gemini (3.1 Pro)32 of 320.9970.899100%
ChatGPT (GPT-5.5)32 of 321.0001.000100%
Claude (Sonnet 5)32 of 321.0001.000100%
DeepSeek (V4 Pro)32 of 321.0001.000100%

The order at the top is decided by a hair. Gemini ranks first only because of one text that scored 0.899 instead of 1.000. In practice all four models landed in the same place: a detector certain that a machine wrote the text, with every sentence marked.

Does the kind of writing matter?

Kind of writingFlaggedAverage AI scoreLowest score
University essay paragraph32 of 321.0001.000
Work email32 of 321.0001.000
Blog post opening32 of 320.9970.899
LinkedIn post32 of 321.0001.000

Casual formats did not help. We expected short emails and chatty LinkedIn posts to slip through more often than formal essays, because people write those loosely too. They did not. The models brought the same polish to a two line thank you note that they bring to an essay.

Telling the AI to sound human does nothing

For three prompts of each kind we ran the same request again with one extra line: write it the way a real person would actually write it, casual and natural, so it does not sound like AI. This is the most common trick people try, and it is repeated in countless prompt guides.

ModelFlagged with the human voice promptAverage AI score
ChatGPT (GPT-5.5)12 of 121.000
Claude (Sonnet 5)12 of 121.000
Gemini (3.1 Pro)12 of 121.000
DeepSeek (V4 Pro)12 of 121.000

It made no difference. 48 of 48 were flagged, and the average score for every model stayed at or next to 1.000. The texts did read more casually, with contractions, short sentences and the odd aside. The detector did not care, because the tone was never what it was measuring.

Why every model looks the same to a detector

Detectors like GPTZero do not look for a model's style or its favourite words. They measure how predictable the text is, word by word, and how evenly that predictability is spread across sentences. Every large model is trained to pick likely words and to keep its sentences smooth, so its writing is predictable in the same way whichever company built it. A casual tone changes the words on the surface but not that underlying pattern.

Human writing is uneven. People start a sentence one way and finish it another, use an odd word because it is the one they know, and write one line of four words next to one of forty. That unevenness is what a detector reads as human, and no model reproduced it when asked.

What actually lowers a detection score

  1. 01
    Write the first draft yourselfUse AI for research, outlines and checking, then write the text in your own words. Every text in this test that a model wrote from scratch was flagged.
  2. 02
    Add what only you knowReal names, real numbers, a specific moment. Detail the model could not have guessed is also detail that breaks its predictable rhythm.
  3. 03
    Edit sentence by sentence, not word by wordSwapping synonyms leaves the pattern intact. Changing how sentences are built, their length and their order, is what moves a score.
  4. 04
    Do not trust a tone promptAs the test shows, asking for a human voice changes how the text sounds to you, not how it scores.
  5. 05
    Check before you submitRun the final version through a detector yourself and rewrite the sentences it marks.

Limits of this test

This is one detector, GPTZero, scoring 176 texts once each. Each model got one plain prompt per task at its default settings, so a person who edits the output, or uses a detailed system prompt, will get different results. We used each company's current flagship or most used model in September 2026; newer versions will come out. Grok and Qwen are not included. The texts are 150 to 200 words; detectors are known to be less certain on very short text and more certain on long text.

The bottom line

If you are choosing an AI model because you think one of them is harder to detect, stop. In our test the difference between the most and the least detectable model was one text scoring 0.899 instead of 1.000. ChatGPT, Claude, Gemini and DeepSeek all write in a way detectors recognise at once, and a prompt asking them to sound human does not change it. The only thing that reliably changes a score is a person rewriting the text in their own way.

Sources

  1. 01GPTZero API documentation
  2. 02OpenAI models (GPT-5.5)
  3. 03Anthropic Claude models (Sonnet 5)
  4. 04Google Gemini API models (Gemini 3.1 Pro)
  5. 05DeepSeek API documentation (V4 Pro)
Cite this article
HumanGPT (2026). Which AI Is Hardest to Detect? We Tested 4 in 2026. Published 27 September 2026. https://humangpt.io/blog/which-ai-is-hardest-to-detect-2026

Frequently asked questions

  • 01Which AI is hardest to detect in 2026?

    None, in our test. ChatGPT, Claude, Gemini and DeepSeek were flagged by GPTZero on all 176 texts. Gemini ranked first only because one of its texts scored 0.899 instead of 1.000.

  • 02Is Claude harder to detect than ChatGPT?

    No. Both were flagged on every text we tested, at the maximum AI score, with every sentence marked as AI.

  • 03Can GPTZero detect Gemini?

    Yes. Gemini 3.1 Pro was flagged on all 44 of its texts. Its lowest score was 0.899, still a clear AI verdict.

  • 04Can GPTZero detect DeepSeek?

    Yes. Every DeepSeek V4 Pro text in our test was flagged at the maximum score of 1.000.

  • 05Does telling ChatGPT to write like a human work?

    No. We added that instruction to 48 prompts across all four models and GPTZero still flagged 48 of them. The texts sounded more casual but scored the same.

  • 06Are casual emails or social posts harder to detect than essays?

    Not in our test. Work emails and LinkedIn posts were flagged just as often as university essay paragraphs, all at or near the maximum score.

  • 07What makes AI writing easy to detect?

    Detectors measure how predictable the text is word by word and how evenly that predictability is spread. Every large model writes predictably in the same way, whatever tone it is asked for.