§01United Kingdom guide

Turnitin AI detection
at UK universities, explained honestly.

If your university uses Turnitin, your assignment is probably being scored for AI writing whether or not anyone told you. This page explains what that score is, what it is not, what your institution can and cannot do with it, and what the measurable evidence actually says. It also tells you plainly where nobody, including us, can give you a real number.

What Turnitin actually reports

Turnitin's AI writing indicator returns a percentage that estimates how much of a submission it believes was generated by a language model. It sits separately from the similarity score, which is the older and much better understood plagiarism measure. The two are frequently confused, including by staff, and they mean entirely different things. A high similarity score means your text matches an existing source. A high AI score means a statistical model thinks the writing pattern looks machine generated, with no source involved at all.

That distinction matters enormously if you are ever asked to explain one. Similarity is evidence you can interrogate, because it points at a specific document you either did or did not copy. An AI score points at nothing. It is a probability derived from sentence rhythm, vocabulary predictability and structural regularity, and there is no source to produce and no passage to compare.

Turnitin itself has been careful in its documentation to describe the indicator as a signal requiring human review rather than proof of misconduct. Institutions vary widely in how faithfully they follow that guidance.

Why the false-positive problem is worse for some students

AI detectors do not measure authorship. They measure predictability. Writing that is careful, formulaic, or produced by someone working in a second language tends to be more statistically regular than writing that is fluent and idiosyncratic, and regularity is the exact signal being flagged.

The most cited evidence on this remains Liang et al. at Stanford in 2023, which found detectors flagged 61.3 percent of TOEFL essays written by non-native English speakers as AI generated, against 5.1 percent for essays by US-born students. That gap has narrowed as detectors have improved, but the underlying mechanism has not changed, because it cannot: a detector that stops penalising predictable prose stops working.

Our own measurement of a different detector points the same way. We scored 100 genuinely human, pre-2022 paragraphs on GPTZero and found a 1 percent false-positive rate overall, but one of those human paragraphs still returned a full 1.00 AI score. A low average rate is little comfort when you are the one paragraph.

The practical consequence for a UK student is straightforward. If you write in a careful, structured academic register, particularly if English is not your first language, your risk of a false flag is higher than the headline numbers suggest, and it has nothing to do with whether you cheated.

What UK universities can and cannot do with a score

Practice varies by institution, but the pattern across published UK academic misconduct procedures is consistent: an AI score is treated as a trigger for investigation, not as a finding in itself. A panel is expected to consider the score alongside other evidence, which typically means your drafting history, your previous work, and an interview in which you are asked to discuss your own submission.

Several Russell Group institutions have gone further and explicitly instructed staff not to open a misconduct case on a detector score alone, precisely because the false-positive research is well known within university academic-integrity offices. That is a meaningful protection and it is worth knowing it exists.

If a case does proceed and you exhaust your university's internal appeals, the Office of the Independent Adjudicator for Higher Education can review whether the process was handled fairly. The OIA reviews procedure rather than re-marking your work, so the quality of your evidence about how you wrote the assignment matters more than any argument about the detector's accuracy.

Reading what you are told

What you are toldWhat it usually meansWhat to do
Similarity score is highText matches an existing source in the databaseCheck your quoting and referencing against the flagged passages
AI writing score is highA model thinks the pattern looks machine generated, no source involvedGather your drafting evidence before responding to anything
Invited to an academic conduct meetingInvestigation opened, not a decision madeBring version history, notes and earlier drafts, and ask what evidence beyond the score exists
Told the detector is conclusiveThis contradicts Turnitin's own guidanceAsk politely for the institution's written policy on detector evidence

How to protect yourself, whether or not you use any tool

Write somewhere that keeps version history

Google Docs, Word with autosave, or any editor that records revisions. A document that visibly grew over two weeks, with false starts and rewrites, is the single strongest piece of evidence a student can produce. It is far more persuasive than any argument about detector accuracy, and it costs you nothing to have.

Keep your notes and sources

Reading notes, annotated PDFs, a messy outline. These show a research process. Someone who submitted generated text cannot produce them retrospectively, and a panel knows that.

Know your institution's actual policy

Find the academic misconduct procedure on your university intranet and read the section on AI. Knowing whether your institution permits a case to be opened on a score alone changes how you respond if you are ever contacted.

Check your own work before submitting

If you want to know how your writing scores, run it through a detector yourself first. Being surprised in a meeting is worse than knowing in advance, and a high score on your own genuine writing is exactly the situation where drafting evidence saves you.

The bottom line

A Turnitin AI score is a probability produced by a model that measures predictability, not authorship. UK institutions generally treat it as a trigger for review rather than as proof, and the OIA exists if a process goes wrong. The most useful thing any student can do, regardless of what tools they use or do not use, is to write somewhere that keeps a version history.

We publish our own measured detector numbers, including the runs that failed, in the GPTZero pass-rate study. Those figures are measured against GPTZero rather than Turnitin, for the reason explained above.

§05Questions

Frequently asked.
Honestly answered.

  • It detects statistical patterns associated with generated text and returns a probability, which is not the same as definite detection. It produces both false positives on genuine human writing and false negatives on edited machine writing. Turnitin's own documentation describes the indicator as requiring human review rather than as proof.

  • There is no universal threshold, and any specific number you see quoted online is someone guessing. Institutions set their own review triggers and most do not publish them. Treat any non-zero AI score as a reason to have your drafting evidence in order rather than as a verdict.

  • Nobody can honestly answer that, and you should be suspicious of any tool that claims to. Turnitin provides no public detector, so no independent party can measure a pass rate against it. Our published figures are measured against hosted GPTZero, which does expose a real probability, and they are labelled as GPTZero figures for exactly that reason.

  • Gather your evidence before you respond to anything. Version history, notes, earlier drafts, reading records. Ask what evidence exists beyond the score. Ask for your institution's written policy on detector evidence. If the internal process concludes unfairly, the OIA can review how it was handled.

  • The evidence says yes. The Stanford study found 61.3 percent of non-native TOEFL essays flagged against 5.1 percent for US-born writers. Detectors penalise predictable, careful prose, and second-language academic writing is often exactly that. This is a known limitation of the technology rather than a reflection on the writer.