CWR

Accused of using AI when you didn’t

A flag is not a finding. Here is what actually helps, in the order it helps — and what to stop arguing about.

Last reviewed

This page is for the person on the receiving end. It is not legal advice, and it cannot tell you your institution’s rules. What it can do is separate the claims that survive scrutiny from the ones that collapse under it.

First: find out what actually flagged you

The single most useful thing you can establish is which of three very different things happened. People conflate them constantly, including the people doing the accusing.

  1. A commercial AI detector — Turnitin, GPTZero, Copyleaks and similar. These are statistical classifiers. They output a probability, they have a documented false positive rate, and that rate is markedly worse for writers whose first language is not English and for anyone with a plain, formulaic style. This is the most common cause of a false accusation and the weakest form of evidence.
  2. A provider watermark check — an actual query against Anthropic’s detection for its own model. This is far more reliable at what it measures, but what it measures is narrow: that text passed through Claude, not that Claude authored it.
  3. Someone’s judgement — a marker who thought it read like AI. Sometimes right, entirely unfalsifiable, and often retrofitted with a detector score afterwards.

Ask, in writing, which one it was. If the answer is a detector score, ask for the score and the tool’s own published false-positive rate. If it is a watermark check, the argument moves to a different question, below.

If it is a watermark hit and you did use AI assistance

Be straightforward about it. If you had Claude proofread, translate or tidy your own writing, the mark is expected and its presence is not in dispute — Anthropic’s documentation states the mark means content may have been processed by Claude, which includes exactly that use.

The real question is whether the assistance you used was permitted, and that is a policy question about your institution’s rules, not a technical one about the watermark. Argue it there. Trying to dispute the detection itself when you did run the text through the model will cost you credibility you need for the part that matters.

If it is a detector score and you used nothing

Then the classifier is wrong, and your job is to make that concrete rather than to argue about AI in the abstract.

What holds up

  • Version history. Google Docs revision history, Word’s AutoRecover versions, an Overleaf or Notion timeline, or a git log. A document that grew over hours with real deletions and reordering is very hard to fake and very easy to show.
  • Process artefacts. Notes, outlines, annotated sources, photographs of handwritten drafts, the browser history behind your research, dead ends you abandoned.
  • Being able to talk about it. Offering to discuss the argument, defend a choice, or explain why you dropped a section. Someone who wrote it can do this cold.
  • A baseline. Earlier work of yours, graded before this became an issue, showing the same voice.

What wastes your time

  • Running the text through more detectors. Contradictory outputs from unreliable tools do not add up to reliability, and a second opinion from a fourth classifier persuades nobody.
  • Rewriting to score lower. It reads as consciousness of guilt and it concedes the premise that the score means something.
  • Stripping invisible characters. This is worth saying plainly: it does nothing here. Commercial detectors do not look at zero-width characters, and Claude’s text mark is not made of characters. You will have altered your document for no benefit, and having done so is worse for you than not having done so.

Questions worth putting in writing

Politely, in an email, so there is a record. These are not gotchas — a well-run process has answers to all of them.

  • What specifically triggered the concern, and what tool produced it?
  • What is that tool’s stated false positive rate, and what threshold is being applied?
  • Is a detector score alone sufficient for a finding under the written policy?
  • What is the appeal route, and what evidence would be considered?
  • Can I present my drafting history and discuss the work directly?

Many institutions have quietly moved away from treating detector output as sufficient evidence, precisely because of the false-positive problem. Asking what the policy actually says is often enough on its own.

One thing this site can check for you

If the accusation concerns an image rather than text, provenance is readable and you can look yourself. Run the file through the C2PA inspector: it will show any signed manifest, what claims to have produced the file, and what actions are recorded against it.

Note the asymmetry before you rely on it. A manifest declaring AI generation is meaningful. An absent manifest proves nothing — metadata is destroyed by screenshots, exports and most platforms — so it cannot clear you, only fail to incriminate you. For text, no third-party tool can check Claude’s mark at all, including this one, because the method is unpublished. Anyone claiming otherwise is selling something.