The Commonwealth Short Story Prize controversy in mid-2025 gave writers a preview of a bad new world. Multiple longlisted authors were told, after the fact, that their entries had been flagged by Pangram, an AI detector, and disqualified. Some, including writer Anna Vangala Jones, said publicly on social media that they had not used AI. If you write, edit heavily, or use YouWrite's refine tool to tighten a draft you already wrote, this is your problem too. The appeal process is worse than you think.
What detectors actually measure
AI detectors do not read intent. They compare your text against a statistical model of what machine-generated prose tends to look like: token predictability, burstiness (variation in sentence length and complexity), perplexity (how surprising your word choices are to the model), and syntactic uniformity. That is it. When Pangram or GPTZero returns a number, it is reporting a probability that your text pattern matches its training distribution of AI-written samples. It is not detecting cheating. It is detecting resemblance.
The marketing suggests otherwise. Pangram's site claims accuracy figures above 99% on internal benchmarks. Independent testing tells a different story. A 2023 Stanford study by Weixin Liang and colleagues, published in Patterns, found that GPTZero and six other detectors flagged 61% of TOEFL essays written by non-native English speakers as AI-generated, versus near-zero false positives on essays by native speakers. The paper is titled "GPT detectors are biased against non-native English writers." Read it before you trust a percentage.
Who gets falsely flagged
Predictable patterns get punished. That includes:
- ESL writers, whose vocabulary tends to be smaller and more common, which lowers perplexity.
- Spare, plain-style prose. Hemingway would fail. So would most technical writing.
- Heavily edited work. Revision smooths burstiness. The more passes you do, the more your sentences approach an average length and structure.
- Genre writing with conventions: romance beats, cozy mystery cadences, YA voice, that trend toward familiar phrasing.
- Formal academic prose, which is trained into models heavily and therefore resembles model output.
Notice what these have in common: nothing to do with cheating. A careful ESL novelist who has revised a story twelve times is exactly the profile a detector loves to flag.
The asymmetry problem
If a detector accuses you, the burden falls on you to prove a negative. The accuser (an editor, prize administrator, or professor) has a number on a screen that looks scientific. You have your word. There is no equivalent of DNA evidence for authorship. This is unjust, and pretending otherwise wastes energy. The best you can do is make the cost of doubting you higher than the cost of believing you.
Build the paper trail before you need it
Writers who survived the Commonwealth situation with their reputations intact tended to have one thing: receipts. Not because receipts prove innocence, but because they shift the conversation from vibes to specifics.
Draft in tools that log history
Google Docs keeps version history indefinitely by default. Microsoft Word with AutoSave on OneDrive does too. Scrivener saves snapshots. If you draft in a plain text editor and save only the final file, you have nothing to show. Change that today. Turn on version history on every document you might one day defend.
Keep the ugly drafts
Do not delete the messy first pass. Do not overwrite. A clean linear progression from "terrible sentence" to "good sentence" across timestamped revisions is the most persuasive artifact you can produce. It shows thinking. AI output does not have a thinking trail. It arrives whole.
Record your process in adjacent artifacts
Notes to yourself. Marginalia. Voice memos where you talk through a plot problem while walking. Photos of a notebook page. Emails to a critique partner asking whether the second act sags. None of this is proof, but together it forms a texture that is hard to fabricate after the fact.
Know what your tools do
If you used Grammarly, ProWritingAid, or YouWrite's refine feature, write that down when you use it, and keep the before-and-after. Refine rewrites at the sentence level and will raise your detector score. That is not cheating, but you should be able to say clearly what you did and when. Vague answers read as guilty.
What to do the day you are accused
Do not send a long defensive email in the first hour. Almost every writer who has done this has regretted the tone.
Ask, in writing, three things: which detector was used, what score your work received, and what the organization's stated threshold and appeal process are. Many organizations have no documented process. Forcing them to articulate one in writing is useful, both for you and for the next writer.
Then produce your artifacts. Share the version history link. Offer to walk through the draft evolution on a call. Point to the Liang et al. study if the accusation seems tied to a style pattern rather than specific evidence. If you are an ESL writer, say so, and cite the false positive rate.
One thing to resist: running your own text through three other detectors to "prove" it is human. The scores will vary wildly, which undermines your point that the scores are not reliable. Pick your argument and hold it.
The limits of self-defense
Honestly: none of this is foolproof. An editor who has decided you cheated will read your version history as elaborate staging. A prize committee protecting itself from embarrassment will not reverse a public decision easily. You may lose. Some writers already have.
The deeper fix is not individual. It is institutional. Publications and prizes need to publish their detection policies before submissions open, disclose which tools they use, commit to human review, and accept that some percentage of their decisions will be wrong. Until then, the accused writer's agency is real but limited. You can make yourself the hardest possible target. You cannot make yourself immune.
A note on YouWrite
We build tools that help writers draft and refine. Our users are exactly the population most likely to be flagged. Refine changes sentences; it will move your detector score. We are working on process-log exports so writers using our tools have a defensible record by default. We do not have that shipped yet, and until we do, the burden is on you to keep your own trail. That is a real limitation, not a marketing sentence.
