
AI in School and University: Tool, not Ghostwriter
Listen to this article
Personal note
SavedStored only locally in your browser — nothing is sent.
Two reading levels: the running text is the through-line — for everyone, no prior knowledge needed. The clearly marked In-depth boxes explain the technology in detail and can be skipped without losing the thread.
The pocket-calculator reflex
When the pocket calculator entered schools, the fear ran high: now the children will unlearn how to do arithmetic. It turned out otherwise. The calculator took the dull crunching off the pupils’ hands and left room for what matters — understanding, modelling, problem-solving. Today it is a matter-of-course tool.
With AI the same reflex suggests itself, and for research, spell-checking or sorting one’s thoughts the analogy holds up: a tool one can no longer imagine everyday life without. But the analogy has a fracture that must be taken seriously. The calculator automates a secondary skill — arithmetic — while the actual purpose of learning, mathematical thinking, stays with the human. With writing it is different — at least in those subjects where the act of formulating is the intellectual work: there, writing is not the packaging of thought but the process of thinking itself. Whoever composes an essay orders their thoughts, weighs, argues. Let the AI do that work, and you are not outsourcing arithmetic but thinking.
Precisely here — not at the tool as such — runs the line at issue.
The line: writing with AI ≠ letting the AI write
German institutions have by now drawn this line with remarkable clarity — and not as a ban. In October 2024 the Standing Conference of the Ministers of Education (KMK) adopted a recommendation for a “constructive-critical” approach and expressly calls for the culture of assessment to be adapted. The German Rectors’ Conference (HRK) considers a blanket AI ban “neither sensible nor practicable”. The consensus at universities in 2024/25 is not prohibition but graduated permission with a duty of disclosure: AI as an aid for research, outlining and language is legitimate; having the AI produce the actual substance and passing it off as one’s own work is deception.
Legally it all hangs on a single, old sentence: the declaration of authorship, by which one affirms that one has written the work oneself and named all aids used. Undisclosed use of AI violates exactly that declaration — that is the offence of deception, not the keystroke. Many universities therefore now require an AI appendix: which tools, for what, in which sections. Whoever discloses may not be marked down. That is the legitimate zone — and it is large.
The real point of contention is thus not “yes or no” but: can the line be policed? And here begins the part about which the most nonsense is told.
The core: can AI-written text be detected at all?
The widespread assumption goes: if a pupil has let the AI write, you can see it — there are detectors, after all, and if all else fails the AI gives itself away through hidden characters. Both are, in essence, a myth. To understand why, one first has to know how these detectors actually work.
At heart they measure one single thing: predictability. A language model produces text by repeatedly choosing the statistically obvious next word. A detector turns this around — it checks how “surprised” a language model would be by each word of the text put before it. If the text is consistently unsurprising and even in rhythm, it counts as machine-made; if it is jumpy, variable, unexpected, as human. That is the whole idea. And from it follows, of necessity, every mistake these tools make.
In-depth · for specialists Three families of detectors, the same foundation. (1) Statistical (the GPTZero layer): perplexity = the mean surprise of the reference model at the choice of words; burstiness = the variance of that surprise across the text. AI is low on both. (2) Trained classifiers (RoBERTa-based detectors, Turnitin, Originality.ai): a transformer is fine-tuned on a labelled corpus {human, AI} and learns surface features that separate the two classes in the training data — which is why every new model and every paraphraser shifts the target. (3) Zero-shot stylometry (DetectGPT, Stanford 2023): uses the log-probabilities of a reference model and the observation that AI text sits on local maxima of the probability function. All three rest on the same assumption — that AI text is “more predictable” — and all thereby inherit the same blind spot.
How bad they really are
The strongest witness is the maker itself. In January 2023 OpenAI launched its own “AI Text Classifier” and shut it down again in July 2023 — because the hit rate was too low. The figures OpenAI itself cited: the detector correctly identified 26 percent of AI texts while at the same time wrongly flagging 9 percent of genuine human texts as AI. The company that builds the models could not manage a usable detector for its own output.
The market leader in education, Turnitin, advertises a false-positive rate of under one percent — but itself conceded that in reality it is higher, and at sentence level around four percent. What a one-percent error rate means in practice, Vanderbilt University worked out when it disabled Turnitin’s AI detector in August 2023: with 75,000 papers a year, that would be roughly 750 papers wrongly flagged as AI — 750 potential cheating accusations against the innocent. Michigan State, Northwestern, UT Austin and others followed suit.
Most serious of all is whom the errors hit. A Stanford study (Liang et al., 2023) tested seven detectors on genuine essays by non-native speakers: 61.3 percent were on average wrongly classified as AI, 97.8 percent by at least one detector, nearly a fifth by all seven unanimously — while texts by native speakers passed through almost error-free. The reason is the same mechanism: whoever has a more limited vocabulary writes “more predictably” — and is mistaken by the machine for AI. For the same reason detectors classify the US Constitution and passages of the Bible as “AI-written”: texts that sit a thousand times over in the training data are statistically smooth. Real students have already been failed unjustly and accused of cheating through such false alarms.
On top of this comes a statistical trap. Many pupils and students do use AI — but the complete ghostwriting that an examination actually means to penalise is the exception in the pile, not the rule. And the rarer the thing sought, the more false alarms: against the great mass of honest texts, even an accurate tool errs more often than it hits the few genuine cases (the arithmetic is in the box further down). What matters, therefore, is not the hit rate but the harm of a single false alarm — of one person wrongly accused of cheating.
One might object that all these findings come from the early days of the detectors. True — the hit rates have improved since. But the real problem does not sit in the rate; it sits in the design. And that has not changed.
The decisive point: a detector never tells you what the AI did
Even if a detector were right — it does not answer the question teachers are really asking. For it delivers only a single probability — a measure of how AI-like the text looks. It cannot distinguish whether a text was generated entirely by AI, merely corrected in its spelling, reworded, or translated from another language. The providers say so themselves, unmistakably: “60 percent AI does not mean that 60 percent of the text comes from AI” (Originality.ai); the percentage value does “not mean that 6 percent of the document contains AI” (GPTZero), only the model probability. Because the software cannot resolve the degree of involvement, in practice it collapses everything into an all-or-nothing verdict — “human-written, heavily AI-edited = AI-generated”. Precisely the distinction that is decisive for education — tool or ghostwriter? — is the one no detector can make.
And even a correct finding is trivially circumvented
The signals on which everything rests are fragile. A paraphrasing model called DIPPER, in one study, pushed the hit rate of the DetectGPT detector down from 70.3 percent to 4.6 percent — through mere rewriting, without altering the content. Light manual reworking or a translation already suffices, because they destroy exactly the even, predictable signature the detector depends on. Whoever wants to deceive circumvents the tools without much effort; those left hanging are the honest ones, whose clean style looks “too smooth”. Only, whoever sees a free pass in this deceives himself: let the AI write, and you buy a grade and lose what the work is for — the detector can be tricked, your own learning cannot.
What about watermarks and hidden characters?
That leaves the hope of a built-in signature — surely the AI gives itself away? Here one must separate two entirely different things that are constantly confused.
The one is a genuine statistical watermark: at the moment of generation, the model provider skews the choice of words according to a secret pattern that a matching detector recognises later. Google presented something of this kind (SynthID) in 2024, and technically it works. But it has three decisive snags. It exists only if the provider actively builds it in — local and open-source models mark nothing at all, so a missing watermark proves nothing. It does not survive thorough rewriting or a translation, and it barely grips on short or factual texts. And only whoever holds the secret key can check it — a teacher is not among them.
In-depth · for specialists The basic scheme (Kirchenbauer et al., 2023): before each token, a pseudo-random function splits the vocabulary into a “green” and a “red” list; the sampler gently favours green tokens. The detector counts the green share and tests it against the 50-percent chance expectation — entirely without model access, but with the key. SynthID uses a refined variant (“tournament sampling”). Added to this is a fundamental base-rate problem for every detector: if only 5 percent of texts are actually AI-ghostwritten, but a tool is 1 percent false-positive and 95 percent accurate, then roughly one sixth of all hits is a false alarm. And the rarer complete ghostwriting is in the pile, the more sharply the ratio tips: with genuinely rare use, most hits are false alarms — the classic base-rate fallacy.
The other is invisible special characters — for instance a narrow space that looks like an ordinary one. In April 2025 it was noticed that some ChatGPT models were scattering such characters, and at once there was talk of a “secret watermark”. OpenAI’s explanation: no watermark, but a training artefact — and in any case meaningless, because such characters are trivially removed. A mark of recognition that vanishes on being copied over is good for nothing. No serious detector relies on it.
And the folk “proofs” — the em dash, buzzwords like “delve”, too-smooth language? They are real clusterings across millions of texts, but worthless as evidence for a single one. Humans have set em dashes for centuries; you can drill the dash out of the AI, and back in, by instruction. And in a direct comparison, humans recognise AI text with the naked eye in only about 53 percent of cases — barely better than a coin toss.
And the EU law?
Since 2 August 2026 the transparency obligation of the EU AI Act (Article 50) has been in force: providers of generative AI must mark their outputs — text, image, sound, video — machine-readably as artificially generated, and whoever publishes AI text in order to inform the public on matters of public interest must disclose it. It sounds like the solution, but it is not one — at least not for the teacher. For one thing, the obligation is addressed to providers and deployers, not to the pupil who has the AI write his coursework and hands it in as his own; a term paper does not “inform the public”. For another, the labelling relies on exactly the watermarks and metadata that a rewording, copy or translation destroys — and the EU Commission concedes in its own guidelines that no single marking technique meets all requirements and that forensic detection is not yet reliable enough. Existing systems have until December 2026, an interoperable watermark detection even until February 2027.
The law makes sense — it strengthens labelling at the source. But it confirms precisely the finding of this essay: even the legislator does not bet on convicting AI text after the fact, but on disclosure. For school and university that means the same.
What follows from this — for school and university
If detection does not work, that does not mean: surrender. It means the control shifts from detection to assessment design — and that is exactly where the universities are already moving. What an AI cannot replace makes one’s own achievement visible: the oral defence, the subject-matter conversation about one’s own work, writing under supervision, the demonstration of the process through drafts and intermediate steps, the mandatory, honest disclosure. That is more inconvenient and more expensive than a click on “check” — it costs time that teachers are not handed for free. But it is the only reliable way, one that catches the right people instead of the wrong ones. Whoever has to explain his coursework in a colloquium cannot have had the AI write it without it showing — with no detector at all.
That does not mean a detector is worth nothing. As an occasion to look more closely somewhere — a smoke alarm that starts a conversation, not one that passes a verdict — it may have its place; and with thousands of papers that no one can examine orally one by one, a rough pre-sorting does help. Detection and assessment design are therefore not opposites. Only, in the end the human being must decide in conversation, never the bare probability — for the number cannot prove who deceived.
And the question of the benefit? It has a surprisingly clear, empirical answer — and it falls exactly on the tool/ghostwriter line. Used as a tutor that asks, explains and gives feedback, an AI purpose-built for it doubled the learning gains against active in-class learning in a randomised Harvard study. Used as a ghostwriter, on the other hand, that delivers the finished text, something insidious happens: in a controlled study the ChatGPT group achieved the best essay grades — but no measurable additional gain in knowledge and transfer over the other groups. Better product, no better student. That is the line, in one sentence: AI you think with makes you cleverer; AI that thinks for you delivers a better piece of work and a poorer human being.
What stays with you
The defensive stance of many teachers and universities is understandable, but it is fighting on the wrong front. AI will not be kept out of school and university, and its use will not be reliably detected either — the tools for it are unreliable, blind to what matters and trivial to circumvent; whoever relies on them ends up accusing the wrong people. The honest answer is more inconvenient and better at once: it is not the machine that is to be convicted, but one’s own achievement that is to be shown. AI is a tool like the pocket calculator — as long as it takes over the arithmetic and not the thinking. Where that line runs, no software decides. We decide it.
If this essay gave you something to think about, feel free to share it — and the next time someone tells you they can “unambiguously detect” AI text, do ask them exactly how.
Sources (selection):
- OpenAI — “New AI classifier for indicating AI-written text” (launched Jan 2023, discontinued July 2023; 26% TP / 9% FP): https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/
- Liang et al. — “GPT detectors are biased against non-native English writers”, Patterns 4(7), 2023 (61.3% false positives): https://www.cell.com/patterns/fulltext/S2666-3899(23)00130-7
- Vanderbilt University — “Guidance on AI detection and why we’re disabling Turnitin’s AI detector” (2023): https://www.vanderbilt.edu/brightspace/2023/08/16/guidance-on-ai-detection-and-why-were-disabling-turnitins-ai-detector/
- Turnitin — higher false-positive rate than advertised (Inside Higher Ed, 2023): https://www.insidehighered.com/news/quick-takes/2023/06/01/turnitins-ai-detector-higher-expected-false-positives
- Krishna et al. — “Paraphrasing evades detectors… (DIPPER)”, NeurIPS 2023 (70.3% → 4.6%): https://arxiv.org/abs/2303.13408
- Originality.ai — “AI Detection Score” (the value is a probability, not a share): https://originality.ai/blog/score-meaning
- GPTZero — “Understanding your GPTZero AI scan”: https://gptzero.me/news/understand-gptzero-ai-scan/
- Kirchenbauer et al. — “A Watermark for Large Language Models”, ICML 2023: https://arxiv.org/abs/2301.10226
- Google DeepMind — SynthID-Text (watermarking, limitations): https://deepmind.google/discover/blog/watermarking-ai-generated-text-and-video-with-synthid/
- Ars Technica — “Why AI detectors think the US Constitution was written by AI” (2023): https://arstechnica.com/information-technology/2023/07/why-ai-detectors-think-the-us-constitution-was-written-by-ai/
- KMK — Recommendation on dealing with AI in school education processes (10 Oct 2024): https://www.kmk.org/aktuelles/artikelansicht/bildungsministerkonferenz-verabschiedet-handlungsempfehlung-zum-umgang-mit-kuenstlicher-intelligenz-1.html
- Kestin et al. — “AI tutoring outperforms in-class active learning”, Scientific Reports, 2025: https://www.nature.com/articles/s41598-025-97652-6
- Fan et al. — “Beware of metacognitive laziness”, British Journal of Educational Technology, 2025 (better score, no transfer): https://bera-journals.onlinelibrary.wiley.com/doi/10.1111/bjet.13544
- European Commission — Code of Practice on Marking and Labelling of AI-generated Content (final version June 2026): https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content
- EU AI Act, Article 50 — transparency obligations (applicable from 2 August 2026): https://artificialintelligenceact.eu/article/50/
