Why AI sounds like "that"
Raters kept rewarding the more familiar good answer, so the model learned the expected shape. Asking it to sound human makes that worse.
The em dash got a bad rap. It was never the culprit, just the canary in the coal mine: the one tell anybody could spot without training, so it became shorthand for "a robot wrote this." The whole internet responded the same way, wrote "never use em dashes" into its prompts and rules files, the dashes dutifully disappeared, and we're all still scratching our heads about why the fuck AI still sounds so... weird.
RIP to the em dash, and to everyone who used one before 2023. You deserved better. Keep the rule anyway: a ski mask is legal too, and you still wouldn't wear one into a bank.
The frontier models will tell you this is a prompting problem. Don't buy it. The problem is us. And when I say "us" I mean humans, the people who train these models. No matter how smart we are or how much cool shit we build, we can't seem to overcome basic human nature, physiology, and psychology: we like things that are familiar. So familiarity is what gets fed back to us. A warm blanket on a cold night. Except the blanket runs too warm, and you get sweaty, and you feel gross, and that's where we are: everybody's got swamp ass from sweating inside the sauna blanket while yelling at it to "keep me warm without making me sweat."
So if the dash was never the tell, what is? The usher: a sentence posted at the door of the point, gesturing at it, adding nothing to it.
- "It's important to note that"
- "The key thing to understand here is"
- "In this section, we'll explore"
- "That said, it's worth considering"
Each one tells you a point is coming and that the point matters, and then the point walks in smaller than its introduction.
Linguists have been cataloguing ushers since before the machines showed up. Ken Hyland calls the category metadiscourse and splits it in two: the interactive kind organizes the text (transitions, signposts, frame markers), and the interactional kind puts a person in it (hedges, opinions, somebody willing to be seen holding a position). The comparisons that exist keep finding the same lopsided ratio. Pham ran the head-to-head, ChatGPT research articles against human ones: signposting everywhere, self-mention and attitude markers missing. Lei Hong found post-2022 student writing drifting the same direction, more transitions, more frame markers, converging on the machine (paywalled, no free copy, and worth it if you have access). Alsharif is the adjacent case, a decade of news articles about AI showing the same organizing-heavy patterns, though it compares sections and years rather than model against human. In the direct comparisons, the model is fluent about where you are in the document and silent about what anyone in it believes.
The "I know it when I see it" feeling survives peer review, too. Shaib and colleagues built a taxonomy of slop from interviews with experts across writing, journalism, linguistics, and NLP, annotated real articles span by span, and found that gut-level slop calls line up with measurable dimensions like coherence, relevance, and vagueness. Kobak and colleagues did the vocabulary version, borrowing the excess-mortality method from epidemiology and pointing it at word frequency to catch which words spiked past their trend line after 2022. Their chart is why everyone now flinches at "delve."
It was trained to do this
To turn a raw model into a product, the lab shows human raters two drafts and asks which is better, thousands of times, and those picks train a reward model that shapes everything the product says afterward. Raters shown two options reliably pick the one that feels more familiar. Psychologists have measured that pull for decades, and the raters weren't being lazy: familiarity is what preference feels like from the inside.
Zhang and colleagues traced the whole chain and named it typicality bias. The raters' preference for familiar text gets baked into the reward model, and the model collapses toward the most typical way of saying anything at all. The most typical way to signal that a fact matters is to announce that it matters. So the model posts an usher.
This is why "sound more human, be conversational, write like a real person" keeps failing you. Typical is what the raters rewarded, so you're asking the machine for a stronger dose of the cause. It's treating your anxiety by worrying about it harder. The prompt itself is not the dead end, and the fixes below are mostly prompts: what fails is the specific request to be less typical, aimed at a machine whose sense of normal came from the raters.
Waiting for the next model fails for the same reason. NoveltyBench measured twenty leading models and found every one less diverse than a human writer, and found that larger models inside a family are often less diverse than their smaller siblings. Capability scores climb while the writing flattens, because the flattening rides in on the alignment, and a bigger model often carries more of it, not less.
The experiment
The research yields four countermeasures. Instead of listing them, I ran them on one essay, and not cleanly: the first draft ran under the string bans alone, later passes added move, structure, and referent checks as each draft failed, and the final version was revised under the full list.
The list went in first
The cheapest fix in the literature is a list that bans the constructions by name, and the best argument for it comes from a training method called Diverse Preference Optimization. DivPO attacks the collapse where it starts, building each training pair backwards from diversity: the chosen example is rare and good, the rejected example is common and mediocre. Persona diversity up 45.6%, story diversity up 74.6%, win rates held, with a little quality given back on one story task.
You'll never run DivPO. The rule inside it transfers anyway:
Reward the rare good example. Penalize the common mediocre one.
DivPO retrains the model; a list only constrains it at the keyboard, which is a real difference, and the cheap version is still worth having. The list borrows the shape of DivPO's rejected half: the common-mediocre corner, labeled by hand, for free, in an afternoon.
It goes wherever your tool keeps standing instructions, which means it works everywhere:
| Tool | Where the rules live |
|---|---|
| Claude Code | CLAUDE.md |
| Codex and most agent CLIs | AGENTS.md |
| Cursor | .cursor/rules |
| Gemini CLI | GEMINI.md |
| Warp, Devin, Copilot | the rules or knowledge panel |
| Anything with an API | the system prompt |
The version that went into the file for this essay, which you should steal and then ruin with your own additions:
- No em dashes.
- No "not X, it's Y." This is the single most recognizable construction in machine writing and it is everywhere.
- No "it's important to note," "it's worth noting," "that said," "in today's."
- No delve, robust, seamless, comprehensive, crucial. Leverage is a noun.
- No opening paragraph that describes what the piece is about to cover.
- No closing paragraph that summarizes what the piece just covered.
Then add the specific thing that made you wince last week. That entry outranks the six above it, because you caught it yourself.
Every opener arrived with odds on it
The group that named typicality bias also built the cheapest working counter, and it is one prompt change. Ask for several outputs with an explicit probability attached to each, then take one from below the top of the distribution. They report 1.6x to 2.1x diversity gains with factual accuracy held, and the gains grow with model capability.
From the second draft on, the essay's titles and openers arrived five at a time with a probability on each, and the likeliest went in the bin unread. The sampled title lost anyway, overruled by the editor, which is the human preference signal doing exactly its job. Paste the same instruction into anything: "Give me five openers for this. Put an explicit probability on each." The highest-probability opener is the one every other firm's model produced this morning.
The examples were outliers on purpose
Everybody who builds an example library collects representative samples of their writing, and a representative corpus teaches the model the average, which is the sound you came to get away from. Collect the outliers that worked instead: the email that got a reply from someone who never replies, the line that got a laugh on a call, the paragraph a client quoted back to you. Four strange successes beat forty adequate samples. The essay's source material worked the same way, a reading list of rare finds assembled over months rather than a stack of typical takes on the topic.
The voice sentences stayed human
Padmakumar and He ran a controlled co-writing experiment and traced the flattening mainly to the model's share of the text. The sentences the humans wrote themselves held their variety, with a model in the loop or without one.
Retrieval, facts, ranking, assembly: hand it all over. The opening line, the judgment call, and the ask stay yours. For the essay, the split meant the angle, the targets, and the jokes came in as a brain dump from the author, and the model worked around them.
While you are in there, leave the temperature knob alone. Turning it up buys variety by spending accuracy, which is the wrong trade in anything that names a real number or a real date, and the sampling research mostly isn't reachable through product APIs anyway.
The grep came back clean
Draft one of the essay passed every check I could automate. Grep against the banned list: zero hits. Then a person read it.
"There is a reason for that, and it turns out to be a good one."
A sentence whose entire job is to promise the next sentence. Call it the drumroll.
"Everybody can spot this shit. Almost nobody can say what it is."
Check the list: "not X, it's Y" is banned. This is the same move wearing two sentences. The seesaw.
"Here is the part almost nobody in the marketing conversation has read, and it is more human than you would guess."
A drumroll with a superiority complex.
"The honest caveat"
A section heading that names its own virtue, which is a thing honest writing never needs to do. The stance label.
"Three sentences. That is the whole difference."
This one was inside a diagram. A fragment whose only job is to inform you that the previous sentence was important. The gravity fragment.
The grep found nothing, and the read found all four moves anyway, rebuilt from words that were still legal. Shaib and colleagues, the same group behind the slop taxonomy, measured this exact mechanism at the syntax level: models reuse templates, repeating sequences of grammatical structure at rates human text never reaches, the words rotate while the template survives, 76% of those templates trace back to pre-training data, and the repetition rides straight through RLHF untouched. A list of strings polices words, and the collapse operates on shapes.
So the draft was rewritten under a list that banned the four moves by function. The rewrite failed differently. The moves were gone, and the fixes arrived as a numbered list under the heading "Four fixes, cheapest first," about the most predictable shape a fixes section can take. The taxonomy paper names lists-as-responses as a slop indicator on its first page. The editor's margin notes from that pass: "I feel like you didn't run DivPO on this at all," which was accurate, since the openers had been sampled five at a time and the outline had been sampled exactly once. Also "what humans?" beside a sentence that said "the humans," and "what the fuck does this mean" beside a metaphor with no referent, which the taxonomy files under vagueness.
In this experiment, each draft failed one level up from the last: words, then sentence moves, then structure and referents. The list grew to match:
- No drumrolls. A sentence that only promises the next sentence gets cut, and the next sentence learns to stand alone.
- No seesaws. A balanced opposite pair built for cadence is the banned construction in a costume. A contrast reporting a measured finding stays.
- No stance labels. Never announce your own honesty, fairness, or candor.
- No gravity fragments. If the point carries weight, the reader will notice without a two-word fragment saluting it.
- Sample the structure too. The outline gets the five-with-odds treatment, same as the openers. The numbered how-to list is usually the one to bin.
- Every referent resolves. "It," "the humans," and every metaphor must point at something the reader can name, or it goes.
Whether the ladder tops out at referents or my editor just stopped climbing, I can't tell from inside it.
And the check has to climb with it. Grep catches strings. Only a reader catches moves and shapes, and asking the model to catch its own typicality has the same circularity as asking it to sound more human, since its sense of normal came from the raters who caused the problem. A person reads the draft with one question per sentence: does this contain information, or does it frame information that lives somewhere else?
What a list can't reach
Everything above is containment. The collapse itself lives in the weights, put there by the preference data, and the repair at that level is a training method you'll never run. The load-bearing fix is the split: sentences the model never writes need no inspection at all. And every catch the read produces is a preference pair, the machine's sentence next to the one that replaced it, the same kind of data DivPO trains on, generated free every time you edit. Keep the diffs. They become the list, and the list becomes the voice.
Where the evidence runs out
Every corpus behind this piece is academic prose, peer review, news, or creative writing. None of it is sales email, and none of it is a pitch, where some signposting earns its place and the norms run different. Read the transfer as directional. The evidence that governs your voice is still a person who knows you, reading your draft, telling you what they hate.
The byline
A model wrote this post.
The essay in the experiment is the one you are reading. The specimens are quotes from its first draft, the margin notes are real, and the editor writing them was me. Two drafts failed, each one level up from the last, and the version in front of you was revised under the full list, then read sentence by sentence against it.
The angle is mine. The examples are mine. The irritation is mine. The swamp ass is mine, near verbatim from a margin note. The swearing is definitely mine. The model did retrieval, structure, and assembly around a brain dump of what I wanted to say and who I was annoyed at.
If it reads like a person wrote it, that's the argument.
Start here
Recognize the problem? Let's look at yours.
Thirty minutes on the loop that costs you the most hours. You leave knowing whether it can be automated and roughly what that takes.