AI didn’t invent the em dash - it just made sameness easier to see.

At some point, the em dash became evidence.

Use one in an email, article or LinkedIn post now, and there’s a decent chance somebody will joke that ChatGPT wrote it. The Oxford comma has started looking a little suspicious too. So have certain words (Delve is probably never recovering).

And I understand why, because once you’ve read enough AI-generated text, you start recognizing the rhythms. The tidy introduction. The evenly sized sections. The sentence that explains what the previous sentence already explained. The conclusion where everything comes together in a tidy little package.

It’s smooth.

And it all sounds the same.

That feeling isn’t entirely anecdotal. There’s some pretty compelling research showing that generative AI isn’t simply adding more content to the web. It’s changing the language of the web itself.

 

The em dash isn’t innocent, exactly

In August 2026, Pew Research Center analyzed 490,000 English-language webpages collected from the Common Crawl archive between 2021 and 2026.


A few of the changes were hard to miss:

  • AI page volume: About 10% of all webpages in Pew’s July 2026 sample showed significant signs of AI authorship or editing. Among pages published after ChatGPT became publicly available, that rose to more than one-third.

  • Punctuation shifts: Em-dash use nearly doubled, from 5.79 to 11.19 instances per 10,000 words. Oxford commas increased by 63%.

  • Vocabulary shifts: Words disproportionately associated with AI writing, including delve, pivotal, intricate, interplay, showcase, and testament, more than doubled in frequency. Even the “it’s not just X, it’s Y” construction nearly tripled.


Pew used an AI-detection model for the analysis, so those numbers shouldn’t be treated as a precise headcount of robot-written pages. The researchers also caution that individual habits like using an em dash or Oxford comma don’t prove that AI wrote something.

So the people side-eyeing em dashes aren’t hallucinating the pattern, but they’re giving the punctuation too much credit.

 

The web really is getting flatter

This is the part I find both fascinating and concerning.

A 2026 study published in Nature Human Behaviour looked at more than 880,000 pieces of writing. The researchers wanted to see what happened when large language models (LLMs) rewrote human text.

With the LLMs at the helm, the differences in writing complexity shrank by 21% to 50%, depending on the dataset and model.

In other words, AI made different writers sound more like each other. The information was still there, but some of the personality had been smoothed out.

Another study found something similar when researchers looked at ideas instead of writing style. They compared 2,200 human and GPT-written college admissions essays and looked at how many genuinely new ideas appeared as they added more essays to the pile.

The human essays introduced new ideas more often than the AI-written ones. And the more essays the researchers compared, the more obvious the difference became. Even prompting GPT in different ways to make its writing more creative didn’t get rid of it.

That’s an important distinction. One AI-written essay can be good. It might even be better written than the human version beside it. But put 100 AI-written essays together, and you start to see how often they arrive at similar ideas, using similar structures, and similar ways of expressing them.

Now multiply that across websites, social posts, newsletters, product descriptions and articles.

The problem isn’t that the web becomes badly written.

It becomes averagely written.

Good grammar isn’t the problem

To be clear, I’m not advocating for sloppy writing.

Typos aren’t proof of humanity. Bad grammar isn’t a brand voice. Putting an unnecessary lowercase letter in a heading isn’t an edgy choice. And AI can produce writing people genuinely like.

A 2026 study involving 1,600+ adults compared human-written short stories with stories generated by ChatGPT. Participants actually rated the AI stories as more absorbing and higher quality.

But humans are complex creatures, so something interesting started happening.

People rated stories more favourably when they were told a human had written them, regardless of who actually had.

Another experiment involving 700+ people found a similar effect in science journalism. Participants all read GPT-generated material, but people who were told the article was AI-written rated both the article and its source as less credible than people who believed a human wrote it.

That isn’t proof that readers prefer awkward sentences, but it does suggest that our judgment of writing now involves something beyond quality. We’re looking for authorship. Effort. Intent. Some indication that an actual person is in there.

And that is where human writing has something AI still struggles with: A little anarchy.

 

Human writing doesn’t always behave itself

Real people aren’t probability engines.

We change our minds halfway through explaining something. We spend 400 words on a sidequest about something we care about and two sentences on the thing we’re supposed to care about.

We have pet words. Catch-phrases. Strange comparisons. Regional expressions. Opinions that aren’t perfectly balanced.

Sometimes a sentence is too long because that’s how the thought came out.

Then this one isn’t.

A writer might interrupt herself to make a point she hadn’t planned on making. She might decide the technically correct wording sounds ridiculous. She might know an oddly specific detail because she spent six years working with the software, grew up in the town, talked to the customer, made the mistake, or watched a specific tv series fifteen times.

None of that makes the writing automatically good, but it does give it shape.

Generative AI, particularly when it’s being asked to “improve,” “polish,” or “make more professional” something, has a strong incentive to smooth those shapes out.

The awkward transition disappears. Sentences settle into familiar rhythms. Paragraphs become the same size. Ideas acquire headings. Strong opinions acquire caveats.

Eventually everything is a nice, smooth circle.

 

And that’s why the AI tells keep changing

We seem to want a forensic test.

Em dash? AI.

“Delve”? AI.

Three things in a list? You’re in jail now.

But that approach was always going to fail.

Pew’s research is useful precisely because it looks across hundreds of thousands of pages. The researchers explicitly caution that an em dash, Oxford comma or particular phrase tells us very little about the authorship of an individual document. These characteristics become meaningful when researchers look at patterns across large collections of writing.

So an em dash isn’t an AI fingerprint - it’s something AI reaches for more consistently than humans do.

And once everyone knows that, people start removing em dashes from AI output because they don’t want to be accused of using AI (let’s talk about that more later).

Then the language models get updated. Prompts change. Detectors adapt. Human writers start avoiding perfectly normal punctuation, and now the definition of “sounds human” moves again.

Which is absurd, actually, because we’ve reached a point where a perfectly competent human writer can make their work seem more authentic by making it less polished.

 

We probably shouldn’t make writing worse

AI is useful.

It can sort information, summarize large amounts of material, help with repetitive work and make complicated tools more accessible. Pretending none of that has value doesn’t get us anywhere.

I’m much more interested in what happens when we hand over the parts of our work that require judgment.

Because the biggest risk I see isn’t that AI will produce terrible writing. It does. The real risk is that AI writing idles on ‘acceptable’.

And 'acceptable' scales.

 

AI is fast.

Much faster than people. Businesses can suddenly publish 100 competent articles a year instead of ten, and every service they offer can have its own page. Every possible keyword can have another page. Soon, everyone is explaining the same subject in slightly rearranged versions of the same language.

Google’s current spam policies make an important distinction here too.

Its concern isn’t simply whether AI was involved. Its scaled-content policy targets large volumes of unoriginal content that exist specifically to boost search rankings and provide little or no additional value, regardless of how the content was produced.

Google specifically lists using generative AI to produce many pages without adding value for users as one example of scaled content abuse.

Generative AI is one way to create a problem very quickly, because the tools make ‘acceptable’ infinitely reproducible.

 

The rough edges may become more valuable

For years, web writing has been moving toward polish.

Make it clearer. Shorter. More professional. Easier to scan. Fix the grammar. Follow the template. Use the proven structure. It’s boring, sensible advice.

But now generative AI can produce polished, organized, grammatically competent web copy almost instantly. And when polish is cheap, everything gets polished.

And when polished becomes the standard, specificity starts mattering more.

So does experience. Point of view. A weirdly good example. A sentence somebody else wouldn’t have written. An opinion that hasn’t been neutralized and rounded into something nobody could disagree with.

Even a slightly strange turn of phrase conveys something important:

There is a person here.

Now, I don’t think readers are consciously sitting around scoring websites and picking service providers based on linguistic chaos. But we’re all becoming more familiar with the look of generated language. The smoother, more predictable the writing we encounter, the more visible the deviations become.

There’s an irony in all of this.

AI learned to write by absorbing an enormous amount of human language. Now that its version of that language is being fed back into the web at scale, some completely ordinary human writing habits have started looking suspicious.

Like the poor em dash.

The em dash didn’t change, but the environment around it did. And now we treat the em dash as a horseman of the AI-pocalypse, but this villain is actually the victim.

And if generative AI continues pulling web writing toward an ‘acceptable’ middle, the most valuable writing isn’t the writing with the cleanest edges...

It’s the writing that’s just a little bit weird.

 

Sources and further reading

Pew Research Center:How Much of the Internet Is Written With AI? August 20, 2026. Read the Pew Research Center analysis

Sourati et al., Nature Human Behaviour:The shrinking landscape of linguistic diversity in the age of large language models. August 24, 2026. Read the study in Nature Human Behaviour

Moon, Green and Kushlev:Homogenizing effect of large language models (LLMs) on creative diversity: An empirical comparison of human and ChatGPT writing. 2025. Read the study on ScienceDirect

Sears and Weisberg:Bot or not: Can people tell the difference between stories written by a human or by an AI system? 2026. Read the study from Cambridge University Press

Henestrosa and Kimmerle:The Effects of Assumed AI vs. Human Authorship on the Perception of a GPT-Generated Text. 2024. Read the credibility study

Google Search Central:Spam Policies for Google Web Search and Google Search’s guidance on generative AI content on your website.
Read Google’s scaled content abuse policy
Read Google’s generative AI content guidance