The Signal FilesOperating in the AI EraResearch Brief

Why everyone suddenly sounds the same

Generative AI makes each writer measurably better. It makes everything written measurably more alike. The trap has receipts now.

~1,800 WordsThree Cited SourcesStop Trying To Be Invisible

Every writer in the experiment who took the machine's help got measurably better. The collection of what they wrote together got measurably more alike. That is the whole trap in two sentences, and since 2024 it has a peer-reviewed citation. The tool that lifts you as an individual is the same tool that averages you as a crowd. Everyone hears the first half of that sentence. Almost nobody prices in the second.

This is a research brief: one peer-reviewed experiment, one web-scale measurement, one preprint. Each is stated with its exact scope, because the three pieces of evidence carry different weights. Together they describe a market condition forming in real time. At the end we say what we think it means for anyone who competes for attention, which is everyone.

Section OneThe experiment: better alone, alike together

In July 2024, Science Advances published a causal experiment by Anil Doshi (University College London) and Oliver Hauser (University of Exeter). The setup was clean: people wrote short stories, and some of them could draw on story ideas generated by AI. Then independent evaluators rated the results.

Science Advances, Doshi & Hauser, "Generative AI enhances individual creativity but reduces the collective diversity of novel content," July 12, 2024.

The first finding is the one the tool vendors will quote. Writers with AI-generated ideas produced stories judged more creative, and the lift was largest for the less creative writers: the technology worked hardest for the people who needed it most. The second finding is the one that should reorganize your content strategy: the AI-assisted stories were more similar to each other than the stories written without help. Individual lift and collective convergence, in the same dataset, from the same tool, at the same time.

Scope, stated plainly: this is a controlled experiment on short-story writing. One task, one setting. It does not by itself prove that the whole web is homogenizing. What it proves is causality: the convergence is produced by the tool, not merely correlated with it. Hauser named the incentive spiral himself: "if individual writers find out that their generative AI-inspired writing is evaluated as more creative, they have an incentive to use generative AI more in the future, but by doing so the collective novelty of stories may be reduced further."

The mechanismThis is a textbook social dilemma. Each writer's choice to use the model is individually rational: the ratings prove it pays. But every writer who accepts the model's suggestions is sampling from the same underlying distribution of ideas. The gain is private; the sameness is collective. Nobody defects, because defecting means writing worse, alone.

In 2024 that was a lab result with a warning attached. The next two measurements show what the warning looks like at web scale.

The tool that makes you faster is the same tool making you average.

Section TwoHalf the web, give or take

The lab result would matter less if adoption were marginal. It is not. In October 2025, Graphite (an SEO firm, so a single vendor with an interest in the topic, which is why we attribute rather than assert) published an analysis of tens of thousands of web articles sampled from Common Crawl, spanning January 2020 to May 2025, classified with an AI detector carrying a stated false-positive rate of 4.2 percent and false-negative rate of 0.6 percent. Its finding: by May 2025, roughly half of newly published web articles were AI-generated. Press coverage rounded it to 52 percent; Graphite's own framing stayed at roughly half, a dead heat.

Graphite, "More Articles Are Now Created by AI Than Humans," October 2025 (Common Crawl sample, Jan 2020–May 2025; single-vendor analysis).

One detector, one vendor, one sample: the exact decimal is arguable, and Graphite itself has since re-run the study with three detectors. The direction is not arguable. In 2020 the AI share of new articles was a rounding error. Five years later a serious measurement puts it at the coin-flip line. Whatever the true figure is this quarter, the supply of machine-written text is no longer a niche; it is the market's default output.

Two honesty notes belong next to the number. First, the window: the measurement ends in May 2025 and was published in October. The share has had time to move since, and Graphite's own follow-up exists precisely because a single snapshot is not a trend line. Second, the unit: the study counts published articles, not read ones. It says half the supply is machine-made, not half of what people actually consume. Both caveats trim the claim. Neither reverses it.

Section ThreeSemantic contraction

What does half a web of model output actually sound like? A preprint from researchers at Stanford, Imperial College London, and the Internet Archive measured it. That preprint is not yet peer-reviewed, so treat every figure here as provisional. AI-generated websites showed pairwise semantic similarity roughly 33 percent higher than human-written ones: machine sites resemble each other far more than human sites do. The same study estimates that by mid-2025 about 35 percent of newly published websites were AI-generated or AI-assisted.

arXiv preprint 2604.26965 (Stanford / Imperial College London / Internet Archive), "The Impact of AI-Generated Text on the Internet" (preprint, not yet peer-reviewed).

The authors coined two terms worth keeping. Semantic contraction: the space of things actually being said shrinks toward a common center, even as the volume of text explodes. And artificial positivity: the machine register defaults to the agreeable, upbeat middle: no rough edges, no risk, no position. More text, fewer meanings, one mood.

The web is not just filling with machine text. It is contracting toward one voice.

Section FourSameness is a market condition, not a style problem

Put the three findings in a row. A peer-reviewed experiment shows the tool causally converges its users' output. A vendor's web-scale measurement puts machine text at roughly half of everything newly published. A preprint measures the result: machine content is a third more alike, and drifting toward one agreeable register. Each piece has its caveats. All three point the same way.

Now run the scarcity math. When competent, pleasant, on-topic prose can be produced in unlimited quantity by anyone, its market price falls toward zero. That is not a tragedy; it is a repricing. What becomes scarce, and therefore valuable, is exactly what the model cannot generate on demand: a position someone is willing to defend, an observation drawn from experience the model has never seen, a voice that could not have come from anywhere else. Distinctiveness used to be a taste preference. The data above makes it a competitive asset with a rising price.

You can test the contraction against your own market in an afternoon. Open the websites of ten providers in any category (agencies, consultancies, software firms) and read the first paragraph of each. If the category has followed the trend, you will find the same promises in the same order in the same register: tailored solutions, passionate teams, measurable results. That paragraph now costs nothing to produce, which is precisely why it buys nothing.

Why it mattersThe audience for your writing is increasingly a machine that compresses. AI assistants and answer engines read many sources and return one synthesis. Here is our working inference from that mechanic (reasoning, not a cited measurement): when ten sources say the same thing in the same register, the summary needs at most one of them, and the other nine contributed training mass, not presence. The content that survives compression is the content that adds something the pile did not already contain. Sameness is not just boring. In a summarized web, sameness is invisibility.

Section FiveUsing the tool without taking the trap

The answer is not abstinence. The experiment's first finding is real: the tool genuinely lifts individual output, most for those who start weakest, and we run our own studio on AI agents. Refusing the leverage would be theater. The answer is deciding which layer of the work you delegate. Three rules we hold ourselves to:

Feed it your material, not its own. The model can average the entire web; it cannot average your client conversations, your numbers, your failures. Writing that starts from what only you have seen cannot converge on the common center, because the common center does not contain it.

Keep the position-taking human. Let the machine draft, structure, translate, tighten. The claim itself is the one component whose value is rising: what you believe, what you would bet on, what you think everyone else has wrong. Delegating it is selling the asset to buy the commodity.

Audit against the sea. Before publishing, ask the one-line test: could a competitor sign this without changing a word? If yes, the piece is contributing to the pile, not standing out of it. Rewrite until the answer is no.

In ClosingDistinct, and legible

Our whole thesis is that machines can only act on what they can clearly read, outside your company and now inside it. The homogenization data adds the sharper corollary: being readable is necessary but no longer sufficient. In a sea of semantically contracted, artificially positive text, the brands that get chosen, by readers and by the models summarizing for them, are the ones that are both distinct and legible. Distinctiveness is the signal. Legibility is the transmission. One without the other is either noise or silence.

The trap is real and it has numbers now. The tool that makes you faster is making everyone the same, which means the point of view you were told to sand off is the most defensible thing you own. Keep it. Sharpen it. Make it machine-readable.

If you want to know how visible your brand currently is to the machines doing the choosing, you can measure it at /signal-index/, or write to us at /contact/.

Figure 01 · The Trade
What the tool gives each writer, and what it takes from all of them
~50%
Share of newly published web articles that were AI-generated by May 2025. Graphite's Common Crawl analysis: one vendor, one detector, date-stamped.
+33%
Higher semantic similarity among AI-generated sites versus human-written ones. Stanford/Imperial/Internet Archive preprint, not yet peer-reviewed.
Two measurements, two confidence levels, one direction: more machine text, and machine text that sounds like itself. Sources: Graphite (Oct 2025); arXiv:2604.26965 (preprint).
Stop trying to be invisible.

Sources

  1. Doshi, A. & Hauser, O., "Generative AI enhances individual creativity but reduces the collective diversity of novel content," Science Advances, Vol. 10, Issue 28, July 12, 2024. science.org/doi/10.1126/sciadv.adn5290
  2. Graphite, "More Articles Are Now Created by AI Than Humans," October 2025 (Common Crawl sample, Jan 2020–May 2025; single-vendor analysis, since updated with additional detectors). graphite.io/five-percent/more-articles-are-now-created-by-ai-than-humans
  3. "The Impact of AI-Generated Text on the Internet," arXiv:2604.26965 (Stanford / Imperial College London / Internet Archive; preprint, not yet peer-reviewed). arxiv.org/abs/2604.26965

The Signal Index

How clearly can the AI era see you?

A free, transparent score of how AI and search find, understand and recommend you. Instant, from your domain.

Get your Signal Index →

The Signal Files

Field notes on visibility, in your inbox.

The research behind how brands get seen now. The Signal Files, the moment they publish. No noise.

Double opt-in. Unsubscribe anytime.