The Signal FilesAI VisibilityThesis

Clarity beats loudness

The machine does not choose the brand that shouts. It chooses the brand it can read. Two peer-reviewed teams just proved it, two years apart.

~1,800 WordsTwo Cited SourcesStop Trying To Be Invisible

Somewhere in a Princeton dataset sits a result that should retire a century of marketing instinct. When a generative engine assembles an answer, it does not favor the source that shouts. Keyword stuffing, the oldest trick of loudness, moved visibility backwards. What lifted a source's visibility, by up to 40 percent, was almost embarrassingly plain: quotations, statistics, cited sources. Material a machine can extract, check, and attribute. We have made one argument since the day this studio opened: buyers now ask the machine before they ask anyone, and the machine recommends what it can unambiguously resolve. The technical field keeps rediscovering that thesis from its own side. Twice now. In peer review. This file lays out the receipts.

Section OneThe thesis, before the receipts

First, the distinction, because the entire argument rests on it. Loudness is everything a brand does to occupy attention: volume, frequency, adjectives, keyword density, superlatives, spend. Clarity is whether a machine reading your public footprint can resolve you: which entity you are, what exactly you do, for whom, on what evidence. Loudness is addressed to a human scrolling past. Clarity is addressed to a model that has to decide, in the middle of composing an answer, which sources it can safely build on and cite.

That decision is the whole game. A model synthesizing an answer must pick a handful of sources out of many, extract what they claim, and stand behind the result. Ambiguity is expensive at every step: an entity it cannot resolve, claims it cannot verify, text it cannot cleanly quote. Our thesis makes two testable predictions. One: tactics that add extractable substance should lift visibility. Two: tactics that add noise should fail, and fail harder as more competitors adopt them. Both predictions have now been confirmed independently, by two research teams, in two different years, at two of the field's most serious venues.

The model does not choose the brand that shouts. It chooses the brand it can resolve.

Section TwoPrinceton, 2024: the anatomy of being chosen

The first team, from Princeton and collaborators, built the field's founding benchmark and published it at KDD 2024. They ran 10,000 queries, drawn from nine source datasets across multiple domains, through generative engines, and measured how visible each source was in the generated answers, weighting not just whether a source was cited but how prominently. Then they systematically rewrote sources with different optimization tactics and measured what actually changed.

Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan & Deshpande (IIT Delhi, Princeton University et al.), "GEO: Generative Engine Optimization," KDD 2024, DOI 10.1145/3637528.3671900 (arXiv:2311.09735).

The abstract's headline number: the right changes "can boost visibility by up to 40%." And the anatomy of those changes is the story. Adding quotations lifted visibility by roughly 40 percent, the strongest single tactic in the benchmark. Adding statistics lifted it by roughly a third. Adding cited sources lifted it by roughly 28 percent; even plain fluency editing, making the text easier to parse, helped by roughly 29 percent. And keyword stuffing, the tactic that built an entire industry, was negative. It made sources less visible than doing nothing.

Why these tactics winLook at what the winners have in common. A quotation is a pre-packaged, attributable unit the model can lift into its answer verbatim. A statistic is a claim with a number attached, checkable and quotable. A citation is provenance the model can pass along to its own reader. Every winning tactic hands the machine something it can extract and stand behind. Every one of them is a form of clarity. Keyword stuffing hands the machine nothing but noise around the signal, and the machine, unlike a 2010 search crawler, can tell.

Section ThreeNeurIPS 2025: the tricks don't survive contact

One study is a finding. The second, from an entirely independent team, is the pattern. In 2025 the field had moved on to "conversational SEO," a wave of tactics promising to make large language models recommend your product: injected instructions, persuasive rewrites, formatting tricks. A team publishing at NeurIPS 2025 built C-SEO Bench to test whether any of it works, across two tasks and three domains each, and, crucially, in a multi-actor protocol that simulates what happens when several competitors adopt the same trick at once.

Puerto, H., Gubri, M., Green, T., Oh, S. J., Yun, S., "C-SEO Bench: Does Conversational SEO Work?," NeurIPS 2025 Datasets & Benchmarks (arXiv:2506.11097).

The verdict is unusually blunt for a benchmark paper. Most current C-SEO methods were found "largely ineffective," and frequently negative, they hurt the very ranking they were sold to improve. What did keep working was the unglamorous thing: traditional SEO, improving the source's actual relevance so it earns its place in the model's context. The paper's second finding is the one we would frame in stone: "as we increase the number of C-SEO adopters, the overall gains decrease." The tricks are congested. Every additional adopter of the same trick shrinks the payoff for all of them.

Loudness is zero-sum. Clarity compounds.

That congestion result deserves a slow read, because it is the economic signature that separates the two strategies. A trick works by grabbing a fixed quantity of the model's attention, so every competitor running the same trick is bidding against you in the same auction; the equilibrium is everyone spending more to stand exactly where they stood. Substance does not congest the same way. A source that is genuinely relevant, verifiable, and cleanly structured is not fighting for a slot; it is the reason the slot exists. This is our reading, stated as our reading. The benchmark measured the congestion; the interpretation is the thesis this file argues.

Section FourEvery mechanic reduces to legibility

Put the two studies side by side and every effective mechanic collapses into three properties. Entity resolution: the machine must be able to determine which thing you are, one name, one identity, across every surface it reads. Consistency: claims that agree with each other across sources raise confidence; contradictions lower it, and a low-confidence source is a risk the model routes around. Extractable provenance: quotations, statistics, citations, the units a model can lift, verify, and attribute, which is precisely the list Princeton found at the top of its benchmark.

Now run the inversion. A brand that is loud but not legible, heavy on adjectives, light on checkable claims, inconsistent across its surfaces, does not get rejected by the machine. Something quieter happens: it gets averaged out. Its unverifiable self-descriptions dissolve into the category's background statistics, and the model describes it, if at all, in the same generic language it uses for every interchangeable competitor. The loud brand does not lose the argument. It never enters it.

Why it mattersTwo teams, two years apart, no shared authors, different engines, different tactics under test, and the results point the same direction: substance a machine can extract wins, noise loses, and noise loses harder as it spreads. Independent replication at peer-reviewed venues is the strongest form of evidence this young field can currently produce. The clarity thesis stopped being a branding opinion. It is now the measured behavior of the machines doing the choosing.

Section FiveWhat these studies cannot tell you

The honest counterweight, because our own rule is to name uncertainty in the text. Both results are controlled benchmarks, not market measurements: fixed query sets, simulated competitive settings, visibility metrics rather than revenue. A 40 percent visibility lift inside a benchmark is not a promise of 40 percent more customers; the multi-actor protocol is a simulation of competition, not competition. And the engines themselves keep changing, so any specific number has a shelf life.

What has a longer shelf life is the direction, because it is the one thing the two teams, and the underlying mechanics, agree on. Models reward what they can extract, verify, and attribute. They punish, or ignore, what they cannot. Every future engine that has to choose sources and stand behind an answer faces the same constraint. Betting on clarity is not betting on a benchmark. It is betting on the geometry of the problem.

In ClosingLegible, or averaged out

The through-line of everything we publish is one claim: machines can only act on what they can clearly read, outside your company, on the surfaces where buyers now ask their questions, and increasingly inside it, where your own agents do the work. The AI-visibility field keeps arriving at that claim from the technical side, tactic by tactic, benchmark by benchmark. Princeton showed what clarity is worth: up to 40 percent. NeurIPS showed what loudness is worth: little, often less than nothing, and less every time a competitor copies you.

A brand that is clear to the machine gets chosen. A brand that is merely loud gets averaged out. That is not our slogan restated; it is the measured behavior of the systems that now sit between you and your next customer. The unglamorous work, one resolvable identity, consistent claims, checkable evidence, quotable substance, is the whole strategy. It always was; now it is peer-reviewed.

If you want to know how legible your brand already is to the machines, the Signal Index measures it, or write to us.

Figure 01 · Two Verdicts, One Direction
What the machine pays for, and what it penalizes
+40%
The price of clarity, paid to you. Adding quotations lifted source visibility by roughly 40 percent, statistics by roughly a third, citations by roughly 28 percent, in Princeton's 10,000-query benchmark (KDD 2024).
< 0
The price of loudness, paid by you. Keyword stuffing turned negative in the same benchmark; at NeurIPS 2025, most conversational-SEO tricks were ineffective or negative, and congested as adopters multiplied.
Two peer-reviewed teams, two years apart, no shared authors, one direction. Sources: GEO (KDD 2024); C-SEO Bench (NeurIPS 2025). Benchmark findings, not market guarantees.
Stop trying to be invisible.

Sources

  1. Aggarwal, P. et al., "GEO: Generative Engine Optimization," Proceedings of KDD 2024, DOI 10.1145/3637528.3671900. arXiv:2311.09735
  2. Puerto, H., Gubri, M., Green, T., Oh, S. J., Yun, S., "C-SEO Bench: Does Conversational SEO Work?," NeurIPS 2025 Datasets & Benchmarks. arXiv:2506.11097

The Signal Index

How clearly can the AI era see you?

A free, transparent score of how AI and search find, understand and recommend you. Instant, from your domain.

Get your Signal Index →

The Signal Files

Field notes on visibility, in your inbox.

The research behind how brands get seen now. The Signal Files, the moment they publish. No noise.

Double opt-in. Unsubscribe anytime.