What actually works in AEO, and what does not

Summary

Very little in answer engine optimisation has controlled evidence behind it: one peer-reviewed study, a few large descriptive datasets, and a lot of confident assertion. Three things have real support — being readable without JavaScript, citing sources and statistics inside the text, and consistent listings. Three are fiction: special AI Overview markup, guaranteed placement, and llms.txt as a citation strategy. This grades every common tactic by the strength of the evidence, names the falsification test for each position I hold, and gives the five things actually worth doing.

Published

Very little in answer engine optimisation has controlled evidence behind it. One peer-reviewed study, a handful of large descriptive datasets, and a great deal of confident assertion. That is not an argument for doing nothing — the mechanism is real and the buyer behaviour is real — but it is an argument for knowing which claims you are paying for.

This article grades the tactics being sold in this category by the strength of the evidence behind each one, names the three that are outright fiction, and says plainly where the honest answer is "nobody knows yet."

Is the demand for this real, or manufactured by agencies?

The buyer-side signal is the stronger of the two, which is worth saying because the supply-side signal is what usually gets quoted.

On the buyer side: G2 surveyed 1,076 B2B decision-makers and reported in April 2026 that 51% of B2B software buyers now start research with an AI chatbot more often than with Google. On the supply side: AgencyAnalytics' 2026 Benchmarks Report, based on 494 agency professionals, found 66% reporting client demand for answer engine optimisation — the most requested new service that year.

Read those as different things. The first is behaviour. The second is agencies reporting that clients are asking, which tells you the category is being talked about, not that it works. Anyone quoting the second as proof of the first is doing something sloppy.

What has actual evidence behind it?

Three things, in descending order of how much weight the evidence bears.

Being readable without JavaScript — direct observation

Vercel, with research partner MERJ, analysed AI crawler traffic in December 2024 and found that the major AI crawlers do not execute JavaScript: GPTBot fetched JavaScript files 11.5% of the time and ClaudeBot 23.8%, and neither ran them, with Gemini and Applebot the exceptions because they reuse existing rendering infrastructure.

This is the strongest evidence in the field because it is not a correlation study — it is a direct measurement of what crawlers did. It is also not really an optimisation. It is a prerequisite, in the same way that having a door is not a retail strategy.

Citing sources, statistics and quotations — one controlled study

Pranjal Aggarwal and colleagues at IIT Delhi and Princeton University tested nine optimisation methods across a 10,000-query benchmark in "GEO: Generative Engine Optimization", published at ACM SIGKDD 2024. Adding quotations, adding statistics and citing sources were the three best performers, at a 30–40% relative improvement on the paper's Position-Adjusted Word Count metric.

What that is: a controlled, peer-reviewed result on a defined benchmark. What it is not: a revenue outcome, a guarantee, or a finding that transfers uniformly — the authors explicitly found effectiveness varied by domain. And it holds only if the statistics are real, because engines cross-reference and an untraceable figure undermines the passage carrying it.

Consistent listings and profiles — strong description, weak causation

Yext analysed 6.8 million citations across 1.6 million AI answers and found in October 2025 that 86% came from sources brands can control or influence, split roughly 44% first-party websites and 42% listings and profiles.

This is where careful reading earns its keep. The study describes where citations came from. It does not demonstrate that creating another listing produces a citation. Treating a descriptive split as a causal instruction is exactly how a company ends up submitting itself to 200 directories and manufacturing 200 slightly different descriptions of itself — which makes the entity harder to resolve, not easier.

What is being sold that is simply not real?

Three claims, each disprovable rather than merely doubtful.

"Special markup for AI Overviews"

There is none. Google's documentation on AI features in Search states there is no additional markup or special optimisation required for AI Overviews and AI Mode, and that ordinary discoverability rules apply. There is no AI-Overviews schema type. A vendor selling one is either not reading the documentation of the platform they claim expertise in, or is counting on you not to.

"Guaranteed placement in AI answers"

Nobody controls source selection at the moment an answer is generated. Retrieval behaviour changes without notice, the same prompt returns different answers to different users on the same day, and no commercial relationship with any provider exists that would allow a placement to be promised. This is the single most efficient disqualifying question you can ask a vendor, and the reasoning behind it is set out in what to ask an AEO agency before you hire one.

"llms.txt will get you cited"

llms.txt is a proposed convention for a Markdown file at a domain's root summarising a site for AI systems. It is a sensible idea and it costs nothing to publish. What it does not have is any published commitment from a major provider to read it in production search — Anthropic's own crawler documentation describes robots.txt controls for its three crawlers and does not mention llms.txt, and Google's AI-features guidance describes no such mechanism either. Publish it as a tidy site summary. Do not pay anyone for it as a citation strategy.

What is plausible but genuinely untested?

The honest middle category, which most vendor content collapses into either "proven" or "myth" depending on whether they sell it.

Common practices with mechanistic logic but no controlled evidence
Practice Why it is plausible What is missing
Self-contained passages of a specific length Follows directly from chunk-level retrieval No engine publishes its chunking parameters. Any specific word count quoted as optimal is inferred, not known
Question-shaped headings The heading travels with the chunk and matches query phrasing No isolated study of heading form; bundled into general structure advice
Publishing more content More passages means more retrieval surface Cuts both ways — thin content dilutes an entity and is discounted by ranking systems that still feed retrieval
Being mentioned on high-authority third-party sites Consistent with the Yext controllable-source split Descriptive only; no experiment separating the mention from the authority that produced it
Entity markup and a corroborated identity Citation is an act of attribution, so resolvable identity is a precondition Mechanistically strong, never measured in isolation

I put four of the five into practice anyway, including on this site, because the mechanism argument is sound and the cost is editorial discipline rather than budget. But "I do this because the mechanism implies it" is a different sentence from "this is proven", and anyone charging you should be able to tell them apart.

How do you spot a renamed SEO package?

Ask what gets reported. This separates the field faster than any other question, because a repackaged offering cannot produce the artefacts a real one does.

  1. A written prompt set, given to you before any work starts. If measurement begins with prompts chosen after the fact, the numbers cannot be compared to anything.
  2. Per-engine reporting, never blended. ChatGPT, Perplexity, Google's AI Overviews, Gemini and Claude retrieve differently. A single averaged "AI visibility score" hides the only pattern worth acting on.
  3. The specific source that earned each citation. "You were mentioned 14 times" is not a finding. "You were cited from your Crunchbase profile, not your website" is.
  4. A stated position on non-determinism. Anyone presenting one screenshot as a measurement has either not run this repeatedly or is hoping you have not.
  5. Something they will not promise. A vendor with no stated limits has not thought about the mechanism.

Where does the evidence run out entirely?

Two places worth naming, because pretending otherwise is how this category loses credibility.

Attribution to revenue. There is no clean chain from a citation in an AI answer to a closed deal. Referral headers are inconsistent, many buyers read an answer and then navigate directly, and Cloudflare's crawl-to-refer analysis from July 2025 shows how lopsided the ratio of crawling to referred traffic has become for some providers. Anyone showing you a tidy attribution model for this channel has built it out of assumptions. The measurement approach I use, and its limits, is in how to measure AI visibility honestly.

How long any of this lasts. Every finding above has a date on it for a reason. The Ahrefs overlap figure moved from about 76% to about 38% in eight months. A tactic validated in 2024 is a hypothesis in 2026. Treat anyone presenting a fixed playbook as someone who stopped reading.

What would change my mind?

A position you cannot state the falsification test for is not a position, it is a preference. So, specifically:

  • I would drop the passage-structure advice if a controlled study varied only section length and self-containment, held everything else constant, and found no effect on citation rate across engines. Nobody has run that study.
  • I would take llms.txt seriously the moment any major provider documents reading it in production search. That is a single sentence in someone's crawler documentation, and it has not appeared.
  • I would stop treating listings as a priority if a study tracked companies before and after fixing profile consistency and found citation rates unmoved. The Yext data cannot settle this, because it looks at outcomes rather than at changes.
  • I would revise the whole ordering if the engines converge on rendering JavaScript. The Vercel finding is a snapshot of December 2024 behaviour, not a law of nature, and Gemini and Applebot already render.

Each of those is a specific, checkable condition. If a vendor cannot answer the same question about their own methodology, you are not buying a method — you are buying confidence.

What pricing signals are worth noticing?

Two, and they are structural rather than about the number itself.

Being quoted separately for SEO, AEO and GEO. These are one discipline with one sequence — the search layer has to work before the passage layer matters at all. Three line items for one piece of work is a pricing decision presented as a technical one.

Performance pricing tied to citations. It sounds aligned and it is not, because the seller does not control the outcome and both parties know it. In practice it pushes effort toward whatever moves most easily — high-volume low-value prompts, mention counts rather than citation rates — which is precisely the reporting you should be trying to get away from.

So what is worth doing?

The short list, in order, and it is shorter than the market implies:

  1. Make the content exist in the initial HTML. Highest evidence, lowest cost, and it gates everything else.
  2. Make it unambiguous which company you are. One canonical description, consistent naming, structured data with a stable identifier, real profile links.
  3. Write in self-contained, answer-first passages with real sources named inline. The one tactic with controlled evidence, and it improves the page for human readers too.
  4. Fix the profiles you already have before creating new ones. Consistency is the mechanism; volume is not.
  5. Measure against a fixed prompt set, per engine, repeatedly. Without this you cannot tell whether any of the above worked.

Everything else in the category is either downstream of those five or unevidenced. That is a less impressive list than most proposals contain, which is rather the point.

What to do about your own situation

Before buying anything, find out empirically whether the engines your buyers use currently name you and which sources they cite when they name someone else. If the answer is that you are already well cited, the correct decision is to spend the money elsewhere.

The service this belongs to is getting found when buyers ask AI, and the underlying mechanism is set out in how AI engines actually choose which sources to cite.

Run a free AI visibility check

Frequently asked questions

Does answer engine optimisation actually work?

The mechanism is real and parts of it are evidenced, but the field has far less controlled research than its marketing implies. The strongest evidence is Vercel's December 2024 crawler analysis showing major AI crawlers do not execute JavaScript, and the Aggarwal et al. KDD 2024 study showing citations, statistics and quotations improved visibility 30 to 40% on its benchmark metric. Most other tactics are mechanistically plausible and untested.

Is there special markup or schema for AI Overviews?

No. Google's documentation on AI features in Search states there is no additional markup or special optimisation required, and that ordinary discoverability rules apply. There is no AI-Overviews schema type. A vendor selling one is either not reading the platform documentation they claim expertise in, or counting on you not to.

Does llms.txt help you get cited?

There is no published commitment from any major provider to read it in production search. Anthropic's own crawler documentation describes robots.txt controls for its three crawlers without mentioning llms.txt, and Google's AI-features guidance describes no such mechanism. Publish it as a tidy site summary if you like — it costs nothing — but do not pay for it as a citation strategy.

Can anyone guarantee placement in an AI answer?

No. Nobody controls source selection at the moment an answer is generated, retrieval behaviour changes without notice, the same prompt returns different answers to different users on the same day, and no commercial relationship exists with any provider that would allow a placement to be sold. It is the single most efficient disqualifying question to ask a vendor.

How do I tell a real AEO practice from a renamed SEO package?

Ask what gets reported. A real engagement hands you a written prompt set before work starts, reports per engine rather than as one blended score, names the specific source that earned each citation, states how it handles non-determinism, and can tell you what it will not promise. A repackaged offering cannot produce those artefacts.

Find out what the engines currently say about you

Send me your company name and website. I run a set of buyer questions across ChatGPT, Perplexity, Gemini, Claude and Google's AI Overviews, and send back a recorded walkthrough of what came out: where you appeared, where you did not, who was named instead, and the two or three structural reasons why. No charge, no obligation, and you keep the findings whether or not you decide to hire me.

I run these myself, so there is a queue. Expect a few working days rather than an instant report.