Answer Engine Optimization: Most of the Work Is Not on Your Site

Here is the position a lot of people are in by now. You read two or three guides to answer engine optimization, and they agreed with each other closely enough to feel like consensus. You restructured the pages so the answer sits under a question heading. You added FAQ schema. You wrote the crisp, extractable, atomic paragraphs everybody recommends. Then you asked ChatGPT the buying question a customer would ask, and four competitors came back, and you were not one of them.

Nothing has gone wrong with the implementation. The checklist was aimed at the part of the problem you control, which is not the part that decides the answer. I read the three highest‑ranking guides for this term before writing, and all three are the same document: define the discipline, contrast it with SEO, hand over an on‑site checklist. All three are written for business software. None of them mentions a product, a retailer or a store. And none of them tells you the most useful thing there is to know, which is where the answers are actually coming from.

You did everything the AEO guides said and you are still not named

If the on‑site work is done and the citations have not followed, the constraint is almost always off‑site, and page‑level changes cannot reach it.

The sequence is predictable enough to be worth naming. Somebody does the structural work over a fortnight, checks a handful of prompts, sees no change, and concludes either that the whole discipline is nonsense or that they need to do more of the same thing harder. Both conclusions are wrong in the same way: they assume the answer is assembled from whichever page is best formatted, and it is not. It is assembled from whichever sources the system retrieved, and your page has to be retrieved before its formatting matters at all.

There is a genuine prerequisite worth ruling out first, because it is binary and it makes everything else moot. If your robots.txt blocks the retrieval crawlers, you cannot be cited, full stop. The ones that matter for citation are OAI‑SearchBot and ChatGPT‑User for ChatGPT, Claude‑User and Claude‑SearchBot for Claude, PerplexityBot and Perplexity‑User for Perplexity, and Googlebot and Bingbot underneath most of it. Those are distinct from the training crawlers, and a great deal of published advice tells people to block both lists as though it were one decision.

Blocking GPTBot, ClaudeBot or Google‑Extended is a decision about training data and has no effect on whether you can be cited today. Blocking OAI‑SearchBot removes you from ChatGPT’s search answers entirely. Anyone who presents those as the same setting has not read the documentation. For a brand that wants to be found, the right default is to name nothing and block nothing.

What is answer engine optimization?

Answer engine optimization is the work of getting a brand named and cited inside a generated answer — in ChatGPT, Gemini, Perplexity, Copilot or Google’s AI Overviews — rather than ranking a link for somebody to click.

You will see the same work called generative engine optimization, GEO, AEO and LLM SEO. They describe one discipline and the proliferation of acronyms is a marketing artefact rather than a technical distinction. I have written about the measurement side of it separately in what actually changes with LLM SEO, and what you cannot measure yet, and this post deliberately does not repeat that ground.

The framing that matters more than the name: this is a subniche inside SEO, not a replacement for it. These systems still pull heavily on Google and Bing results, so conventional ranking work keeps paying and now pays in two places. Anyone telling you SEO is dead and a new discipline has replaced it is selling the new discipline.

Why does the standard AEO checklist not move anything?

Because it optimizes the page and the evidence points at the brand. Almost everything on a standard AEO checklist is on‑site, and the factors that correlate most strongly with being mentioned are things that happen somewhere else entirely.

It is worth being honest about why the guides are shaped this way. On‑site work is what a guide can give instructions for. Nobody can write step four as “become a brand people discuss” and have it read as advice. So the genre optimizes for what it can tell you to do, and the result is a set of tasks that are cheap, checkable, and aimed at the smaller lever.

This is not an argument for skipping the on‑site work. Question headings with the answer in the first sentence underneath genuinely help, because a model lifting two sentences needs something liftable, and the opening section of a page is weighted heavily in what gets extracted. It is an argument about proportion. If the on‑site checklist is a fortnight of work and it is finished, the remaining constraint is not on the site.

What actually correlates with getting cited?

Off‑site brand presence, by a wide margin over anything on the page. Ahrefs ran a correlation study across 75,000 brands and the ranking of factors is not close.

FactorCorrelation with AI mentionsWhere the work happens
YouTube mentions0.737Off‑site. The strongest single factor in the study.
Branded web mentions0.664Off‑site. Being written about by name, with or without a link.
Branded anchors0.527Off‑site.
Branded search volume0.392Off‑site, and largely a consequence of the three above.
Domain Rating0.326Off‑site.
Referring domains0.295Off‑site.
Backlinks0.218Off‑site, and notably weaker than plain unlinked mentions.
Site page count0.170On‑site, and the weakest thing measured. Publishing more pages is not a citation strategy.

Two readings of that table matter, and Ahrefs states the caveats itself: correlation is not causation, every factor measured came out moderate to weak, and brands with no AI mentions at all were excluded from the set. So this is a signpost, not a formula.

The first reading is the useful one. Brand mentions beat backlinks by roughly three to one, which inverts the priority order of most SEO budgets, and unlinked mentions count. Muck Rack’s analysis of 25 million links found the large majority of AI citations came from earned media rather than from brand‑owned pages, which points the same way from a different dataset.

The second reading is the uncomfortable one. Most of what that table measures is downstream of already being a known brand, which a small company cannot simply decide to be. The honest conclusion is about where effort goes, not a promise that effort converts.

Does schema markup help with answer engine optimization?

There is no credible evidence that schema markup improves your chances of being cited in a generated answer, and it sits near the bottom of every factor analysis that has tried to measure it. Build it anyway, for a different reason.

The different reason is that structured data earns rich results in conventional search, which is proven, well documented and worth having on its own terms. FAQPage markup on pages that genuinely have questions and answers, Organization or ProfessionalService markup tying the entity together, Product markup on product pages, Breadcrumb markup for structure. All of that is real work with real returns. It is simply not an answer‑engine lever, and selling it as one is the most common piece of overclaiming in the field.

One rule if you do build it: generate the markup from what is visible on the page. Schema that contradicts the page is the one kind of structured data that actively gets penalised, and hand‑maintained markup drifts away from the page it describes within a few edits.

The same applies to llms.txt, which appears on a lot of checklists. Google’s own guidance states it is not used for AI Overviews or AI Mode, server log analyses find the retrieval crawlers effectively never request it, and when SE Ranking modelled citations across 300,000 domains the model improved once the llms.txt variable was removed. It earns a place on developer documentation, where coding agents are part of the audience. It does nothing for a store.

How do you find out who is getting cited instead of you?

Ask the engine and then read the sources it names. The pages it cites are the real competition for the citation, and they are frequently not the pages outranking you in Google.

  1. Ask the buying question, not the brand question. The category and the use case, phrased the way a customer would — best waterproof duffel for sailing, whatever yours is. Never include your own name.
  2. Write down every brand named. Then ask the same question twice more in fresh, logged‑out sessions, because the answer moves and a single run tells you almost nothing.
  3. Ask it why those. It will usually name specific pages. This is the step everybody skips and it is the one that produces a task list.
  4. Sort the cited pages into three piles. Pages you own, pages you could plausibly appear on, and pages you cannot influence. The middle pile is the work.
  5. Repeat in Gemini and Perplexity. The overlap between engines is smaller than people expect, and a source set that repeats across all three is worth more than one that appears in a single engine.

What comes back from that exercise is nearly always the same shape. Review sites, roundups and best‑of listicles, forum threads, retailer category pages, the occasional YouTube video. Brand‑owned pages show up, but they are rarely the majority, and the brands being recommended are frequently not the ones with the best product pages. They are the ones other people have written about.

Which turns the middle pile into a recognisable job: getting into the roundups that already rank, being present where the category is discussed, having the review coverage the engines keep reaching for. It is outreach, it is slow, and it is the highest‑leverage work available to a site without authority. It is also the part no checklist wants to be the answer.

What this looks like for a store rather than a software company

The local half of the standard advice does not apply and the product half is missing entirely. Every guide ranking for this term was written for software or for a business with a physical location, and a direct‑to‑consumer store is neither.

Standard AEO adviceWhether it applies to a storeWhat to do instead
Keep your Google Business Profile currentNo. There is no profile, no local pack and no proximity query in play.Put the equivalent effort into the review platforms and category forums where products get discussed.
Publish original research and dataRarely practical for a store, and it is the advice most often given to people with no research function.Specifications, materials, sizing and comparison data written as text. It is the concrete detail that gets extracted, and stores already own it.
Build topical authority with volumeWeakly. Page count was the lowest factor Ahrefs measured.Fewer pages that answer real buying questions completely, including the comparison and alternative pages most stores refuse to write.
Add schema markupYes, but for rich results rather than citations.Product, Organization and Breadcrumb markup generated from the live page. Treat it as conventional SEO work, because that is what it is.
Get mentioned in the pressYes, and it is the highest‑value item on the list.Category roundups, review sites and the forums for your niche. Less glamorous than press and considerably more achievable.

There is one store‑specific trap worth naming. A lot of product detail lives inside images — the specification table baked into a graphic, the sizing chart as a JPEG, the material breakdown in a lifestyle shot. A system fetching your page reads the text. If the specifications that would qualify your product for a recommendation exist only as pixels, they are not there as far as the answer is concerned, and this is common enough on Shopify themes to be worth checking before anything else. It overlaps almost exactly with ordinary Shopify SEO work, which is the point: the two disciplines share most of their task list.

What nobody can promise you

Nobody can promise you a citation. There is no index to submit to, no support queue, no verification step and no mechanism by which anybody guarantees a brand appears in a generated answer.

The discipline is roughly eighteen months old, which is not long enough for anybody to have durable evidence about what works. Anyone quoting a precise uplift figure is describing a sample too small and too recent to support it. Treat the correlation data above the same way — as the best signpost currently available, not as a mechanism.

  • A single check proves nothing. The same question asked three times returns three different answer sets, so a favourable screenshot is the tail of a distribution rather than a result.
  • Attribution is worse here than in paid media. Somebody who reads about you in an answer and then buys arrives as direct traffic or a branded search. There is no referrer that says a model recommended you.
  • A model update can undo a quarter of work, in either direction, with no announcement and no changelog.
  • Every tool in this category is measurement, not optimization. They sample prompts and report whether you were named. Useful if you are publishing often enough for a monthly readout to change next month’s work, and an expensive number you cannot act on if you are not.

The reasonable position is that the on‑site work is cheap, sensible and worth doing once; that the off‑site work is where the evidence points and where the effort should go; and that the whole thing pays for itself regardless, because the same work ranks pages in Google and Bing, which is still where most of the money is. That is the case for treating this as a wedge into an SEO engagement rather than a separate discipline with its own promises.

Questions people ask

What is AEO vs SEO?
SEO gets a link ranked so somebody clicks it; answer engine optimization gets a brand named inside a generated answer. In practice they share most of their task list, because the systems producing those answers still pull heavily on Google and Bing results. AEO is a subniche inside SEO rather than a replacement for it.
How do you do answer engine optimization?
Confirm the retrieval crawlers are not blocked, write question headings with the answer in the first sentence beneath them, then spend most of the effort off your own site. The factors that correlate most strongly with being mentioned are brand mentions and third-party coverage, not page-level formatting.
Does schema markup help with answer engine optimization?
There is no credible evidence that it does, and it scores near the bottom of every factor analysis that has tried to measure it. Build it anyway for conventional rich results, which are proven, but do not treat it as an AEO lever or pay anyone who sells it as one.
What is the best answer engine optimization tool?
Every tool in this category measures rather than improves anything, so the question is whether a monthly readout would change what you publish next month. If you are shipping content regularly, a cheap tier is a reasonable feedback loop. If the site was set up once and left, you are paying for a number you cannot act on.
Is answer engine optimization worth it for a small store?
The on-site portion is worth doing once because it is cheap and it improves conventional rankings at the same time. The off-site portion is genuinely slow and nobody can promise a citation at the end of it, so treat it as a long-term brand presence investment rather than a channel with a forecast.

Want this run properly?

I am Greg Asuncion, an ecommerce paid media consultant. I run Google Ads, Meta, Amazon and Microsoft for direct‑to‑consumer brands on Shopify — one person, published pricing, and a free audit after a call. Here is how I approach AI search visibility.