LLM SEO: What Actually Changes, and What You Cannot Measure Yet

Here is the sequence almost everybody goes through. You open ChatGPT, you ask it the buying question a customer would ask — best waterproof duffel for sailing, whatever your category is — and five brands come back, none of them yours. So you ask again in a fresh window to make sure, and this time six brands come back, four of them different from the first five. One of your competitors was in both. You have now got two contradictory readings and no idea which one describes reality.

That instability is not a bug in your test. It is the central fact of this whole discipline, and it is the thing the guides ranking for this term do not tell you. I read the three pages sitting above this one before writing; all three hand you a checklist, and not one of them explains why you cannot tell whether the checklist worked.

The symptom: you asked, you were not named, and the answer keeps changing

If your checks keep contradicting each other, nothing is broken — you are sampling a probabilistic system one draw at a time.

The usual reaction is to take a screenshot of the good answer and treat the bad ones as flukes, or to take a screenshot of the bad answer and conclude the site is invisible. Both are the same error in opposite directions. A single answer from a model tells you roughly what a single coin flip tells you about a coin.

The second symptom people bring is subtler: organic traffic soft, rankings unchanged, revenue flat or down, and no obvious cause in Search Console. That pattern is consistent with answers being read instead of clicked, and it is also consistent with four other things, which is why it needs measuring rather than assuming.

What is LLM SEO?

LLM SEO is the work of getting a brand named and cited inside generated answers — in ChatGPT, Gemini, Perplexity, Copilot and Google’s AI Overviews — rather than ranking a blue link for someone to click.

You will also see it called generative engine optimization, answer engine optimization, GEO, AEO and LLMO. They describe the same work. The proliferating acronyms are a marketing problem rather than a technical distinction, and the term you search for does not change what has to be done.

The most useful framing: this is a subniche inside SEO, not a replacement for it. These systems still pull heavily from Google and Bing results, so classic ranking work keeps paying, and it now pays twice. Anyone telling you that SEO is dead and a new discipline replaces it is selling the new discipline.

Why does the same question give a different answer every time?

Because the model is sampling, not looking up. A generated answer is constructed fresh each time, and the retrieval step that feeds it can pull a different set of sources on each run.

Three separate sources of movement stack on top of each other, and it is worth knowing which is which.

Source of variationWhat it looks likeHow long it lasts
Sampling inside the modelSame question, same day, same account, different brands named. The most common and the most confusing.Every single run
Retrieval and query fan‑outThe system rewrites your question into several searches of its own and reads whatever comes back. Different rewrites, different sources.Every run, and it shifts with wording
Personalization and contextMemory, prior chat, location and account settings steering the answer. A logged‑in test on your own machine is the least reliable test there is.Per account, persistent
Model and product updatesThe whole picture moves at once, in either direction, with no announcement and no changelog you can read.Can undo a quarter of work

Put those together and the conclusion is unavoidable: one check proves nothing, and a favourable screenshot proves less than nothing because of who tends to send them. Any agency deck built on screenshots of good answers is showing you the tail of a distribution.

Attribution is worse here than anywhere in paid media. Somebody who reads about you in an answer and then buys arrives as direct traffic or as a branded search. There is no referrer that says a model recommended you, so expect aggregate lift in branded demand before you ever see per‑conversation proof.

How do you measure LLM SEO without fooling yourself?

By sampling. Ask the same question a fixed number of times in fresh, logged‑out sessions, and record the share of runs you appear in rather than whether you appeared.

That single change — from a yes or no to a rate — is what turns this from theatre into measurement. A brand named in four runs out of ten has a real number attached to it, and a number can move. A brand with a screenshot has an anecdote.

  1. Fix a question set. Eight to fifteen buying questions in your category, phrased as a customer would ask them, never including your brand name. Write them down and do not change them, because the wording is part of the measurement.
  2. Fix a run count. Five to ten runs per question per engine. Fresh chat, logged out, memory off. Consistency matters more than the size of the number.
  3. Record the rate, not the event. Share of runs in which you are named, per question, per engine. Competitors too, in the same pass — their rates are the only benchmark that means anything, and you are collecting them for free.
  4. Record the cited sources. When an answer names its sources, those pages are the actual competition for the citation, and they are frequently not the pages outranking you in Google. This is the most actionable output of the whole exercise.
  5. Repeat monthly, on the same set. Month‑over‑month movement in the rate is the signal. Anything measured once is noise wearing a suit.
  6. Watch branded search and direct traffic alongside it. Given the attribution problem, aggregate branded demand is the closest thing to a business outcome you will get.

Do this for your top three competitors at the same time and you also get the answer to the only strategic question that matters here: is this category winnable for you, or is it locked up by brands the models already know?

What actually changes for a store

Less than the guides imply, and the biggest change is where effort goes rather than what the effort is.

Ahrefs ran the largest correlation study on this, across seventy‑five thousand brands, and published it with the caveat that every factor was weak and correlation is not causation. Read carefully, its ordering is still the most useful thing available: mentions of the brand elsewhere on the web, and on YouTube in particular, tracked far more closely with being named in AI answers than backlinks or domain authority did. Off‑site presence beat on‑site optimization, and it was not close.

The honest reading of that is uncomfortable for anyone selling a technical audit. Much of what those correlations measure is simply being a known brand, which a small store cannot decide to be. The usable conclusion is about allocation — where the next hour goes — not a promise that the hour converts.

WorkWhy it matters hereWhere it sits
Let the AI crawlers fetch the siteBinary, and it makes every other item on this list moot. Check robots.txt for GPTBot, OAI‑SearchBot, ClaudeBot, PerplexityBot, Google‑Extended and CCBot. Plenty of stores block these without knowing it.First, always
Specs and details as text, not baked into imagesThe ecommerce‑specific one, and the item every general guide omits. Materials, dimensions, compatibility and care instructions inside a product image cannot be read, quoted or compared.High, and cheap
Question‑shaped headings with the answer in the first sentenceModels quote answers to questions. A page of statement headings gives them nothing to lift cleanly.High, and it helps human readers too
Comparison and alternative pagesThe questions people actually ask an assistant are comparative. If nothing on your site compares your product to the obvious alternative, you are absent from the answer that matters most.High for a store
Being present on third‑party sites and forumsReviews, category roundups, Reddit threads, YouTube. This is where the correlation data points hardest, and it is outreach rather than a code change.Highest leverage, slowest, least glamorous
Renders without JavaScriptA model fetching your page sees roughly what a raw request sees. Content that only appears after scripts run may not be there at all.Check it once, fix it if broken
Schema markupBuild it for classic rich results, which are proven. Do not buy it as an AI citation lever — the evidence for that specific claim is thin.Worth doing, oversold

For a store specifically, the first two rows and the comparison pages are where I would start, because they are cheap, they are entirely within your control, and they improve the page for people as well. The technical layer of this overlaps almost completely with ordinary Shopify SEO work, which is the practical argument for not treating it as a separate budget.

What has no evidence behind it

Two things are being sold hard right now with very little behind them, and one of them has been publicly contradicted by Google.

The llms.txt file. Google’s own AI optimization guidance states plainly that it is not needed for AI Overviews, AI Mode or any other generative search feature, and Google staff have compared it to the old keywords meta tag and confirmed they are not pursuing it. Server log analyses find the major AI crawlers rarely or never request the file. It costs ten minutes, so build one if it makes you feel better, but nobody should be charging you for it as an AI search deliverable.

Publishing volume. Page count is one of the weakest factors in the correlation data, which makes a programmatic content push a poor answer to this problem. It was already a poor answer to the old one.

Nobody can guarantee a citation. There is no index to submit to, no ranking to check, and no support queue to escalate to. Anyone promising you placement in AI answers is promising something that has no mechanism behind it. That sentence is on my services page too, for the same reason.

The ten minute check you can run yourself

Run this before you pay anyone, including me. It takes ten minutes and it tells you whether there is a problem worth spending on.

  1. Open a fresh, logged‑out chat and ask the buying question for your category — the product and the use case, never your brand name.
  2. Write down every brand named, in order.
  3. Do it twice more in new sessions. Compare the three lists. The overlap between them is the real result; the differences are the variance you now know to expect.
  4. Ask the model why it named those. It will usually cite specific pages. Those pages, not the sites outranking you, are your actual competition for the citation.
  5. Repeat the whole thing in Gemini and Perplexity. The overlap between engines is smaller than most people expect, which is itself useful to know.
  6. Check your robots.txt for the AI crawler names above. If any are blocked, stop here and fix that first, because nothing else counts until it is true.

If you are named in most runs across most engines already, you do not have a problem and you should spend the money elsewhere. If you are named in none of them and a competitor is named in all of them, look at what the model cites when it names them — that is the brief, and it is usually a review site, a roundup or a forum thread rather than anything on their own domain. How I approach AI search visibility starts from that same check, because it is the one piece of evidence in this field you can gather yourself before spending anything.

Questions people ask

Is SEO dead now with AI?
No, and the mechanics argue the opposite. Generated answers still draw heavily on Google and Bing results, so ranking work feeds both the classic result and the AI answer. What has changed is that ranking is no longer sufficient on its own, because a meaningful share of AI citations come from pages that are not in the top ten at all.
What is the LLM equivalent of SEO?
It is called generative engine optimization, answer engine optimization, LLM SEO, GEO, AEO or LLMO depending on who is writing, and they all describe the same work: getting a brand named and cited in generated answers rather than ranking a link. The naming churn reflects an eighteen-month-old field, not five different disciplines.
How do I check if ChatGPT recommends my brand?
Ask the buying question for your category in a fresh logged-out chat with memory off, never naming your brand, and repeat it five to ten times. Record the share of runs you are named in. A single check tells you almost nothing, because the same question returns different brands on different runs.
Does schema markup help with AI search?
There is no strong evidence that it drives AI citations specifically, and it is routinely oversold on that basis. Implement it anyway for classic rich results, where the benefit is well established, but do not treat it as the lever that gets you named in an answer.
Do I need an llms.txt file?
Not for AI search. Google’s own AI optimization guidance says it is not needed for AI Overviews, AI Mode or any other generative feature, and the major AI crawlers rarely request it. It does earn its place on developer documentation and API references, where coding agents genuinely read it.

Want this run properly?

I am Greg Asuncion, an ecommerce paid media consultant. I run Google Ads, Meta, Amazon and Microsoft for direct‑to‑consumer brands on Shopify — one person, published pricing, and a free audit after a call. Here is how I approach AI search visibility.