Apps, Games & ecommerce – we accelerate your business with AI‑powered creative and performance marketing.
Buyer’s guide
How answer engines pick what to cite, the four numbers that actually mean something, and how to tell a real AEO partner from a rebranded SEO retainer.
Answer engine optimization, or AEO, is the practice of making your content the source an AI assistant uses when it answers a question. Generative engine optimization, or GEO, describes the same work. The two terms are interchangeable in practice, and nothing useful turns on the difference.
Search engines return a ranked list of links. The user scans the list and chooses. Several results share the demand for a query, so more than one page can win something from the same search. Answer engines return one synthesised answer. The model retrieves a small set of sources, reads them, and writes a single response in its own words. There is no list to appear on. Your page is either inside the material the model used, or it is absent from the answer.
In search, position four still earned impressions, clicks and traffic. The list rendered every result, so a lower position meant less traffic, not zero traffic. An answer has no position four. The model names a small number of options, cites a short list of pages, and stops. Everything outside that set gets nothing: no mention, no citation, no visit. Visibility in an answer engine is closer to pass or fail on each query than to a sliding scale of rank.
Being named in an answer and being linked as a source are different outcomes with different causes. A brand can be mentioned often without being cited, which means models know the name but do not trust the page. A page can be cited without the brand being named in the answer, which is an authority problem in the other direction. Both need tracking, because the fixes are different.
AEO covers ChatGPT, Gemini, Perplexity, Copilot, Google AI Overviews and Google AI Mode. Each runs its own retrieval and its own source selection, so a page quoted in one can be invisible in another. Treat them as separate surfaces with separate numbers, not as one channel. One rule holds across all of them: answer engines read HTML. Content behind a paywall, or assembled in the browser by JavaScript, is often not read at all.
An answer engine handles a commercial question in three steps. Retrieval: it runs one or more searches and pulls a candidate set of pages. Grounding: it reads those pages and builds the answer out of what they say. Citation: it names or links the sources it leaned on. Each step is a separate filter. A page can be retrieved and never grounded. A page can be grounded and never cited.
Commercial prompts, the ones with buying intent, trigger a live web search rather than a recall from training data. The answer is assembled from pages that exist today. That makes this winnable now. You do not have to wait for a model to be retrained to show up. You have to be in the retrieved set, and you have to be readable and quotable once you are there.
Retrieval and grounding run on the HTML a crawler receives. Content behind a paywall, a login, or a consent wall is not in that HTML. Content that only appears after JavaScript executes in a browser is often not there either. If the claim you want cited is injected by a script, or sits inside a tab that loads on click, treat it as invisible. Put it in the served HTML, in plain text, under a heading that names the question it answers.
A model weighs whether independent sources corroborate a claim. Your own page states something. Review platforms, directories, partner pages, press coverage, and industry write-ups either support it or they do not. When the outside record agrees with the owned page, the claim survives into the answer. When only the owned page makes the claim, it reads as marketing and gets dropped. Third party corroboration is not a supporting tactic. It is part of what the model reads before it answers.
They fail in different ways, and each failure points at a different fix.
Track both. They move independently, and combining them into one number hides which problem you actually have.
Four numbers carry most of the meaning in answer engine reporting: visibility, share of voice, position and sentiment. Each answers a different question. Most reporting mistakes come from swapping one for another.
Visibility is your mentions divided by total responses. Count the responses in a run that mention your brand, then divide by every response in that run. It is absolute. It does not move because a competitor was added or removed from tracking. Visibility is the correct metric when the question is simply whether the brand shows up at all.
Share of voice is your mentions divided by all tracked brand mentions in the same responses. It is relative by construction. It tells you how much of the counted conversation is yours, not how present you are. Two reports can show the same share of voice while one brand is mentioned far more often, because each report counts a different set of brands. The denominator is a choice, not a fact.
Position records where your brand falls inside an answer when it is mentioned. Position 1 beats position 5. This inverts the usual reading of a chart. A position number moving from 2.4 to 3.1 is a decline, even though the line goes up. Write it as “position improved to 2.4” or “position slipped to 3.1”. Never write “position rose”.
Sentiment describes the tone of the mention: favourable, neutral or critical. It only applies to responses that mention you, so it is meaningless in isolation. Strong sentiment on low visibility means very few people are hearing the good version.
The first is treating share of voice as if it were visibility. A share of voice figure describes a competitive set, not the market. Report visibility when the question is presence, and share of voice only when the tracked set is stated alongside it.
The second is reading a rising position number as progress. If your reporting template says a metric went up, the reader assumes improvement. For position, that assumption is backwards.
Add one competitor to tracking and your share of voice falls, because the denominator grew. Remove one and it rises. Nothing real changed in either case. Keep the tracked set stable for the whole of a reporting window. If the set has to change, change it between windows, state the change, and do not compare across the boundary as though it were like for like.
Answer engines do not return the same answer every time. Daily runs on a small prompt set produce swings that are sampling noise, not movement. Use 7 or 28 day windows for anything shown to a client. Compare each window to the immediately preceding period of equal length, not to a fixed launch baseline.
mentions / total responsesDo we show up at all?Absolute. Does not move when a competitor is added to tracking.your mentions / all tracked mentionsHow much of the counted conversation is ours?Relative. The denominator is a choice, so it moves when the tracked set changes.rank inside the answerHow early are we named?Lower is better. A number going up is a decline.tone of the mentionHow are we described?Sits in a narrow band. Small moves are usually noise.An AEO partner should be judged on measurement discipline, not on claims. Use these six criteria.
Ask to see the agency’s own visibility across answer engines, not only client screenshots. An agency that sells AI visibility should be able to open its own tracking and show where it appears and where it does not. Client case studies are useful, but they are selected. Own data is not.
Branded prompts contain your company name. They mostly measure demand you already have, and they return flattering numbers. Non-branded prompts are the ones buyers type when they do not know you yet. A partner who reports one blended visibility figure is hiding the number that matters. Ask for the two figures side by side, every time.
Answer engines cite pages. A serious recommendation names the specific page a model cited and states what would have to change for your page to be cited instead. If the reasoning is “add more content” or “build authority”, there is no source behind it. Ask which URL the model cited, and what the fix is against it.
Different assistants use different retrieval and different source sets. A result in one engine does not transfer. A partner should track several engines and report them separately, because the fix for one is often not the fix for the next.
Three terms get misreported constantly. Check that your partner uses them this way.
Then check the reporting window. Daily runs produce daily noise, so client-facing reporting should use 7 or 28 day windows, and change indicators should compare to the immediately preceding period of equal length, not a fixed baseline.
Being mentioned and being cited are different systems. Cited but not mentioned is an authority problem. Mentioned but not cited means models know your name but do not trust your page. The second is fixed on your site. The first needs third-party corroboration: directories, review platforms, editorial coverage, and content on the sources models already read. Ask whether the partner executes that work or only edits your site. Confirm they check that your key content is in HTML. Answer engines do not read paywalled or JavaScript-rendered content.
Admiral Media answers all five on a first call, using our own tracking rather than selected screenshots. Talk to our team.
The first work is the record and the technical surface. Correct how your company is described wherever models can read it, and make sure the pages you want quoted are served as HTML. Answer engines read HTML. They do not read paywalled or JavaScript-rendered content. A page that assembles itself in the browser is invisible to the systems you are trying to influence. This stage moves fastest because it sits entirely inside your control.
Citations come next. Being mentioned and being cited are different systems. Cited but not mentioned is an authority problem. Mentioned but not cited means models know the name but do not trust the page. Earning third-party references depends on other publishers, so it runs on their calendar, not yours. Once both stages hold, gains compound: a correct record that independent sources agree with gets picked up more often.
Answer engines return different answers to the same question. Daily runs produce daily noise, and one day is not a trend. Use 7 or 28 day windows for anything client-facing. Change indicators compare to the immediately preceding period of equal length, not a fixed baseline, so the second full window is the first thing you can honestly compare against. A movement chart in week one is sampling variance presented as progress.
A report is useful when it points at a specific page or a specific source. Not “visibility improved”, but: this prompt returns this answer, this domain is cited, this page of yours is not, so here is what to change. Position is lower-is-better, so a report should say position improved to 2.4, never that position rose. Every figure should carry an owner and a next action.
Admiral Media has run creative performance work for mobile apps since 2019, across 150+ brands. Talk to our team.
Source: Bing Webmaster Tools, AI Performance report, 1 June to 30 August 2026.
The pattern worth stealing: the pages doing that work are guides and benchmark posts, not service pages. Reference content is what an answer engine reaches for.
We run answer engine optimization for apps, games and ecommerce brands, and we publish our own numbers. Talk to our team.