Understanding this mechanism matters because almost every inflated AEO claim depends on the reader not knowing it. The sequence is not secret and the parts that are undisclosed are undisclosed for everyone, which is precisely why claims of control should be treated sceptically.
Step one: the question is rewritten
A question asked in natural language is not used as a search query directly. The system generates several queries from it, covering different phrasings and sub-questions. Google describes this behaviour in general terms in its documentation on AI features.
The practical consequence is that optimising for the exact phrasing of a question is not the right target. The system is generating its own phrasings, and a page that covers the subject thoroughly is more likely to match one of them than a page tuned to a single string.
Step two: retrieval
Documents are retrieved against those queries. The candidate set comes from an index, and being in that index is a hard prerequisite. This is where publisher influence is real and unambiguous: a page that is blocked, unrendered, uncrawled or unindexed is not a candidate, and no amount of writing quality changes that.
Step three: passage selection
From the retrieved documents, passages are selected for use. This is where writing style has a measurable role, because a passage that states a complete fact with its subject named is usable without repair, and a passage that depends on the three paragraphs above it is not.
The system's criteria for selection are not published. What can be reasoned about is what any such system must be able to do: identify a span of text that answers the query, that is self-contained enough to be quoted or paraphrased, and that is attributable to a source. Writing to make that easy is legitimate, and it is the same discipline that makes a page useful to a person skimming it.
Step four: composition
The answer is composed from the selected material. Paraphrasing is normal, several sources may be blended, and the wording of the answer is generated rather than quoted. A publisher cannot control the wording, and a claim to be able to control it should be treated as a claim about a mechanism nobody has documented.
Step five: attribution
Citations are attached. Which sources are shown, how many, and in what order is a system decision. Not every source used is necessarily shown, and not every source shown contributes equally. Attribution behaviour also differs between products and changes between releases, which is the strongest single reason to be sceptical of any tool claiming a stable measurement of AI visibility.
AI answers and assistants
LOW CONTROL- Whether your content is crawlable by the relevant user agents
- How clearly a page states the fact you want quoted, and where
- Whether the same fact is consistent across your site, profile and listings
- Structured data that removes ambiguity about who and what you are
- Whether a model chooses to cite you
- Which passage it lifts, or how it paraphrases it
- Whether an answer appears at all for a given prompt
- The training data of any model already released
- Run the same prompt repeatedly and record what is cited, with dates
- Check server logs for AI user agents rather than guessing at access
- Verify the fact you want quoted appears verbatim on a crawlable page
Variance is a property, not a bug
Ask the same question twice and the citations may differ. Ask it from a different account, a different location or on a different day and they may differ again. Any measurement therefore needs repetition to mean anything. A single observation that a clinic is cited is not evidence of a stable position, and a single observation that it is not is not evidence of a problem.
It rules out guaranteed placement in an AI answer. It rules out a technique that reliably makes a specific page the cited source. It rules out any percentage claim about AI visibility improvement from a supplier who did not establish a repeated baseline first. If the mechanism does not permit control, a service selling control is selling something else.
What a publisher can legitimately do
- Be retrievable: crawlable, server-rendered, indexed, fast, not blocked at the network layer.
- Be clear: self-contained passages, named subjects, question-shaped headings, one claim per paragraph.
- Be consistent: identical facts across the site, the profile and the listings.
- Be verifiable: named authors, professional registrations, links to public registers, visible dates.
- Be complete: answer the follow-up questions on the same page, since a document that answers a cluster is a candidate for more of the generated queries.
- Be current: maintained pages with genuine review dates.
That list is unexciting and it is the whole of what the mechanism permits.
Being described versus being cited
Two different objectives are often conflated. Being cited on a general clinical question is a publishing outcome and is competitive with major health publishers. Being described correctly when someone asks about your clinic by name is an entity outcome and is much more achievable, because you are the primary source about yourself.
For most clinics the second is worth considerably more, and it is reached by consolidating the entity and publishing complete facts rather than by chasing citations on generic questions.
Measuring without inventing
Keep a simple record: a fixed set of questions, run monthly, on each of the assistants that matter to you, recording whether your clinic was mentioned, whether it was cited, whether the facts stated were correct, and the date. That record is defensible. A dashboard reporting a single AI visibility score is a model of a system whose behaviour is neither published nor stable, and should be described that way whenever it is used.
