Skip to content
SectionAnswer engines
Reviewed2026-08-01
Words1,299
Sources4
Answer engines

How AI answers choose and attribute sources

The retrieval and composition mechanism behind generated answers, described without invented detail, and what it implies for publishers.

Short answer

A generated answer is produced by rewriting the question into several searches, retrieving candidate documents, selecting passages from them and composing a response with citations. Publishers influence retrieval eligibility and passage quality. Selection, composition and which URL is displayed are decided by the system, and no published method exists for controlling them.

Understanding this mechanism matters because almost every inflated AEO claim depends on the reader not knowing it. The sequence is not secret and the parts that are undisclosed are undisclosed for everyone, which is precisely why claims of control should be treated sceptically.

Diagram 04 / prompt to citation
1. Prompt arrives 2. Query fan-out 3. Retrieval 4. Passage selection 5. Attribution you control nothing here the model rewrites the question into several search queries crawlable, indexed pages are fetched; blocked pages are not candidates a self-contained passage that states the fact plainly is easier to lift the URL shown is chosen by the system, not by the publisher
The only two steps a publisher influences are retrieval eligibility and passage quality. Everything else is decided by the system. That is the whole of what answer engine optimisation can legitimately claim.

Step one: the question is rewritten

A question asked in natural language is not used as a search query directly. The system generates several queries from it, covering different phrasings and sub-questions. Google describes this behaviour in general terms in its documentation on AI features.

The practical consequence is that optimising for the exact phrasing of a question is not the right target. The system is generating its own phrasings, and a page that covers the subject thoroughly is more likely to match one of them than a page tuned to a single string.

Step two: retrieval

Documents are retrieved against those queries. The candidate set comes from an index, and being in that index is a hard prerequisite. This is where publisher influence is real and unambiguous: a page that is blocked, unrendered, uncrawled or unindexed is not a candidate, and no amount of writing quality changes that.

Step three: passage selection

From the retrieved documents, passages are selected for use. This is where writing style has a measurable role, because a passage that states a complete fact with its subject named is usable without repair, and a passage that depends on the three paragraphs above it is not.

The system's criteria for selection are not published. What can be reasoned about is what any such system must be able to do: identify a span of text that answers the query, that is self-contained enough to be quoted or paraphrased, and that is attributable to a source. Writing to make that easy is legitimate, and it is the same discipline that makes a page useful to a person skimming it.

Step four: composition

The answer is composed from the selected material. Paraphrasing is normal, several sources may be blended, and the wording of the answer is generated rather than quoted. A publisher cannot control the wording, and a claim to be able to control it should be treated as a claim about a mechanism nobody has documented.

Step five: attribution

Citations are attached. Which sources are shown, how many, and in what order is a system decision. Not every source used is necessarily shown, and not every source shown contributes equally. Attribution behaviour also differs between products and changes between releases, which is the strongest single reason to be sceptical of any tool claiming a stable measurement of AI visibility.

SURFACE 05

AI answers and assistants

LOW CONTROL
What controls it
  • Whether your content is crawlable by the relevant user agents
  • How clearly a page states the fact you want quoted, and where
  • Whether the same fact is consistent across your site, profile and listings
  • Structured data that removes ambiguity about who and what you are
What does not
  • Whether a model chooses to cite you
  • Which passage it lifts, or how it paraphrases it
  • Whether an answer appears at all for a given prompt
  • The training data of any model already released
How to test it
  • Run the same prompt repeatedly and record what is cited, with dates
  • Check server logs for AI user agents rather than guessing at access
  • Verify the fact you want quoted appears verbatim on a crawlable page

Variance is a property, not a bug

Ask the same question twice and the citations may differ. Ask it from a different account, a different location or on a different day and they may differ again. Any measurement therefore needs repetition to mean anything. A single observation that a clinic is cited is not evidence of a stable position, and a single observation that it is not is not evidence of a problem.

What this rules out

It rules out guaranteed placement in an AI answer. It rules out a technique that reliably makes a specific page the cited source. It rules out any percentage claim about AI visibility improvement from a supplier who did not establish a repeated baseline first. If the mechanism does not permit control, a service selling control is selling something else.

What a publisher can legitimately do

  1. Be retrievable: crawlable, server-rendered, indexed, fast, not blocked at the network layer.
  2. Be clear: self-contained passages, named subjects, question-shaped headings, one claim per paragraph.
  3. Be consistent: identical facts across the site, the profile and the listings.
  4. Be verifiable: named authors, professional registrations, links to public registers, visible dates.
  5. Be complete: answer the follow-up questions on the same page, since a document that answers a cluster is a candidate for more of the generated queries.
  6. Be current: maintained pages with genuine review dates.

That list is unexciting and it is the whole of what the mechanism permits.

Being described versus being cited

Two different objectives are often conflated. Being cited on a general clinical question is a publishing outcome and is competitive with major health publishers. Being described correctly when someone asks about your clinic by name is an entity outcome and is much more achievable, because you are the primary source about yourself.

For most clinics the second is worth considerably more, and it is reached by consolidating the entity and publishing complete facts rather than by chasing citations on generic questions.

Measuring without inventing

Keep a simple record: a fixed set of questions, run monthly, on each of the assistants that matter to you, recording whether your clinic was mentioned, whether it was cited, whether the facts stated were correct, and the date. That record is defensible. A dashboard reporting a single AI visibility score is a model of a system whose behaviour is neither published nor stable, and should be described that way whenever it is used.

No commercial links on this page

This article contains no affiliate links, no sponsored placements and no links to any agency, supplier, clinic or commercial brand. Nobody paid for it, nobody previewed it and no directory advertiser had sight of it. This publication does not rank or recommend agencies anywhere on the site.

Nothing here is medical or legal advice. Regulatory material summarises published guidance. Read the source before relying on it. Our full position is in the editorial standards.

Sources

Primary documentation and regulators only. We do not cite opinion surveys as though they were measurements of a system.

Frequently asked questions

Can we guarantee being cited in AI answers?

No. Selection and attribution are system decisions, undocumented and variable between runs and releases. Any guarantee is a claim about a mechanism nobody has published.

Why do citations change when I ask the same question twice?

Variance is a property of these systems. Retrieval and generation are not deterministic, and personalisation, location and product version all contribute. That is why measurement needs repetition.

Is it better to be cited or to be described correctly?

For a clinic, being described correctly on branded and local questions is usually worth more, and it is far more achievable, because you are the primary source about your own business.

Does being cited send traffic?

Sometimes, and less than an equivalent organic position typically would. Treat citation as a visibility and credibility outcome rather than as a traffic channel, and measure enquiries rather than assuming a click.

Do AI systems use the same index as web search?

Different products draw on different sources, and providers do not publish full detail. What holds across all of them is that content which cannot be fetched cannot be retrieved.

The surfaces move. We write when they do.

A short briefing on documented changes to the surfaces clinics are discovered on. One labelled sponsor slot per issue, no rankings, no invented numbers.