Skip to content
SectionAnswer engines
Reviewed2026-08-01
Words1,341
Sources4
Answer engines

Measuring visibility in AI answer engines

A method for tracking whether a clinic appears in generated answers, built from observations you can defend rather than scores you cannot.

Short answer

Measure AI visibility by running a fixed set of questions repeatedly across the assistants that matter, and recording four things each time: whether the clinic was mentioned, whether it was cited with a link, whether the facts stated were correct, and the date. Repetition matters because these systems vary between runs, and a single observation is not a measurement.

The demand for an AI visibility metric arrived long before any defensible way to produce one. The result is a market of tools reporting scores derived from undisclosed sampling of systems that behave non-deterministically. Those scores are not useless, but they are estimates of a moving target and they should never be presented as measurements.

What follows is a method a clinic can run itself, which produces a record that holds up when someone asks how it was produced.

Step one: fix the question set

Choose fifteen to twenty-five questions and do not change them, because changing the questions destroys comparability. Cover four families.

  • Branded: what is Example Clinic, where is it, what does it treat, is it regulated.
  • Local: where can I get a specific treatment in a named town, who does mole checks near a named area.
  • Clinical: the treatment questions you have built pages to answer.
  • Comparative: the choice questions patients actually ask between two options.

Write them the way a patient would type them, not the way a keyword tool would.

Step two: decide what you record

FieldValuesWhy
Date and productDate, assistant, version if shownBehaviour changes between releases
MentionedYes or noNamed in the answer text
CitedYes or no, with URLLinked as a source
Facts correctAll, partly, noThe outcome that actually matters
Incorrect factsFree textEach one has a findable origin
Other sources citedDomainsTells you who owns the answer today

The fourth and fifth fields are the reason to do this at all. Every incorrect statement about your clinic came from somewhere, and finding the source is a concrete task with a concrete fix.

Step three: repeat on a schedule

Monthly is sufficient. Run every question on every assistant you care about, in a fresh session, without prior context in the conversation, and record the result. Three runs of each question in the same session is not three observations, because context carries over.

Variance discipline

Because output varies, treat a single change as noise. A clinic that appears in month one, vanishes in month two and returns in month three has not experienced a collapse and a recovery. Look at the proportion of questions where you are mentioned across the whole set, over several months, and treat that trend as the signal.

Step four: check the server logs

Server logs are the one place with hard evidence. They show which crawlers requested which URLs and when. Filter for the AI user agents you know about, confirm they are receiving 200 responses rather than being challenged by a security layer, and see which pages they fetch. Google publishes its full crawler and user agent list, and other providers publish theirs.

This is the most under-used measurement available. It answers a question no dashboard can: are these systems able to read the site at all.

Step five: watch referrals, carefully

Some assistants send referral traffic and it appears in analytics with identifiable referrers. The volume is typically small and the intent is often high. Two cautions: attribution is incomplete because many answers produce no click at all, and referrer data depends on consent and on the product preserving it. Use it as corroboration, never as the headline number.

SURFACE 05

AI answers and assistants

LOW CONTROL
What controls it
  • Whether your content is crawlable by the relevant user agents
  • How clearly a page states the fact you want quoted, and where
  • Whether the same fact is consistent across your site, profile and listings
  • Structured data that removes ambiguity about who and what you are
What does not
  • Whether a model chooses to cite you
  • Which passage it lifts, or how it paraphrases it
  • Whether an answer appears at all for a given prompt
  • The training data of any model already released
How to test it
  • Run the same prompt repeatedly and record what is cited, with dates
  • Check server logs for AI user agents rather than guessing at access
  • Verify the fact you want quoted appears verbatim on a crawlable page

Step six: connect it to enquiries

The measurement that matters is whether more of the right patients contact you. Add one question to the enquiry form or the phone script: how did you come across us. It is imperfect and it is the only direct evidence available. Recorded consistently over a year it becomes a genuinely useful series, and it costs a sentence.

On buying a tool

Commercial AI visibility tools automate the question runs, which saves real time for a group with many locations. Ask three questions before buying: how many runs per question per period, which products and which locales, and whether the raw observations are exportable. A tool that will not show you the underlying observations is asking you to trust a score whose method you cannot inspect.

Reporting it without overclaiming

Report the raw counts: questions asked, mentions, citations, factual accuracy, sources dominating the answers, and what changed. Do not report a single index. Do not compare across products as though the numbers are equivalent. State the sampling method next to the result every time, because the method is what makes the number mean anything, and a number without its method is the thing that gets quoted back at you a year later.

Four ways this measurement goes wrong

  1. Asking leading questions. A prompt that names your clinic will produce a mention of your clinic. That is not visibility, it is a lookup. Keep branded and unbranded questions in separate groups and never let a branded result stand in for an unbranded one.
  2. Running in a session with history. Context carries over, so the second question is answered with knowledge of the first. Every observation needs a fresh session.
  3. Changing the question set. The moment the wording changes, the series restarts. Fix the set, and if a question must be replaced, record the date and treat it as a new series.
  4. Measuring from one location. Local answers vary with where the request appears to come from. If your catchment matters, note the location for each observation and keep it constant.

Who should run it

This works best in-house, because it takes about an hour a month and because the person who knows what the clinic actually offers is the person best placed to notice that a stated fact is wrong. Outsourcing it produces a report; running it produces the more valuable output, which is a list of incorrect facts and where each of them came from.

No commercial links on this page

This article contains no affiliate links, no sponsored placements and no links to any agency, supplier, clinic or commercial brand. Nobody paid for it, nobody previewed it and no directory advertiser had sight of it. This publication does not rank or recommend agencies anywhere on the site.

Nothing here is medical or legal advice. Regulatory material summarises published guidance. Read the source before relying on it. Our full position is in the editorial standards.

Sources

Primary documentation and regulators only. We do not cite opinion surveys as though they were measurements of a system.

Frequently asked questions

How many times should I run each question?

At least once per product per month, in a fresh session. More runs give a better picture, and consistency of method matters more than volume. Never change the question wording between periods.

Do AI visibility tools work?

They automate sampling, which saves time. They cannot see inside the systems they measure, and their scores are estimates from undisclosed sampling. Insist on access to the raw observations.

Can I see AI traffic in analytics?

Partially. Some assistants pass identifiable referrers, subject to consent and product behaviour. Many answers produce no click at all, so referral data understates influence and should not be the headline measure.

What should I do when an assistant states something wrong about us?

Find the source. Incorrect facts almost always exist on a record somewhere: an old directory entry, a stale profile, a page on your own site nobody updated. Correct the source, then re-test over the following months.

Is there a benchmark for how often a clinic should be cited?

No credible published benchmark exists, and any figure offered as one should be treated as marketing. Measure your own trend against your own baseline, which is the only comparison with a defensible method.

The surfaces move. We write when they do.

A short briefing on documented changes to the surfaces clinics are discovered on. One labelled sponsor slot per issue, no rankings, no invented numbers.