The demand for an AI visibility metric arrived long before any defensible way to produce one. The result is a market of tools reporting scores derived from undisclosed sampling of systems that behave non-deterministically. Those scores are not useless, but they are estimates of a moving target and they should never be presented as measurements.
What follows is a method a clinic can run itself, which produces a record that holds up when someone asks how it was produced.
Step one: fix the question set
Choose fifteen to twenty-five questions and do not change them, because changing the questions destroys comparability. Cover four families.
- Branded: what is Example Clinic, where is it, what does it treat, is it regulated.
- Local: where can I get a specific treatment in a named town, who does mole checks near a named area.
- Clinical: the treatment questions you have built pages to answer.
- Comparative: the choice questions patients actually ask between two options.
Write them the way a patient would type them, not the way a keyword tool would.
Step two: decide what you record
| Field | Values | Why |
|---|---|---|
| Date and product | Date, assistant, version if shown | Behaviour changes between releases |
| Mentioned | Yes or no | Named in the answer text |
| Cited | Yes or no, with URL | Linked as a source |
| Facts correct | All, partly, no | The outcome that actually matters |
| Incorrect facts | Free text | Each one has a findable origin |
| Other sources cited | Domains | Tells you who owns the answer today |
The fourth and fifth fields are the reason to do this at all. Every incorrect statement about your clinic came from somewhere, and finding the source is a concrete task with a concrete fix.
Step three: repeat on a schedule
Monthly is sufficient. Run every question on every assistant you care about, in a fresh session, without prior context in the conversation, and record the result. Three runs of each question in the same session is not three observations, because context carries over.
Because output varies, treat a single change as noise. A clinic that appears in month one, vanishes in month two and returns in month three has not experienced a collapse and a recovery. Look at the proportion of questions where you are mentioned across the whole set, over several months, and treat that trend as the signal.
Step four: check the server logs
Server logs are the one place with hard evidence. They show which crawlers requested which URLs and when. Filter for the AI user agents you know about, confirm they are receiving 200 responses rather than being challenged by a security layer, and see which pages they fetch. Google publishes its full crawler and user agent list, and other providers publish theirs.
This is the most under-used measurement available. It answers a question no dashboard can: are these systems able to read the site at all.
Step five: watch referrals, carefully
Some assistants send referral traffic and it appears in analytics with identifiable referrers. The volume is typically small and the intent is often high. Two cautions: attribution is incomplete because many answers produce no click at all, and referrer data depends on consent and on the product preserving it. Use it as corroboration, never as the headline number.
AI answers and assistants
LOW CONTROL- Whether your content is crawlable by the relevant user agents
- How clearly a page states the fact you want quoted, and where
- Whether the same fact is consistent across your site, profile and listings
- Structured data that removes ambiguity about who and what you are
- Whether a model chooses to cite you
- Which passage it lifts, or how it paraphrases it
- Whether an answer appears at all for a given prompt
- The training data of any model already released
- Run the same prompt repeatedly and record what is cited, with dates
- Check server logs for AI user agents rather than guessing at access
- Verify the fact you want quoted appears verbatim on a crawlable page
Step six: connect it to enquiries
The measurement that matters is whether more of the right patients contact you. Add one question to the enquiry form or the phone script: how did you come across us. It is imperfect and it is the only direct evidence available. Recorded consistently over a year it becomes a genuinely useful series, and it costs a sentence.
On buying a tool
Commercial AI visibility tools automate the question runs, which saves real time for a group with many locations. Ask three questions before buying: how many runs per question per period, which products and which locales, and whether the raw observations are exportable. A tool that will not show you the underlying observations is asking you to trust a score whose method you cannot inspect.
Reporting it without overclaiming
Report the raw counts: questions asked, mentions, citations, factual accuracy, sources dominating the answers, and what changed. Do not report a single index. Do not compare across products as though the numbers are equivalent. State the sampling method next to the result every time, because the method is what makes the number mean anything, and a number without its method is the thing that gets quoted back at you a year later.
Four ways this measurement goes wrong
- Asking leading questions. A prompt that names your clinic will produce a mention of your clinic. That is not visibility, it is a lookup. Keep branded and unbranded questions in separate groups and never let a branded result stand in for an unbranded one.
- Running in a session with history. Context carries over, so the second question is answered with knowledge of the first. Every observation needs a fresh session.
- Changing the question set. The moment the wording changes, the series restarts. Fix the set, and if a question must be replaced, record the date and treat it as a new series.
- Measuring from one location. Local answers vary with where the request appears to come from. If your catchment matters, note the location for each observation and keep it constant.
Who should run it
This works best in-house, because it takes about an hour a month and because the person who knows what the clinic actually offers is the person best placed to notice that a stated fact is wrong. Outsourcing it produces a report; running it produces the more valuable output, which is a list of incorrect facts and where each of them came from.
