To measure AI citations for a clinic, ask each AI engine the same patient question several times. Count only the links inside the answer, and score each clinic by how many runs cite it. I tested this on 25 September 2026 with 126 answers from ChatGPT, Claude and Perplexity. If an engine cited a clinic in any of five identical runs, a single check still missed it 37% of the time.

Why is one AI check not enough?

Because the answer changes every time you ask.

Most clinic owners check their AI visibility once. They type a question into ChatGPT, see whether their clinic comes up, and draw a conclusion. Some are relieved. Some are worried. Both are reading one roll of the dice.

This is not a new worry. In a study published in January 2026, SparkToro and Gumshoe had 600 volunteers run 12 prompts through ChatGPT, Claude and Google's AI a combined 2,961 times. Their finding was blunt. There is "a <1 in 100 chance that ChatGPT or Google's AI, if asked 100X, will give you the same list of brands in any two responses".

That study tracked brands named in the answer text, like headphones and chef's knives. It asked US questions. It did not test the thing a clinic cares about: whether an AI engine links to your own website when a local patient asks for help.

So I ran that test on Australian clinic questions. Then I turned what I learned into a method you can repeat without buying a tool.

What did I test?

Six patient questions, three engines, and every question asked seven times per engine.

The questions are the kind a patient types when they want a provider nearby. Each names a suburb and a city:

  • A physiotherapist in Brunswick, Melbourne, for knee pain.
  • Emergency dentists in Parramatta.
  • A psychologist in Fortitude Valley, Brisbane, for anxiety.
  • A skin cancer check in Fremantle, WA.
  • GP clinics in Glenelg, Adelaide, taking new patients.
  • A podiatrist in Hobart for heel pain.

I asked each question five times with identical wording, then twice more in different words. That was done on ChatGPT, Claude and Perplexity. The total is 126 answers, and every one was saved.

For each answer I pulled out every link and reduced it to a website. I then read the homepage of all 136 websites the engines returned and sorted them by hand. A website was counted as a clinic only if it was a provider's own site: a clinic, a practice, a chain or a private hospital. Directories, booking sites, review sites, social media and government pages were counted separately.

What happens when you ask the same question five times?

You get a different set of clinics, most of the time.

Here is the core result. Take every clinic an engine linked to at least once across the five identical runs of a question. That gives 97 clinic and question pairings across the three engines. Then count how many of the five runs linked to each one.

Runs out of 5 that linked to the clinicClinic and question pairings
All 5 runs.27
4 runs.18
3 runs.13
2 runs.20
Only 1 run.19

Only 27 of 97 appeared every time. 19 appeared once and never again.

Now put yourself in the position of a clinic owner doing a single check. If the engine linked to your clinic in some of those runs, what are the odds your one check catches it? Across all three engines, one check missed the clinic 37% of the time. More than a third of the clinics an engine does cite would have been told they were invisible.

The reverse error happens too. A clinic that shows up once may take it as a sign it is safely visible. 19 of 97 pairings were a one-off.

Which engine is the most consistent?

Perplexity, by a wide margin. ChatGPT was the least consistent of the three.

This chart shows the chance that one check misses a clinic the engine linked to in at least one of five identical runs.

ChatGPT (GPT-5 mini, web search)44%
Claude (Sonnet 4.5, web)41%
Perplexity (Sonar)24%
EngineClinic and question pairingsLinked in all 5 runsLinked in only 1 run
ChatGPT.3958
Claude.3078
Perplexity.28153

ChatGPT is the one to watch. It linked to 39 different clinic and question pairings, the widest spread of the three. But only 5 of the 39 came back in all five runs. On ChatGPT, being cited is closer to a lottery ticket than a ranking.

Perplexity is steadier. 15 of its 28 pairings appeared in every run. That makes it the easiest engine to check. It does not make it the most important one, and this test does not measure how many patients use each engine.

What counts as a citation?

Only the links a patient can see in the answer. That sounds obvious. It changed the numbers more than anything else.

Perplexity sends back a long list of sources with every answer, often around 20. Most of them are never mentioned in the answer itself. Across all its answers, Perplexity returned 810 sources and referred to only 307 of them (38%) in the text. Claude referred to 164 of the 210 it returned. ChatGPT placed every link inside its answer.

If you count the full source list, Perplexity looks almost perfectly consistent, and far more clinics look cited than really are. Any tracker that counts the full source list will overstate both.

One example shows why it matters. For the Hobart heel pain question, Perplexity returned the same podiatry clinic in every one of its five identical runs. The clinic is in Hobart, Indiana, in the United States. Its page is titled "Podiatrist Hobart". Perplexity never mentioned it in any of the five answers. A tool reading the source list would have logged a steady citation. A patient never saw it.

There is a second trap in the other direction. In one ChatGPT answer about Brunswick physios, a clinic was named in the text. The link beside it went to a review site's page about the clinic. It did not go to the clinic's own website. A name is not a link. Record both, in separate columns.

This also connects to something my earlier test of four AI engines found. An engine that cannot tell which Newcastle, or which Hobart, you mean will not guess in your favour. Put your suburb, state and country in the visible sentences on your pages.

Does rewording the question change who gets cited?

Yes. Rewording finds clinics that repeating never does.

For each question I also asked two rewordings, once each, on every engine. For example, "Can you recommend a podiatrist in Hobart for heel pain?" became "Who treats heel pain in Hobart, Tasmania? I need a podiatrist."

Most of what the rewordings linked to was already in the five-run pool. The share was 81% on ChatGPT, 85% on Claude and 84% on Perplexity. So the core of each answer held up when the words changed.

But the rewordings also turned up 18 clinic websites that five identical runs never linked to: 7 on ChatGPT, 5 on Claude and 6 on Perplexity. Patients do not all type the same sentence. SparkToro found the same thing at much larger scale: across 142 prompts written by real people for one intent, "there were barely two prompts that, if you squinted, looked similar".

The lesson for the method is simple. Repeat each question, and also word it more than one way.

If ChatGPT cites you, do Claude and Perplexity?

Often not. Each engine keeps its own list.

Across all 126 answers, the engines linked to 62 different clinic websites in the answer text. Here is how many engines linked to each one.

Linked by one engine only29 of 62
Linked by two engines13 of 62
Linked by all three engines20 of 62

Almost half the clinics were visible on one engine and absent on the other two. That matches the pattern in my September test of four engines, where 74.3% of cited clinics appeared on only one. The exact share moves with the questions and the day. The direction does not.

So a check on one engine tells you about that engine. It says very little about the others.

How many times should you ask each question?

At least three times per engine. Five is better.

To answer this, I took the five identical runs and asked a simple question. If you had only run the question once, twice, three or four times, how much of what five runs found would you have seen? I averaged that over every possible order of the five runs.

Checks per questionChatGPTClaudePerplexity
1 check.58%60%82%
2 checks.79%78%91%
3 checks.90%88%96%
4 checks.96%95%99%

Three checks found about 90% of what five checks found on ChatGPT and Claude. One check found under 60%.

Be clear about what this means. It is a share of what five runs found, and five runs are a small sample. SparkToro's advice, for anyone who wants an engine's full list, is to ask "usually at least 60-100X". Nobody running a clinic will do that. Three to five runs per question is the honest middle: enough to stop one lucky or unlucky answer from deciding what you believe.

What is the method, step by step?

Here is the whole method. It needs a spreadsheet and some patience. It does not need a paid tool.

  1. Write three to six questions before you look. Use your patients' words, not your service names. Put the suburb and the state in each one. Writing them first stops you from choosing the questions you already win.
  2. Write one or two rewordings of each. Change the words and keep the need. A symptom version and a service version is a good pair.
  3. Pick your engines. Check each one separately. Almost half the clinics in this test were linked by one engine only.
  4. Start fresh each time. Use a new chat for every run, so an earlier answer cannot shape the next one.
  5. Ask each question three to five times with identical wording. Then ask each rewording once.
  6. Record four things per answer. Is your clinic named in the text? Is your own website linked in the text? Which page is linked? Is a third-party page about you linked instead?
  7. Ignore the source panel. Count only the links the answer shows inside its text.
  8. Check where every link points. Suburb, state and country. A Hobart in Indiana is not your Hobart.
  9. Score each question as a count out of the runs. Linked in 5 of 5 is a solid position. Linked in 1 or 2 of 5 is an edge position that can vanish tomorrow. Zero is zero.
  10. Save every answer and repeat quarterly with the same questions. Compare the counts, not single answers.

If that feels like a lot, start small. Three questions, three engines and three runs each is 27 answers. On this data, three runs found about nine tenths of what five runs found on ChatGPT and Claude.

The count out of five is the number to track. It is the same idea SparkToro landed on after its study: "visibility % across dozens to hundreds of prompts run multiple times is a reasonable metric". A ranking position is not.

What you do with the result is a separate job. My guide on how AI search picks its answers covers what engines pull from a page. The questions patients ask AI before they ask a clinic is a good source of wording for step 1.

Can you advertise that an AI recommends your clinic?

Be very careful. On this data, it would not even be true twice in a row.

It is tempting to screenshot a good answer and put "recommended by ChatGPT" on your homepage. Think about what the screenshot proves. On ChatGPT, 5 of 39 clinic and question pairings appeared in every run. A patient who asks the same question tomorrow may well get a different list.

Ahpra's advertising guidelines say advertising for a regulated health service must not "be false, misleading or deceptive, or likely to be misleading or deceptive". They also list advertising that "makes claims about providing a superior regulated health service" as an example of how advertising can mislead.

An AI answer is not an award or a review. It is one draw from a set of candidates. Treat your citation check as private market research. Use it to decide which pages to improve. Keep it off your marketing.

The same caution applies to what an engine says about you. My post on what Google's AI tells patients before they click covers why the answer about your treatment is worth checking, even when the link is yours.

How did I measure this, and what can it not tell you?

The method is short, and every raw answer was kept so it can be repeated.

Method. 6 Australian patient questions, each asked 5 times with identical wording and twice more in other words, on 3 engines: 126 answers, run 25 September 2026. ChatGPT through the OpenAI Responses API (gpt-5-mini, web search, user location Australia). Claude Sonnet 4.5 with a web search plugin and Perplexity Sonar, both through OpenRouter. Every engine got the same neutral instruction to answer as it would for a member of the public. 136 websites were returned in total. Each was classified by hand from its homepage: 90 were clinic websites, 32 directories or platforms, 11 public or peak body pages, 2 outside Australia, and 1 could not be told. A link counted as a citation only if it appeared in the answer text. No clinic is named here.

Now the limits.

  • This is 126 answers on one day. Engines change often. The percentages will move. The method is the point.
  • I used the engines' developer interfaces, which are not identical to the apps patients use. SparkToro reports early data suggesting the gap between the two may be smaller than feared, and calls it an open question.
  • Claude's web plugin returned at most five sources per answer, which caps how many clinics it could cite.
  • Google's AI Overviews and Gemini were not tested this time. My earlier four-engine test covers AI Overviews.
  • Five runs is a small sample. "Found in 5 of 5" means steady across five tries, not guaranteed.
  • This test counts links. It does not judge whether what an engine said about a clinic was accurate.
  • This is general information about measuring AI answers. It is not legal advice about your practice.

If you want your whole site read the way ChatGPT and a regulator read it, that is what the All Clear Audit does. It quotes every flagged line, names the rule, and writes the replacement. If you would rather talk it through first, book a call. For the page-level work that follows a check like this, see SEO and GEO for clinic websites.

Prefer the short version? The findings are in an eight-slide web story.


Related reading: