The citation data finally arrived, and it redraws the map
A year into the AI-search era, visibility advice ran mostly on vibes. This year the measurement caught up: large-scale citation studies tracking which sources the major engines actually surface for health and commercial queries. The findings are specific enough to act on:
Institutions dominate health answers. Government health bodies appear in roughly four of ten health citations from the leading engines; the major consumer-health encyclopedias and academic medical centers take most of the rest of the head terms. For "what is" and "is it safe" questions, the engines have chosen their sources, and they are not brands.
The engines barely overlap. Only around one in ten cited domains appears across both of the leading answer engines for comparable queries. Visibility on one engine says almost nothing about visibility on another; they are separate citation economies with separate tastes.
Community content carries real weight. Discussion-forum content, one platform above all, shows up in roughly forty percent of relevant answers across engines, because engines treat lived-experience threads as evidence for exactly the questions institutions answer generically.
Freshness compounds. One major engine cites content under thirty days old at rates above eighty percent for evolving topics. Dated, recently-updated pages beat authoritative-but-stale ones on any question where the ground is moving.
Our foundational playbook, Generative Engine Optimization for Telehealth, covered the mechanics. This is the strategy update the data demands.
The query map: concede the head, own the territory
The institutional-dominance finding sounds like bad news until you segment the queries. Institutions win the definitional head, what is this drug, what are the side effects, and no brand budget changes that. But the engines still need sources for everything institutions do not publish:
| Query territory | Institutional coverage | Brand opportunity |
|---|---|---|
| "What is X" definitions | Total | Concede; link through, do not compete |
| Comparisons: option vs option, program vs program | Thin, generic | Own with structured, even-handed comparison pages |
| Process reality: "what actually happens when" | Absent | Own with experience-grounded walkthroughs |
| Cost, coverage, and access mechanics | Sparse and stale | Own with current, dated explainers |
| "How to choose" and evaluation criteria | Absent | Own with honest frameworks |
| Fresh developments: new approvals, program changes | Slow | Own with fast, structured coverage |
| Lived experience: "what does it feel like" | Absent by design | Own via community presence and patient-voice content |
That right-hand column is the entire brand play, and it maps precisely onto what a real operator knows and institutions never will: how programs actually work, what patients actually experience, what changed this month. The comparison and process territories are where posts like The Pill-Curious Patient and the Bridge coverage earn citations, and the freshness territory is why the H2 2026 Metabolic Calendar approach, dated, structured, updated, is engine bait by design.
Three engines, three plays
The low-overlap finding kills the write-once-rank-everywhere assumption. The practical translation:
The retrieval-heavy engine rewards exactly what classic technical SEO rewards, crawlable, structured, schema-marked pages, plus aggressive freshness. Play: dated explainers on moving topics, updated on a visible cadence, with the update itself noted. This is where the eighty-percent-fresh finding lives, and where a monthly refresh habit outperforms a yearly masterpiece.
The conversational engines lean on training-era authority plus their own browsing, favoring sources that read as reference material: comprehensive, even-handed, definition-forward. Play: the glossary-and-framework layer, pages structured as answers with the question in the heading, the direct response in the first sentences, and the nuance beneath, the pattern running through The DTC Telehealth Glossary.
The community-weighted layer runs through forums, and it cannot be bought, only earned: genuine, disclosed, helpful participation where patients already discuss your category, plus patient-voice content on your own surfaces that engines read as experience rather than marketing. The compliance guardrails from Affiliate and Creator Programs for DTC Telehealth apply to any incentivized presence, and undisclosed astroturf is both an FTC problem and, increasingly, something engines detect.
The unifying requirement across all three: machine-legible trust. Named clinical authorship, review dates, published governance, the artifacts from Clinical Governance Is a Growth Asset, because every engine's health filter looks for exactly those signals before citing a commercial domain.
The freshness cadence as operating system
The freshness finding deserves its own operational answer, because it converts content from a project into a rhythm:
- A dated-page inventory. Every page competing on a moving topic carries a visible last-reviewed date and an owner. Stale dates are now measurable liabilities.
- A monthly refresh pass. The moving pages, program mechanics, coverage explainers, pipeline coverage, get a genuine update monthly: new data in, dead facts out, date advanced. Engines detect cosmetic re-dating; the update has to be real.
- Event-triggered publishing. The calendar-driven pattern: pre-written coverage for known dates, approvals, program launches, policy changes, shipped within hours. The engines' freshness bias turns every industry event into a citation land-grab that the prepared win, which is the standing argument of the H2 calendar.
- Measurement by mention. The tracking loop from the foundational playbook, sampled brand-mention checks across engines, now segmented per engine, because the low-overlap finding means a single blended score hides everything useful.
Operators running this cadence describe the compounding honestly: each fresh, structured page is a small bet, and the engines' own biases roll the winnings forward.
FAQ
Which sources do AI engines cite most for health questions? Institutional ones: government health bodies appear in roughly 40% of health citations from leading engines, with major consumer-health encyclopedias and academic centers taking most remaining head-term citations. Brands win by targeting the query territories institutions leave open rather than competing on definitions.
Do different AI engines cite the same sources? Rarely: studies find only about 11% domain overlap between leading engines for comparable queries. Each engine is a separate citation economy, retrieval-heavy, reference-favoring, or community-weighted, and visibility must be built per engine.
How much does content freshness matter for AI citations? Enormously on evolving topics: one major engine cites content under 30 days old more than 80% of the time. Dated pages with genuine monthly updates and fast event coverage systematically outperform authoritative-but-stale sources.
Why does Reddit appear so often in AI health answers? Engines weight lived-experience discussion as evidence for questions institutions answer generically, what treatment feels like, how programs compare in practice, surfacing forum content in roughly 40% of relevant answers. Brands participate only through genuine, disclosed helpfulness; astroturf is detectable and penalized.
What health questions can a brand actually win in AI answers? The territories institutions never cover: structured comparisons, process walkthroughs, current cost-and-access mechanics, honest evaluation frameworks, and fast coverage of fresh developments, all published with named clinical authorship and visible review dates.
Precision beats volume in the citation era
The first wave of AI-visibility advice said publish more. The data says publish precisely: the right territories, per-engine formats, real freshness, machine-legible trust. That is not a bigger content budget; it is a smarter content system, and it happens to be buildable by a two-person team with the right infrastructure and a calendar.
The institutions own the definitions. Everything patients ask after the definition is still up for grabs, and it is the part that sends them somewhere. Be the somewhere.