The AI Citation Paradox

By Jeff Jenkins

The AI Citation Paradox

The organizations with the most accurate information are often the least likely to be cited by AI. Here's why regulated institutions lose citations to intermediary sites.

<p><em>The organizations with the most accurate information are often the least likely to be cited by AI.</em></p><p>On July 22, 2026, I asked Gemini a simple question: if my fintech app fails, is my money still FDIC insured?</p><p>The answer was good.</p><p>It correctly explained pass-through deposit insurance. It distinguished between bank failure and fintech failure. It explained why incomplete record-keeping can leave people waiting months for their own money.</p><p>Then it linked to the FDIC exactly once.</p><p>Not to explain the rule. To verify a bank using BankFind.</p><p>A place to check the answer. Not the source of it.</p><p>The knowledge was there. The institution was treated as a place to verify the answer, not the place that explained it.</p><p>If AI consistently explains regulated topics using secondary interpreters rather than the organizations that write the rules, that is bigger than a visibility problem. That becomes an information quality problem for everyone.</p><p><strong>TL;DR</strong></p><ul><li><p><strong>The paradox:</strong> The institutions with the most authoritative information are often the hardest for AI to cite.</p></li><li><p><strong>The reason:</strong> This is structural, not negligence. They publish for liability and permanence, not for answers.</p></li><li><p><strong>Where it hides:</strong> The information usually exists. It sits in PDFs, fragmented across documents, or behind logins. In many organizations, it has also been written by committee rather than as a direct answer.</p></li><li><p><strong>The consequence:</strong> When primary sources are hard to retrieve, AI assembles answers from whoever explained them more clearly.</p></li><li><p><strong>The takeaway:</strong> This is a publishing problem, not an SEO problem, which is why SEO advice has not fixed it.</p></li></ul><h2>Three Vocabularies, One Problem</h2><p>Part of the reason this goes undiagnosed is that nobody involved uses the same word for it.</p><p>Marketers talk about citations. AI researchers talk about retrieval and grounding. Buyers use neither word.</p><p>They ask why ChatGPT keeps mentioning a competitor.</p><p>They are all describing the same phenomenon from different positions. <a target="_blank" rel="noopener noreferrer nofollow" class="text-primary underline" href="https://cloud.google.com/use-cases/retrieval-augmented-generation">Google's own documentation defines retrieval-augmented generation</a> as "a technique (also known as grounding)" that retrieves relevant, up-to-date web pages before generating a response.</p><p>Retrieval.</p><p>Grounding.</p><p>Citation.</p><p>One event. Three dialects.</p><p>The result is that a compliance officer, a marketing lead, and an engineer can all see the same problem without realizing they are talking about the same thing. So nobody owns it. And it does not get fixed.</p><h2>Why This Is Structural, Not Negligence</h2><p>The easy explanation is that these organizations are behind. That they do not care about marketing, or have not noticed AI, or cannot get budget approved.</p><p>That explanation is wrong, and it is easy to check.</p><p>Search Google for whether Medicare covers a hospital bed for home use. <a target="_blank" rel="noopener noreferrer nofollow" class="text-primary underline" href="http://Medicare.gov">Medicare.gov</a> ranks first, and Google's AI Overview cites it as a source.</p><p>That is not a failure. That is the system working.</p><p>Medicare publishes a page called <a target="_blank" rel="noopener noreferrer nofollow" class="text-primary underline" href="https://www.medicare.gov/coverage/hospital-beds">Hospital Beds</a>. It sits at a stable address, answers the question directly, and is written as a web page rather than buried inside a policy document. So it gets found, and it gets credited.</p><p>Which means the problem is not discoverability, and it is not that AI refuses to cite institutions.</p><p>The outcome is uneven. And the unevenness tracks something specific: how the information was published.</p><p>It starts with what these documents were built to do.</p><p>A bank publishes a disclosure so it survives an audit. A regulator publishes a rule so it holds up in court. A hospital publishes guidance so it reflects clinical consensus. These documents are written for permanence, for precision, and for defensibility. They are reviewed by people whose job is to ensure nothing in them can be misread.</p><p>That is the correct output for the obligation.</p><p>A regulator's job is to be right, not to be quotable.</p><p>But the format that satisfies an auditor is rarely the format that yields a clean answer to a question someone typed thirty seconds ago. Neither system is wrong. They simply want different things from the same document.</p><h2>The Three Questions</h2><p>Before an answer gets generated, something has to retrieve the material that answer will be built from. It happens before a single word of the response is written.</p><p>You do not need to know how any particular system does it. It helps to think of it as three questions.</p><p>Can I retrieve this?</p><p>Can I understand this?</p><p>Can I trust this?</p><p>Retrieval is about access. Is the content reachable, and is it in a form that can be pulled apart into usable pieces?</p><p>Understanding is about clarity. Once a passage is pulled out, does it still make sense on its own, or does it depend on forty pages of surrounding context?</p><p>Trust is about authority. Is this a source worth relying on?</p><p>Here is where the paradox stops being mysterious and becomes mechanical.</p><p>Regulated institutions pass the third question easily. They are among the most trustworthy sources that exist on their own subjects. No blog outranks the FDIC on trust.</p><p>They often struggle with the first two.</p><p>That is worth sitting with. The organizations most likely to be trusted are often the least likely to be retrieved and understood. And a source that cannot be retrieved does not get to be trusted, because it never enters the conversation at all.</p><p>You can see this priority in real retrieval systems.</p><p><a target="_blank" rel="noopener noreferrer nofollow" class="text-primary underline" href="https://github.com/openai/openai-knowledge-retrieval">OpenAI's reference implementation</a> for knowledge retrieval splits documents by heading hierarchy so sections remain coherent, and it evaluates responses for groundedness.</p><p>In other words, structure is not decoration. It determines whether a passage survives retrieval.</p><p>Every failure mode in the next section is really just one of these three questions failing.</p><h2>The Failure Modes</h2><p>Here is the part that surprises people.</p><p>None of these failures are about AI.</p><p>Nielsen Norman Group has studied how people use online documents since the 1990s. They ran the same research on PDFs in 2001, 2003, 2010, and 2020, and reached the same conclusion every time. <a target="_blank" rel="noopener noreferrer nofollow" class="text-primary underline" href="https://www.nngroup.com/articles/pdf-unfit-for-human-consumption/">Their 2020 summary</a> is blunt: "Do not use PDFs to present digital content that could and should otherwise be a web page."</p><p>In one of those studies, a user was trying to find where Wells Fargo was headquartered. He ended up in a PDF, gave up, and said: "All of a sudden that downloads a PDF, ew! At this point, I would just Google it."</p><p>A bank. A simple question. The answer existed. It was in a PDF. The user left.</p><p>That was 2020, before AI search existed in any meaningful form.</p><p>These failure modes have been costing regulated institutions readers for twenty-five years. AI did not create the problem. AI made it expensive.</p><p>Search engines could still succeed by returning a list of documents. Answer engines have to extract a passage they can understand and stand behind.</p><h3>Failures of Retrieval</h3><p><strong>PDFs.</strong> The answer exists, but it sits inside a container designed for printing. Nielsen Norman Group found that most PDF files have no internal navigation and that their structure is effectively hidden. A document built to be printed and archived is not built to have one paragraph lifted cleanly out of it.</p><p><strong>No canonical answer page.</strong> The information is public, but it is spread across a policy document, an FAQ, a rate sheet, and a footnote, with no single page that answers the question the way a person actually asks it. Technical writing solved this decades ago and named the solution: single sourcing. Maintain one authoritative version of each topic and reuse it everywhere.</p><p>Nielsen Norman Group recommended the same fix in 2001 and called it a gateway page: a plain web page stating the key information, with the full document offered as a download. Most regulated publishing still does the opposite.</p><p><strong>Gated content.</strong> The most complete explanation sits behind a login, a registration form, or a portal. Google explicitly states that pages must be crawlable and indexable to be eligible for its generative features. Gated content does not fail at the trust question. It never reaches it.</p><h3>Failures of Understanding</h3><p><strong>Committee writing.</strong> When six people must approve a sentence, the sentence stops making a claim. Each reviewer softens one more word. What survives is technically unobjectionable and says almost nothing. It reads as consensus rather than as an answer.</p><p><strong>Dense legal drafting.</strong> Precision written for a court is not clarity written for a reader. Both are legitimate. Only one of them can be quoted.</p><p><strong>Context dependence.</strong> The paragraph is clear, but only if you have read the previous six pages. Pulled out on its own, it means very little. Regulated documents are often built as arguments that accumulate, which is exactly what makes any single passage hard to lift.</p><h3>Failures of Trust</h3><p>There are almost none.</p><p>That is the finding hiding inside this list. Go back through the failures and nearly every one breaks retrieval or understanding. Regulated institutions are not struggling to be credible. Credibility is the one thing they have in abundance.</p><p>Trust was never the bottleneck.</p><h2>The Same Pattern, Across Industries</h2><table style="min-width: 75px;"><colgroup><col style="min-width: 25px;"><col style="min-width: 25px;"><col style="min-width: 25px;"></colgroup><tbody><tr><th colspan="1" rowspan="1"><p>Institution</p></th><th colspan="1" rowspan="1"><p>What makes it trustworthy</p></th><th colspan="1" rowspan="1"><p>What makes it hard to retrieve</p></th></tr><tr><td colspan="1" rowspan="1"><p>Government agency</p></td><td colspan="1" rowspan="1"><p>Statutory authority</p></td><td colspan="1" rowspan="1"><p>Guidance split across dozens of pages</p></td></tr><tr><td colspan="1" rowspan="1"><p>Bank</p></td><td colspan="1" rowspan="1"><p>Regulatory oversight</p></td><td colspan="1" rowspan="1"><p>Terms buried in disclosure documents</p></td></tr><tr><td colspan="1" rowspan="1"><p>Hospital</p></td><td colspan="1" rowspan="1"><p>Clinical review</p></td><td colspan="1" rowspan="1"><p>Guidance published as PDFs</p></td></tr><tr><td colspan="1" rowspan="1"><p>Law firm</p></td><td colspan="1" rowspan="1"><p>Precision of language</p></td><td colspan="1" rowspan="1"><p>Answers written as legal drafting</p></td></tr><tr><td colspan="1" rowspan="1"><p>Insurer</p></td><td colspan="1" rowspan="1"><p>Compliance requirements</p></td><td colspan="1" rowspan="1"><p>Policy detail behind a login</p></td></tr></tbody></table><p>Notice that the second column is not the opposite of the first.</p><p>It is the same trait, viewed from a different angle. The regulatory oversight that makes a bank trustworthy is what buries its terms in disclosure documents. The precision that makes a law firm credible is what makes its answers hard to extract. The clinical review that makes a hospital authoritative is what pushes its guidance into PDFs.</p><p>The strength and the obstacle are the same property.</p><h2>Who AI Chooses to Cite, and Who Gets Left Out</h2><p>If the institutions are hard to retrieve, someone else answers the question.</p><p>Across systems, AI-generated answers tend to draw from a handful of recurring source types. Encyclopedic references. Publishers and news media. Independent explainers and specialist blogs. Industry associations. Product documentation and support centers. Government and institutional pages, when the answer sits on a page rather than inside a document. Community discussion.</p><p>That grouping reflects what shows up in the source panels of AI answers and in the domain analysis published in <a target="_blank" rel="noopener noreferrer nofollow" class="text-primary underline" href="https://arxiv.org/abs/2603.16138">Answer Bubbles</a>, a 2026 study of roughly 11,000 queries across four systems.</p><p>The mix is not the same everywhere. That study compared vanilla GPT, SearchGPT, Google AI Overviews, and traditional Google Search, and found that each system favored different sources for the same queries. The researchers called the result answer bubbles: identical questions producing different information realities depending on where they were asked.</p><p>Which matters before anyone optimizes for a single engine. Being cited in one is not the same as being cited in another.</p><h2>Who Gets Cited Instead</h2><p>The intuitive answer is that thin content wins. That shallow, fast, SEO-driven pages beat careful institutional writing.</p><p>The data does not support that.</p><p>The study also found that Wikipedia and longer-form sources were disproportionately represented in AI-generated answers relative to traditional search results. The sources winning citations are often more comprehensive than what they displaced, not less.</p><p>So the winner is usually not less rigorous.</p><p>It is simply more retrievable.</p><p>And very often, the winner is an intermediary whose entire business is explaining the institution.</p><p>Intermediaries do not win because they are more trustworthy. They win because they pass the first two tests. Easy to retrieve. Easy to understand. And trusted enough.</p><p>NerdWallet explains banks. Healthline explains hospitals. Investopedia explains regulation. Their competitive advantage is not greater authority. It is translation.</p><p>The institution publishes the rule.</p><p>The intermediary publishes the translation.</p><p>Which brings us back to that Gemini answer. The FDIC's rules were the whole explanation. The FDIC itself appeared once, as a tool to go check something with.</p><p>One publishes documents.</p><p>The other publishes answers.</p><h2>What It Costs When the Most Accurate Sources Cannot Be Cited</h2><p>Everything above this point is a marketing problem. A visibility problem. Something a team could be assigned to fix next quarter.</p><p>This part is not.</p><h3>The Qualifications Disappear First</h3><p>Regulated writing is full of conditions. Up to $250,000 per depositor, per bank, per ownership category. Subject to underwriting. Coverage varies by state. Talk to your physician.</p><p>Those conditions are not padding. In a regulated answer, they frequently are the answer. The difference between "your money is insured" and "your money is insured up to $250,000 per depositor, per bank, per ownership category" is the difference between a true statement and a false one.</p><p>The Answer Bubbles study found that incorporating search into generative systems reduced hedging language by up to 60 percent, while confidence language was preserved.</p><p>Read that twice.</p><p>The uncertainty gets stripped. The certainty stays.</p><p>An answer can become less complete and more confident at the same time. In consumer finance, healthcare, and law, that is the combination that hurts people.</p><h3>The Institution Is Not Absent. It Is Uncredited.</h3><p>Here is the part that took me a while to see.</p><p>The institution's knowledge usually is in the answer. The FDIC's rules were in the response I read. The coverage limits were right. The mechanism was explained accurately.</p><p>It simply arrived secondhand.</p><p>Which means a reader who wants to verify it cannot easily trace it back. The chain from claim to source is broken, not because anyone hid anything, but because the source was offered as a place to check rather than credited as the origin of the answer.</p><p>Institutions are not being ignored. They are being used without attribution.</p><p>That distinction matters, because it changes what is actually at risk. It is not the institution's reputation. It is the reader's ability to check.</p><h3>Someone Acts On It</h3><p>In 2024, <a target="_blank" rel="noopener noreferrer nofollow" class="text-primary underline" href="https://www.yalejournal.org/publications/the-synapse-collapse">Synapse</a>, a banking-as-a-service company sitting between fintech apps and their partner banks, collapsed. Thousands of people lost access to their money for months because no insured bank had failed and the records showing who owned which dollars were incomplete.</p><p>People who believed their money was FDIC insured discovered that the real answer was more qualified than the version they had been given.</p><p>That is what a stripped qualification costs. Not a ranking. Not a dashboard.</p><p>The person who read the confident version and acted on it.</p><p>This is why the paradox is more than a marketing story. When the organizations that write the rules cannot be cited, the answers people depend on are assembled from translations of translations. Most of the time the translation is good. The problem is that nobody, including the reader, can tell which time is the exception.</p><h2>This Is Fixable</h2><p>Nothing in this article argues that regulated institutions should be less careful.</p><p>The disclosures exist for reasons. The review process exists for reasons.</p><p>The problem is not precision.</p><p>The problem is treating precision and retrievability as opposites.</p><p>They are not.</p><p>An institution can keep every word its lawyers require and still publish something a machine can lift cleanly. The technical work is straightforward. The organizational decision is harder.</p><p>Lead with the answer.</p><p>Put the qualifier after it.</p><p>Give the question its own page.</p><p>Take the answer out of the PDF.</p><p>I wrote about those mechanics in detail in <a target="_blank" rel="noopener noreferrer nofollow" class="text-primary underline" href="https://discoveraio.com/articles/winning-ai-citations-when-legal-reviews-every-sentence">Winning AI Citations When Legal Reviews Every Sentence</a>.</p><p>What comes next in this series is the other half of the story. The trust assets these institutions already own but have never needed to think about as trust assets. And why the constraint they have resented for years may turn out to be the advantage.</p><h2>FAQ</h2><h3>How do AI engines decide which sources to cite?</h3><p>They retrieve candidate material before generating an answer, then build the response from passages they can extract cleanly. Authority matters, but a source has to be reachable and understandable on its own before it is considered at all. Different systems weigh this differently, which is why the same question can produce different citations in different tools.</p><h3>Why is my website not being cited by AI?</h3><p>The most common reasons are failures of retrieval and understanding, not trust. The answer may sit inside a PDF, split across several documents, behind a login, or written so that it only makes sense with the surrounding pages. Start by checking whether a single public page answers the question the way someone would actually ask it.</p><h3>What are AI Overview sources?</h3><p>They are the web pages an AI-generated answer draws from and links to. They tend to include encyclopedic references, publishers, specialist explainers, industry associations, product documentation, and institutional pages, though the mix varies by question and by system. Appearing in one system's sources does not mean appearing in another's.</p><h3>What content gets cited most often by AI?</h3><p>Research on generative search has found that encyclopedic references and longer sources are disproportionately represented. Thin content is not the winner. Content that states an answer directly and holds up when read on its own tends to do better.</p><h3>How can I see which sources AI uses for my industry?</h3><p>Ask the questions your customers ask across several systems and record which domains appear. Ten or fifteen questions is usually enough to see the pattern, including whether the primary sources in your field are being cited at all. Tools exist for tracking this at scale, but the manual version will show you the shape of the problem.</p><h2>Who This Is For</h2><p>This article is written for people inside organizations whose accuracy is their credibility. Banks, hospitals, insurers, agencies, and firms where being right is the product, and where the review process exists because the stakes are real.</p><p>If that is your world, the paradox is not a criticism of your standards.</p><p>It is a description of what those standards currently cost you in a system that reads differently than a person does.</p><p>None of this is an argument for lowering the bar.</p><p>The goal is not to make institutions less rigorous.</p><p>It is to publish their rigor in a form that can be retrieved, understood, and cited.</p><p>Because when authority cannot be retrieved, someone else translates it.</p><p>And the translation becomes the answer.</p>