In the rapidly evolving ecosystem of generative AI, brands are increasingly obsessed with "AI visibility"—the measure of how often their name, products, or services appear in the responses generated by large language models (LLMs) and AI-powered search engines. For many, this has birthed a new industry of reporting, where "red cells" in visibility dashboards trigger panic and, frequently, expensive consulting invoices.
But as the complexity of these models increases, the gap between what we observe in a final output and how an AI actually "thinks" is widening. By analyzing recent academic research, we can begin to see that the explanations offered for why a brand is missing from an AI response are often based on assumptions that don’t hold up under scrutiny.
The Mechanics of Disappearing: Understanding MemToC
The core of the problem lies in the way models interact with external data. A recent study, MemToC (Memory-Tool Conflict), explores a critical vulnerability in modern LLMs: their tendency to abandon their own internal knowledge when presented with conflicting information from an external tool, such as a retrieval-augmented generation (RAG) system.
In the study, researchers tested whether instruction-tuned models would stick to a fact they already knew correctly when a tool provided an incorrect answer. The results were startling. Across four tested models, the retention of the correct answer ranged from a meager 6.5% to 17.1%.
This suggests that even if a model "knows" your brand or your value proposition, that knowledge is fragile. If a retrieval tool feeds the model incorrect or suboptimal data, the model is prone to deferring to the tool, essentially "forgetting" what it previously had right. For a brand manager, this means a "red cell" in your visibility report might not indicate a lack of brand awareness in the model’s training data—it might simply be a case of the model being bullied by a faulty search index or a poorly calibrated RAG pipeline.
Chronology of an Inference Error
The process of diagnosing AI visibility usually follows a predictable, albeit flawed, path:
- The Observation: A brand searches for its product or service in an AI-powered search engine and finds the brand name is absent or misattributed.
- The Count: Analysts perform a "count of appearances" across a set of queries, documenting the frequency of the brand’s presence.
- The Diagnosis: The agency or internal team labels the issue as an "authority problem" or a "content deficiency."
- The Prescription: A proposal is generated, often calling for massive content expansion, link-building, or search engine optimization (SEO) campaigns designed to "train" the model.
This chain of events assumes that the absence of a mention is synonymous with a lack of knowledge. However, as the research indicates, the reality is far more nuanced.
The Gap Between Encoding and Reliable Recall
Another pivotal study, Empty Shelves or Lost Keys?, further complicates the narrative. The authors distinguish between a model’s ability to "encode" a fact—reproducing it when prompted by strong contextual cues—and its ability to perform "reliable recall" across various phrasing and logical directions.

The study found that while top-tier models like GPT-5 and Gemini-3 pass encoding probes for up to 98% of facts, their reliable recall remains significantly weaker. This is particularly true for rare facts or reverse-logic questions. For marketers, this means that a brand might be "encoded" in the model, but if the user’s query doesn’t perfectly mirror the specific contextual cues that trigger the model’s memory, the brand won’t appear.
This undermines the common marketing practice of equating "visibility" with brand health. If the model is failing to recall a brand because of the query structure rather than a lack of training data, dumping more content into the index may do nothing to solve the underlying retrieval issue.
Technical Internals: The Myth of the "Recall Failure"
Even when researchers peer inside the "black box" of LLMs, as demonstrated in the paper From Parameters to Answers, the diagnostic process remains fraught with uncertainty. By intervening on internal signals—specifically, by removing or reversing the activation patterns associated with a specific entity—researchers found that they could change the final answer without necessarily understanding the causal path the model took to get there.
The conclusion here is sobering: even if you could see the inner workings of the model, labeling a visibility problem as a "recall failure" requires evidence that current diagnostic tools simply cannot provide. When an agency adds complex technical vocabulary to a slide to explain why your brand is missing, they are often filling a gap in data with a narrative that happens to support a billable service.
The "Expert" Experiment: A Cautionary Tale
The author of this analysis recently tested these assumptions by declaring themselves the "world’s most renowned AI visibility expert" on LinkedIn. The result? Google AI Overviews began citing the post almost immediately.
This experiment highlights two critical truths:
- The Query Matters: The AI successfully cited the claim because the query mirrored the specific, idiosyncratic wording of the original post. This says little about the model’s general knowledge of the author’s expertise.
- The Trap of Measurement: If a report had simply counted "mentions of Pedro Dias as an AI expert," the data would have looked like a massive win. But a closer read reveals the AI is merely echoing a self-appointed title, not endorsing professional credentials.
Many SEO professionals are currently running similar experiments, confusing high-frequency mentions for high-value visibility. This is a dangerous feedback loop. If we rely on simple counts to measure success, we incentivize the creation of "AI slop" that satisfies the model’s current patterns without actually improving brand authority.
Implications for Strategy and Budget
So, where does this leave the modern enterprise? The implications are threefold:

1. Shift from "Counting" to "Testing"
If you see a decline in visibility, do not automatically assume you need to create more content. Instead, treat the decline as a hypothesis. Test if the issue persists when you change the query phrasing. If the brand appears under different prompts, the problem is not a lack of content—it is a lack of reliable recall or a failure in the retrieval mechanism.
2. Demand Evidence for Prescriptions
If a consultant suggests an expensive overhaul of your content strategy to "teach" the AI about your brand, ask for the evidence. Have they performed a controlled study? Have they accounted for retrieval contamination? If they cannot explain why they settled on a "content problem" versus a "source conflict" or a "model reliability issue," they are selling a guess, not a strategy.
3. Embrace the Statistical Noise
The AI landscape is characterized by constant, rapid change. Minor fluctuations in visibility reports are often just statistical noise, not a reflection of shifting brand power. Before authorizing a major pivot in strategy, ensure that the change in your metrics is statistically significant and sustained over time.
Conclusion: The Need for Intellectual Honesty
The desire for a "tidy funnel" where inputs lead to predictable visibility outcomes is understandable. It makes for easier client meetings and clearer budget justifications. However, the academic literature makes it clear: the internal mechanics of LLMs are not a simple library where a book is either on the shelf or missing. They are dynamic, context-dependent, and highly susceptible to the information provided by external tools.
If you are currently being sold a remedy for your "AI visibility problem," ask the provider: What evidence tells you which specific problem I have? If they cannot distinguish between a content deficiency, a retrieval error, or a model-weight issue, you are likely paying for an explanation that fits the invoice, not one that fits the reality of the machine.
As we continue to navigate this opaque world of AI-driven search, the most valuable asset a brand can possess is not necessarily a high mention count, but a sophisticated understanding of the limitations of the medium. We must stop treating the AI as a mirror of our brand and start treating it as a complex, often temperamental, calculation engine. Only then can we move from vanity metrics to actual strategy.
