Reddit Is to GEO What DMOZ Was to SEO?

Every few years, search develops a favorite. A single source it trusts disproportionately, that an entire optimization industry then organizes itself around. For roughly a decade, that source was a volunteer-edited web directory. Today it is a discussion forum founded in 2005.
The uncomfortable part is not the comparison itself. It is what happened to the last one.
This article makes an argument that will irritate a good portion of the GEO community: Reddit occupies the same structural position in generative engine optimization that DMOZ occupied in search engine optimization, and it is likely to follow a similar trajectory over the next several years. That is not a reason to abandon Reddit. It is a reason to understand what Reddit is actually standing in for, because the thing underneath it is the part that survives.
What DMOZ actually was, and why search engines leaned on it
DMOZ, formally the Open Directory Project, launched in June 1998 under the name GnuHoo, was renamed NewHoo, then acquired by Netscape and shortly afterward by AOL. It was a hierarchical, human-edited directory of the web, maintained by volunteer editors who reviewed submissions by hand.
For over a decade it carried extraordinary weight. Search engines used its curated descriptions for result snippets because volunteer-written summaries were frequently better than what site owners supplied themselves. A DMOZ listing functioned as a trust marker. Getting one became a recognized objective, and a submission industry formed around it, complete with agencies selling placement as a line item and editors fielding a steady stream of appeals.
Then it unwound. Google's own search representatives began publicly discounting directory links as a ranking factor, and on March 14, 2017, DMOZ closed permanently after AOL declined to continue supporting it. The project's full history is worth reading as a case study in signal lifecycles.
Here is the part most retrospectives get wrong. DMOZ was not a mistake, and search engines were not foolish to rely on it. It was a genuinely useful approximation of a question engines could not yet answer directly: did a competent human look at this site and conclude it was legitimate? Early ranking systems had no reliable way to assess quality at scale, so they borrowed the judgment of people who had already done the work.
The signal expired the moment the engines could answer that question themselves. The proxy was retired because the underlying capability caught up, not because human curation stopped mattering.
Reddit is playing the same role, for a nearly identical reason
Generative engines face a version of the same problem, inverted. They are extremely capable at assessing whether text is well-formed and topically relevant, and comparatively poor at determining whether a claim reflects genuine, disinterested human experience. On a web filling rapidly with machine-generated marketing copy, that second question has become the expensive one.
Reddit is the cheapest available approximation of an answer. A pseudonymous account with years of comment history, arguing with strangers about a product it has no financial stake in, is a reasonable proxy for authentic opinion. Models lean on it heavily for exactly that reason.
The data supports how heavily. 5W's State of AI Citations 2026 research, which consolidated more than 680 million citations across six studies conducted between August 2024 and April 2026, found Reddit ranking as the single most-cited source across every major engine, appearing at roughly 40 percent frequency. Separate analysis from Otterly.ai across more than one million citations found community platforms capturing a majority of citations against brand-owned domains.
Both findings point the same direction. Engines are outsourcing the authenticity question to a platform that appears to have already solved it.
That is a proxy. A useful one, and still a temporary one.
The numbers have already started to move
The decline is not hypothetical or predicted. It is in the tracking data.
Reddit's share of total LLM citations fell by roughly half between October 2025 and January 2026, dropping from around 2.02 percent to 1.01 percent, according to Conductor's analysis of citation source trends. More recently, YouTube overtook Reddit as the leading social source for AI citations, appearing in roughly 16 percent of LLM answers against Reddit's 10 percent, a reversal from a year earlier.
Those figures look like they contradict the 40 percent number above. They do not, and the distinction matters enough to spell out, because comparing incompatible statistics is how most GEO reporting goes wrong.
The 40 percent figure measures how frequently Reddit appears somewhere in an AI answer. The 1 to 2 percent figure measures Reddit's share of total citation volume, where the denominator includes every URL cited across every query in the sample. A source can appear in a large share of answers while representing a small fraction of all citations, because most answers cite multiple sources. Before you compare any two AI visibility statistics, confirm whether each one measures presence in answers, share of total citations, or top-ten placement.
One further nuance from the same research deserves attention: when engines do cite Reddit now, they increasingly cite it as the sole source. That pattern suggests a shift toward intent-matched sourcing, where community content is selected because nothing more authoritative exists for that specific question rather than because it is a default trusted domain.
Why the decline is likely to continue
The strongest argument that Reddit's weighting keeps falling has nothing to do with model preference. It is that the proxy is being contaminated.
Roughly 15 percent of Reddit posts were likely AI-generated as of 2025, up from 13 percent the prior year, with some subreddits running far higher. More pointedly, MediaPost reported in July 2026 on organized campaigns flooding Reddit with machine-generated posts engineered to look like authentic user reviews, specifically to influence what ChatGPT and Google AI systems cite.
Follow that logic to its conclusion. Engines cite Reddit because it reads as authentically human. That citation weight creates a commercial incentive to manufacture Reddit content that reads as authentically human. As the manufactured share grows, the platform's value as a proxy for authenticity declines, regardless of how much genuine discussion remains.
This is a specific, local instance of a well-documented general problem. The 2024 Nature paper on model collapse demonstrated that training generative systems on recursively generated data degrades output diversity and accuracy. The same dynamic applies to citation sources: a source loses its informational value once its content is substantially derived from the systems citing it.
There is a commercial dimension as well. Reddit's data relationship with the major AI platforms is contractual, not permanent. Google signed a reported 60 million dollar annual licensing agreement in February 2024, with a comparable OpenAI arrangement following months later. Those agreements come up for renewal, and both sides have public leverage. A GEO strategy anchored to Reddit is partly a bet on contract negotiations you have no visibility into.
The counterargument, taken seriously
The strongest objection to this thesis is that Reddit differs from DMOZ in a way that matters: DMOZ was a static, volunteer-maintained catalog that stopped being updated, while Reddit is a live, high-volume conversation that regenerates continuously. Directories became obsolete because the web outgrew them. Human discussion does not become obsolete.
That objection is largely correct, and it is why the honest framing is a question rather than a declaration.
The counter to it is that DMOZ's value was never really its catalog structure. Its value was the human judgment embedded in it, and that judgment did not disappear either. Engines simply developed better ways to detect the same underlying quality without routing through one intermediary. The likely path for Reddit is the same: not that authentic human opinion stops mattering, but that it stops being sourced predominantly through one platform.
Note also the direction the current data already points. Video transcripts absorbing citation share is exactly what a proxy shift looks like at an early stage. Nobody announces these transitions. The weighting simply moves.
What survives every proxy shift
Something has remained constant from the directory era through the present, across every algorithm change and every reweighting of sources.
Being a clear, consistent, well-structured, authoritative source about your own subject has never once stopped working. Directory listings, link graphs, structured data, community sentiment and generative citations are all different instruments for measuring roughly the same thing. The instruments get replaced. The property they measure does not.
This is also the point where research becomes genuinely actionable. The foundational Princeton and IIT Delhi study on generative engine optimization, presented at KDD 2024 and tested across approximately 10,000 queries, found that specific content properties produced visibility gains in the range of 22 to 41 percent. The highest-performing methods were adding relevant statistics, incorporating credible quotations, and citing reliable sources. None of those techniques are platform-specific. They describe content that is verifiable, and verifiable content travels across every proxy shift.
This is also why treating SEO and GEO as competing budget lines is a mistake we see constantly. Most domains carrying serious AI citation weight today earned that position through years of conventional SEO fundamentals, and Google's own documentation on AI features makes clear that its generative surfaces draw on the same index and the same quality systems as conventional search. GEO sits on top of that foundation. It does not replace it, and it cannot compensate for its absence. They are one system with two reporting surfaces.
Control what you can actually control
Here is the practical asymmetry. You cannot control Reddit's weighting inside a model. You cannot control whether a licensing agreement renews. You cannot control whether video transcripts absorb another five points of citation share next quarter.
You can control whether AI systems have a clean, current, machine-readable account of who you are and what you sell.
This is the thinking behind the Silverback AI Readiness Kit: 18 files deployed to your site root that give ChatGPT, Claude, Perplexity and Gemini a direct, structured, first-party description of your business, following emerging conventions including the llms.txt proposal. It is free, and deployment takes an afternoon.
Two things it deliberately does not do, stated plainly. It does not change who you are or what services you offer. It describes accurately what already exists, which is precisely why it works. And it does not replace agencies, vendors or people. It handles the mechanical layer so that human effort goes toward the part no tool can produce: knowing what your company is genuinely the best answer for and saying it in language a real buyer finds credible.
The value becomes clearest in the failure case. We regularly see engines quote a promotional price a business ran a year ago, pulled from an old forum thread, long after the actual price changed. Nobody on that thread did anything wrong. The engine simply used the most specific source available, because the business never published a more current one in a form the engine could read. Given a stale third-party mention and an authoritative first-party statement, engines generally prefer the latter. Most businesses have never supplied one.
Organic mentions and earned community presence still matter enormously, and we would never argue otherwise. But earned signals age in place with no correction mechanism, and there is no reason to let a year-old comment speak for your pricing when you can state it yourself. If you want that foundation built deliberately alongside your search program, that is the core of our SEO and GEO work.
The practical version
Do not abandon Reddit. Do stop treating a citation source as a strategy.
Three things worth doing this quarter:
- Audit which sources actually feed AI answers in your category right now, not eighteen months ago, and re-run that check quarterly, because the composition shifts by engine and by month.
- Verify that every material fact about your business, pricing, positioning, service scope and location, is published somewhere current, structured and machine-readable under your own control.
- Invest in the earned signals that hold value regardless of which platform is favored: original research, documented outcomes, genuine expertise published openly.
DMOZ taught the SEO industry an expensive lesson about mistaking a proxy for the thing it measures. The lesson is available again, at a discount, for anyone willing to notice the pattern the second time around.
Frequently Asked Questions
Common questions about GEO, SEO, and AI-driven search visibility.
Yes, but less than it was. Reddit remains one of the most frequently cited domains across ChatGPT, Perplexity, Gemini and Google AI Overviews, and consolidated research still places it among the highest-frequency sources. However, tracking data shows its share of total LLM citations fell by roughly half between October 2025 and January 2026, and YouTube has overtaken it as the most-cited social source. Treat Reddit as one input to a broader strategy rather than the strategy itself.
DMOZ, the Open Directory Project, was a human-edited web directory launched in 1998 that search engines used for over a decade as a trust and description signal. An entire SEO submission industry grew around it before Google downplayed directory links and the project closed in March 2017. It matters to GEO because it demonstrates the full lifecycle of a proxy signal: adopted because it approximated human judgment, then discarded once engines could evaluate quality more directly.
Three forces are working together. Reddit is being contaminated with AI-generated posts, with roughly 15 percent of posts likely machine-written as of 2025 and documented campaigns seeding fake reviews specifically to influence AI answers. Models are also shifting toward intent-matched sourcing rather than volume-based sourcing. Finally, competing formats such as video transcripts are absorbing citation share that community forums previously held.
They often appear to, because different studies measure different things. Some report how often Reddit appears somewhere in an AI answer, which lands near 40 percent. Others report Reddit's share of total citation volume, which is closer to 1 to 2 percent because that denominator includes every URL cited across every query. Always check whether a statistic measures presence in answers, share of citations, or top-ten placement before comparing two numbers.
No. The correct response is to stop treating a single citation source as a strategy. Genuine community presence still produces real value and still gets cited, particularly in categories where no authoritative alternative exists. The mistake is building a GEO program whose results depend entirely on one platform's weighting inside models you do not control.
The domains carrying the most AI citation weight generally earned that position through years of conventional SEO fundamentals: authority, crawlability, structure and consistency. Generative engines draw heavily from conventional search indexes, so ranking well feeds AI visibility directly on several platforms. Splitting SEO and GEO into competing budgets usually means funding a moving target while starving the foundation underneath it.
References
All statistics and data points cited in this article link to their original sources.
- Search Engine Land: RIP DMOZ, the Open Directory Project is closing
- Wikipedia: DMOZ
- 5W Public Relations: The State of AI Citations 2026
- Otterly.ai: The AI citation economy, what 1 million data points reveal
- Conductor: Reddit's AI citation decline and how brands win back visibility
- PikaSEO: YouTube overtakes Reddit as the top social source for AI citations
- Originality.AI: 15 percent of Reddit posts are likely AI generated
- MediaPost: Reddit infiltrated by stealth AI in brand citation race
- Nature: AI models collapse when trained on recursively generated data
- arXiv: GEO, Generative Engine Optimization (Princeton and IIT Delhi)
- Tom's Guide: Google strikes 60 million dollar deal with Reddit for AI training data
- Quartz: Reddit stock drops as Google AI content deal nears expiration
- Google Search Central: AI features and your website
- llmstxt.org: The llms.txt proposal