AI Amplifies the Foundation You Give It. Your Website Is That Foundation.

There is a diagram going around that puts two buildings side by side. On the left, a gleaming tower marked AI, balanced on cracked concrete and rusted props, with a dial reading UNCERTAIN. On the right, the same tower on solid pillars labeled governance, lineage, ownership, data quality, metrics and trust. Underneath, one line: AI amplifies the foundation you give it.
It is a data engineering picture. Warehouses, pipelines, dashboards. But it describes something most marketing teams have not fully registered yet, which is that your website has quietly become a data layer. Not a brochure, not a campaign asset. The source material that AI systems read, summarize and repeat when someone asks about your company.
That changes the job. A typo used to be embarrassing. A contradiction between your pricing page and your PDF used to be untidy. Now those are data quality defects in a system that will confidently restate them to a prospect who never visits your site at all. This piece is about what that means in practice, which of the six pillars actually translate to SEO, and where humans still have to stand in the loop.
The evidence that AI amplifies whatever it is given
We should be specific rather than dramatic, because the numbers are strong enough on their own.
In October 2025, the European Broadcasting Union coordinated research led by the BBC in which professional journalists from 22 public service media organizations across 18 countries evaluated more than 3,000 AI assistant responses from ChatGPT, Copilot, Gemini and Perplexity. The findings, published in full in the News Integrity in AI Assistants report, were that 45% of answers contained at least one significant issue. Thirty-one percent had serious sourcing problems, meaning missing, misleading or incorrect attribution. Twenty percent had major accuracy issues, including hallucinated details and outdated information. Gemini fared worst, with significant issues in 76% of responses, largely on sourcing.
Sit with what that sample was. News from public service broadcasters is among the most carefully structured, clearly bylined, well sourced content on the internet. If AI assistants misattribute and garble that material at those rates, the odds that they handle your product specifications, pricing tiers and compliance language cleanly are not good.
Meanwhile the audience keeps growing. The Reuters Institute's Digital News Report 2026 found weekly use of AI chatbots for news rose from 7% to 10% globally in a year, and that 52% of 18 to 24 year olds now name social media, video networks and AI chatbots as their main way of getting news. Trust in chatbot answers sits at 20%, against 37% for news overall, which is a useful reminder that people are using these tools while not entirely believing them.
There is a business version of this concern too. Stanford's AI Index consistently finds inaccuracy ranked among the risks organizations cite most about their own AI use. The interesting part is that most of those organizations are worried about their internal models. Very few have looked at the public content that external models are already reading about them.
The six pillars, translated
The graphic's labels come from data governance. Here is what each one actually means when the data set is your website.
| Pillar | In a data team | On your website |
|---|---|---|
| Governance | Defined ownership and approval paths | Who is allowed to publish a claim, and which claims need review first |
| Lineage | Every field traceable to its source | Every statistic cited inline, with a source and a date |
| Ownership | A named data steward | A named author and editor who answer for the page |
| Data quality | Accuracy, freshness, deduplication | Figures checked against reality and retired when they go stale |
| Metrics | Pipelines monitored end to end | AI visibility measured beyond Search Console, starting with server logs |
| Trust | The output of everything above | One canonical version of every fact, stated the same way everywhere |
Governance: who is allowed to publish a claim
In a data team, governance means defined ownership and approval paths. On a website it usually means nothing at all. Anyone with CMS access can publish a statistic, and nobody records where it came from.
The fix is unglamorous and effective. Decide which categories of claim need review before publishing, typically anything involving numbers, pricing, compliance, safety or competitor comparison. Everything else can move fast. Google's own spam policies target content produced at scale without adding value, and the practical guardrail against drifting into that territory is a human deciding what is worth publishing at all.
There is at least one place where this is a hard requirement rather than good practice. Google Merchant Center policy states that AI-generated images must carry IPTC DigitalSourceType TrainedAlgorithmicMedia metadata, and that AI-generated product data such as titles and descriptions must be specified separately and labeled as AI-generated. If you sell products and you are generating feed content with AI, that is a compliance obligation sitting in your content workflow right now.
Lineage: can a fact be traced back
Lineage is the pillar most websites fail hardest, and the one AI systems punish most visibly. The EBU and BBC found sourcing failures in nearly a third of responses. If assistants struggle to attribute correctly even when sources exist, publishing unsourced assertions and hoping they are carried accurately is optimistic.
Cite inline, link to primary sources, and name the date a figure refers to. This is not just defensive. It gives a model something concrete to ground against, and it gives a reader something to verify. When we publish a statistic at Silverback, it carries a link and a date, which is why you can check every number in this article.
Ownership: a name and a person behind it
Ownership in data terms is a steward. In content terms it is a named author, a named editor and a real bio. Google's people-first content guidance has emphasized first-hand experience and clear authorship for years, and that has not softened in the AI era.
The deeper point is accountability. An author byline is a person saying they will answer for this. No model can do that, which is precisely why it still matters.
Data quality: accuracy and decay
Content does not stay true. Prices change, regulations shift, statistics age. The number you published in 2024 is still on the page, still being crawled, and still being repeated as current by systems with no way of knowing it expired.
This is the failure mode we see most often in audits, and it is invisible in analytics because nothing breaks. The page still ranks. It is just wrong. For anything in a YMYL category, finance, health, legal, insurance, a stale figure is a trust and compliance risk, not a housekeeping issue.
Metrics: measuring what you cannot see
Search Console now reports impressions from Google's generative AI features, and Google's own guidance points you there as the place to measure AI visibility. That is genuinely useful, and it tells you nothing about ChatGPT, Perplexity or Claude, because Google cannot see them.
It is also worth knowing that ranking is a weaker proxy for citation than it used to be. Ahrefs data reported by Search Engine Journal across 863,000 keywords and 4 million AI Overview URLs found 38% of cited pages also ranked in Google's top 10, down from 76% in the earlier version of the study, though Ahrefs notes its detection method improved between the two and the figures are not strictly comparable. Either way, "we rank well" no longer answers the question "are we being cited."
Your server logs remain the most honest thing you own. They show which AI agents requested which pages and what they got back. We wrote about the access side of this in more detail in our piece on optimizing for AI search beyond Google.
Trust: consistency across every surface
Trust is the output of the other five, plus one thing: saying the same thing everywhere. One canonical version of your address, your pricing model, your product names, your founding date. Structured data helps here by stating those facts unambiguously rather than leaving them to be inferred from prose.
When your own sources disagree, a model has to pick. It will not flag the conflict to the reader. It will simply choose, and you will not know which one it chose.
Where humans have to stand in the loop
This is the part the graphic leaves out, and it is the part we care most about.
Every pillar in that picture looks like infrastructure, as though governance and trust are things you install. They are not. Governance is a person deciding what gets published. Lineage is a person choosing to cite. Ownership is a name. Data quality is someone checking a number against reality. Trust is someone taking responsibility. The machinery in the diagram is drawn by people.
Google's guidance on generative AI content lands in the same place. It says AI is useful for research and for adding structure to original content, then makes the standard clear: your work has to meet Search Essentials and the spam policies, and it points readers at the quality rater criteria covering content made with little effort, little originality and little added value. Nothing about that is anti-AI. It is a statement that the tool does not carry the responsibility.
The Reuters Institute puts it plainly in its 2026 report, observing that across newsrooms adopting generative AI deeply into workflows and products, the mantra of human in the loop still holds. If that is true for organizations with dedicated standards teams and legal review, it is true for a marketing team shipping a product page on a Thursday.
Our own position, from doing this work daily: AI is genuinely excellent at the parts humans are bad at. It will read 400 articles and flag every decaying statistic in an afternoon, which no person will ever volunteer to do. It will spot that three pages describe your pricing differently. It will draft a structure in a minute that would have taken an hour. Used that way it is one of the most useful things to arrive in this industry in a decade.
What it will not do is know that the number it found is wrong, that the claim needs legal sign-off, or that the phrasing is technically accurate and commercially disastrous. Almost everything we ship gets meaningful human intervention before a client sees it, and the sequence is always the same. Machine finds, human decides. Reverse that order and you get volume, speed, and the confident restatement of things that are not true.
Where to start, honestly
You do not need a governance program. You need a short list of facts that matter and one true version of each.
Pull the claims that appear most often and cost most when wrong: pricing, specifications, contact and location details, compliance statements, and the headline statistics your sales team repeats. Find every place each one appears. Make them agree. Publish the canonical version somewhere authoritative on your own site. In our experience this single exercise resolves more AI misattribution than a quarter of content production, because it removes the contradictions models are currently choosing between.
Then work through the decay. Every page with a figure, a date or a rate gets checked against a current source and stamped with the date it was verified. Use AI to find them, use people to fix them.
Then, and only then, worry about producing more. A larger site built on the same shaky foundation just gives the amplifier more to work with.
- List the facts that cost most when wrong. Pricing, specifications, contact and location details, compliance statements, and the headline statistics your sales team repeats.
- Make every mention agree. Pick one canonical version of each fact, publish it somewhere authoritative on your own site, and bring every other mention in line with it.
- Work through the decay. Every page with a figure, a date or a rate gets checked against a current source and stamped with the date it was verified.
- Machine finds, human decides. Use AI to crawl the site and flag decaying numbers and contradictions, and use people to verify and rewrite what it flags.
That is the whole argument, and the graphic said it in six words before we did. AI amplifies the foundation you give it. The foundation is not a technology decision. It is a series of small human decisions about what you are willing to put your name to.
If you would like a second pair of eyes on what your site is currently telling AI systems, that is exactly what we do.
Frequently Asked Questions
Common questions about GEO, SEO, and AI-driven search visibility.
More often than most people assume. Research coordinated by the European Broadcasting Union and led by the BBC, published in October 2025, had professional journalists evaluate more than 3,000 responses from ChatGPT, Copilot, Gemini and Perplexity. It found that 45% of answers had at least one significant issue, 31% had serious sourcing problems such as missing or incorrect attribution, and 20% contained major accuracy issues including hallucinated details and outdated information. That study covered news, which is among the best structured and best sourced content on the web. Content about your business is unlikely to fare better.
It means recognizing that your published content is the source material AI systems use to generate answers about your business, so it should be governed like data rather than managed like marketing copy. In practice that means every factual claim has an owner, a source and a review date, that the same fact does not appear three different ways across your site, and that stale figures are actively retired. AI does not fact-check your site before repeating it, so inconsistency in your content becomes inconsistency in the answers people receive about you.
Not for general web content, but there are hard requirements in specific places. Google Merchant Center policy states that AI-generated images must carry IPTC DigitalSourceType TrainedAlgorithmicMedia metadata, and that AI-generated product data such as titles and descriptions must be specified separately and labeled as AI-generated. For general content, Google's guidance is that using generative AI to produce many pages without adding value for users may violate its scaled content abuse spam policy, and it recommends giving readers context about how content was created.
No, but unreviewed AI content is. Google's published guidance says generative AI can be useful for research and for adding structure to original content, and that the test is whether the work meets the standards of Search Essentials and the spam policies. The failure mode is not the tool, it is publishing at scale without expertise, verification or editorial responsibility. Google's guidance points explicitly at content created with little effort, little originality and little added value, which describes unreviewed output regardless of whether a human or a model typed it.
Because the models cannot verify claims about your business against a source of truth they do not have. The EBU and BBC study found sourcing failures in 31% of responses even from leading assistants in October 2025, and the Reuters Institute's Digital News Report 2026 notes that the human in the loop principle still holds across newsrooms adopting AI. A model can draft, summarize, restructure and spot inconsistencies at a speed no team can match. What it cannot do is take responsibility for whether a price, a claim or a regulatory statement is true today.
Start with the content that carries the most risk rather than the most traffic. Pull every page containing figures, dates, prices, rates or regulatory thresholds, then check each against a current source and record the check date. On large sites this is where AI genuinely earns its place, because a model can crawl a sitemap, read every article and flag decaying numbers far faster than a person can, leaving humans to verify and rewrite the flagged items. The order matters: machine finds, human decides.
Start with the facts that appear most often and cost most if wrong: pricing, product specifications, contact and location details, regulatory or compliance statements, and any headline statistic you repeat in sales material. Get one canonical version of each, published somewhere authoritative on your own site, and make every other mention agree with it. That single exercise usually resolves more AI misattribution than a year of content production, because it removes the contradictions the models are currently choosing between.
Sources
- European Broadcasting Union: AI assistants misrepresent news content 45% of the time (opens in a new tab)
- EBU and BBC: News Integrity in AI Assistants report (opens in a new tab)
- Reuters Institute: Digital News Report 2026 executive summary (opens in a new tab)
- Google Search Central: Guidance on using generative AI content (opens in a new tab)
- Google Search Central: Optimizing for generative AI search (opens in a new tab)
- Google Search Central: Creating helpful, reliable, people-first content (opens in a new tab)
- Google Search Central: Spam policies, scaled content abuse (opens in a new tab)
- Google Merchant Center: Policies for AI-generated content (opens in a new tab)
- Google Search Central: Introduction to structured data (opens in a new tab)
- Search Engine Journal: Google AI Overview citations from top-ranking pages drop sharply (opens in a new tab)
- Stanford HAI: 2026 AI Index Report, Responsible AI (opens in a new tab)