top of page
davydov consulting logo

DAVYDOV CONSULTING BLOG

HOME  >  NEWS  >  POST

AI Search Citations Can Be Manufactured: What 215,000 Fake Pages Mean for Your Business

On 2 September 2026 a research outfit called Trellner published a study that deserves more attention than it got. The team asked Perplexity for software recommendations across 380 categories and then checked where every answer came from. Nearly 60% of the AI search citations pointed to websites that barely register in global traffic rankings, and a good share traced back to three connected sites that had published over 215,000 pages written by machines, for machines.

If your customers use AI assistants to research products and services, and by now most of them do, this matters to you directly. The study shows that the sources behind AI answers can be manufactured at industrial scale, cheaply, and that at least one major AI search engine is quoting them right now.

In this post we go through what the researchers found, why the trick works so well, and what a business that wants real AI visibility should take away from it. Let's sink in!


Business professional reviewing AI search citations on a large office monitor

What the researchers actually found

The setup was plain. Trellner queried Perplexity's sonar and sonar-pro models through the API with 380 questions of the “best X software” type, 760 calls in total, asking for top-five product recommendations each time. They logged every URL the models retrieved and cross-checked each domain against the Tranco top-million ranking list and the Wayback Machine archives.

The results were not flattering. Of 7,534 citations analysed, 59.8% pointed to domains ranked worse than #100,000 in the world, and 23.4% pointed to domains that are not in the top million at all. The unranked sources were also suspiciously young. Their median first archive date was around 2020, against 2011 for ranked domains, and 16.6% of them first appeared online in 2025 or later.

The study has limits, and the authors list them openly. It covers one AI engine on a single day, and the chosen categories leaned towards niche software where strong sources are scarce. Still, the core pattern is hard to argue with: when an AI engine needs a source, it does not check reputation the way a human reader would.


Three websites, 215,128 pages, one purpose

The most striking part of the report is a small network of sites: worldmetrics.org, wifitalents.com and gitnux.org, with zipdo.co alongside them. All were registered within a few months of each other in late 2023 and 2024, share the same nameservers, use matching page templates, and run tiny blogs that link only to each other. Together they published 215,128 machine-generated “best software” guides, all under the same /best/category-software/ URL pattern. That is far more pages than there are actual software categories in existence.

These pages do not even pretend to serve human readers. Homepage titles literally say “Facts & Grounding Page”, and the meta descriptions advertise “machine-readable records”. This is content addressed to retrieval systems, not to people. Some pages still contain unrendered template placeholders and invented author bylines.

And the scheme works. Pages from the network show up as sources in Perplexity's product recommendations across hundreds of categories, and the operators monetise the position by selling research reports priced from €499 to €2,500 apiece.


Why AI search citations are so easy to game

AI search engines answer questions by retrieving pages that look relevant and grounding the reply in them. Retrieval favours pages that match the query closely, load cleanly and present data in a structured way. Domain reputation plays some role, but a much smaller one than in classic Google ranking, which had two decades to learn what spam looks like.

So if you generate a tidy, statistic-shaped page for every conceivable query in a niche, you become the best-matching source for thousands of long-tail questions almost by definition. Nobody links to you and no human ever visits you, and none of that matters, because the AI engine is the only reader you need to convince.

We have seen this movie before. Early search engines rewarded keyword stuffing and doorway pages until Google spent years building defences against them. Generative engines are at the doorway-page stage of that story right now, and this study is the clearest proof of it so far.


Automated printing line producing endless identical pages, symbolising machine-generated content farms

It also helps to remember what these engines are optimising for. Perplexity alone answers hundreds of millions of queries a month, and every answer needs sources within a couple of seconds. At that speed the system takes the best structured match it can find and moves on. Quality control of the kind a journalist or a careful buyer applies simply is not part of the pipeline yet.


The part that should worry buyers, not just marketers

There is a second angle here that has nothing to do with marketing. If your team asks an AI assistant to shortlist software, a CRM, a help desk tool, an analytics platform, some of those recommendations may already be shaped by manufactured sources. Trellner also checked the 1,502 vendor homepages the models recommended: 1.1% did not exist at all and 6.1% redirected somewhere else entirely. One recommended domain led to an Indonesian gambling site, another to a casino hotel in Monaco.

The third most cited domain overall was guideflow.com, the marketing blog of a product demo tool, which supplied sourcing for 96 of the 380 categories, most of them unrelated to its own product. A well-run content operation can quietly become the reference library for an entire industry's AI answers.

The practical rule is simple: treat any AI-generated shortlist as a starting point, not a verdict. Check vendor documentation, community discussions and established review platforms before any money moves.

One more habit worth adopting: when a stat or a comparison in an AI answer feels convenient, click through to the source. If you land on a page with no author, no date and a title addressed to machines, you have your answer about how much weight it deserves.


What this means for businesses building digital products

For companies that build and sell digital products, the study cuts both ways, and it is worth being honest about each side.

First, AI visibility is real and winnable. When 60% of citations go to tiny domains, the door is wide open in most niches. A mid-sized company with a solid content base can become a cited source for its category far more easily than it could ever rank first in classic Google search. If AI assistants recommend products in your market and you are not among the sources, someone else is, possibly a site built in an afternoon.

Second, structure now matters as much as substance. The spam sites won citations because their pages are clean, specific and machine-readable. You do not need their volume, you need their legibility: clear headings, direct answers near the top, dated statistics, comparison pages that actually compare things. When we build sites for clients we increasingly treat “will an AI engine parse and quote this page correctly” as a design requirement, right next to speed and accessibility.

Third, do not build a farm of your own. The temptation is obvious and the counter-measures are predictable. Perplexity and its competitors will respond the way Google responded to doorway pages, with index-level purges of coordinated networks. Everything built on that foundation disappears with one policy update, and it can drag your main domain's reputation down with it.


Business team verifying software recommendations against printed documents with a magnifying glass

How to earn AI citations without becoming spam

The defensible version of this playbook looks a lot like good content work with a technical twist. Publish original data you actually have: anonymised benchmarks from your projects, survey results, honest comparisons with named criteria. AI engines love citable numbers, and real ones age much better than invented ones.

Make each important page answer one question fully. Specific long-tail pages get retrieved, sprawling pages that cover everything do not. Add schema markup, date your facts and update them. Keep your brand present on independent surfaces too, review platforms, GitHub, professional communities, because engines cross-reference sources, and the same study shows Reddit and G2 still sitting near the top of the citation table.

Finally, measure it. Once a month, ask the main assistants the questions your buyers actually ask and note who gets cited. It costs almost nothing and it tells you exactly which page is worth building next.


Final notes

The Trellner study will not be the last of its kind. Somewhere between 215,128 fake buying guides and a Monaco casino recommended as a software vendor, it became clear that AI search inherited the web's oldest problem before it inherited the web's defences.

For business owners the conclusion is oddly optimistic, though. Citations are winnable, the bar in most niches is low, and honest, well-structured content wins them today. If you would like to know how AI engines currently see your website, an audit of that kind is exactly what we do at Davydov Consulting. Get in touch and we will take a look together.

 
 
 

Comments


​Thanks for reaching out. Some one will reach out to you shortly.

CONTACT US

bottom of page