Back in July, I had a question I could not stop thinking about: does a page need to rank on Google before an AI assistant can find it and cite it?

I decided to test it in the most direct way I could. I built a directory for marketing tools and published 4,953 pages on a new domain. Most of the pages targeted bottom-of-funnel questions such as alternatives to a specific tool, comparisons between tools, and recommendations for particular teams.

One quick note before I go further: I am keeping the name of the directory private because the experiment is still running. Naming it now could bring new searches, links, submissions, and attention that would change the thing I am trying to observe.

The short answer: this experiment suggests that meaningful Google visibility is not a strict requirement for AI discovery or citation. It does not show that Google is irrelevant, and it definitely does not make mass publishing safe for an established business website.

Results at a glance

SignalObserved result
Pages published4,953
Initial Google visibilityNearly 11,000 impressions per day within days of launch
Later Google visibility918 impressions on July 20, then below 100 on many days
Recorded GSC totals107,000 impressions, 63 clicks, 0.1% CTR, and average position 52.5
Typical crawler activityRoughly 4,000 to 5,000 requests per day two months after the decline
Peak captured traffic26,000 total requests in one daily window
Bing AI Performance121,900 citations in the selected three-month view
Commercial responseAbout five submission requests per week
A summary of the first-party signals observed during the experiment.

These results come from different tools and reporting windows, so they should not be read as one unified time series.

The question I wanted to answer

People often talk about being indexed by an LLM as if it works like Google's index. That language bundles several different events into one:

  • An AI or search crawler requests a page.
  • An answer engine stores, retrieves, or processes the page.
  • An AI answer mentions the business.
  • The answer includes a grounded citation to an exact URL.
  • A person clicks that citation and visits the website.

Those events can be related, but they are not interchangeable. A crawler request is not a citation. A citation is not automatically a visit. A visit is not automatically a customer.

The experiment was designed to see whether those signals could keep moving even if Google stopped giving the site meaningful search exposure.

What I built

I used Claude to research and draft the directory. The information was generally accurate, but I did not simply generate thousands of pages and walk away. I manually checked a few dozen pages, corrected the approach, and then scaled the same structure across the rest of the site.

The setup followed familiar SEO practices:

  • Clear heading structure and page templates.
  • Internal links between tools, alternatives, comparisons, and categories.
  • Content clusters around products and buyer use cases.
  • Structured data and crawlable server-rendered pages.
  • Bottom-of-funnel topics rather than broad informational articles.

The domain was new and had almost no external authority. That was useful for the experiment because it let me watch how discovery systems treated the content without an established brand or backlink profile carrying it.

Google arrived fast, then the exposure disappeared

The site went live on July 8, 2026. Within days, Google Search Console was showing close to 11,000 impressions per day. The curve looked exciting until you looked at what those impressions actually represented.

Across the period shown in Search Console, the site recorded 107,000 impressions, 63 clicks, a 0.1% click-through rate, and an average position of 52.5. This was broad page-five exposure, not a site winning valuable rankings.

Google Search Console chart showing 107,000 impressions, 63 clicks, 0.1 percent click-through rate, and an average position of 52.5
Google Search Console, July 7 to September 17, 2026. The initial test exposure fell sharply after the first two weeks.

On July 20, daily impressions dropped to 918. Over the following weeks they fell below 100 on many days. More importantly, the number of distinct queries showing the site fell from 6,212 to 136 while average position stayed broadly similar.

That changes the interpretation. Google did not simply move thousands of rankings down by a few positions. It stopped testing the site across most of the long tail. The site remained indexed and crawlable, so I do not describe this as proven deindexing or a confirmed penalty. The evidence fits an initial new-site evaluation window ending much better.

Bing went in the opposite direction

While Google exposure was collapsing, Bing's AI Performance report started moving up.

In the selected three-month view, Bing reported 121,900 citations and an average of 44 cited pages. On September 18, the report showed 4,200 citations and 121 cited pages for that day.

Bing AI Performance chart showing 121,900 total citations and an average of 44 cited pages over three months
Bing Webmaster Tools AI Performance. The labels and totals above use Bing's own definitions.

This does not represent every AI assistant, and it does not tell us that each citation created a human visit. It does show that one AI search ecosystem was increasingly citing pages from a domain that Google was barely showing.

Crawlers kept coming

The server logs told a similar discovery story. Two months after the Google decline, the directory was still receiving roughly 4,000 to 5,000 crawler requests on an average day.

In one captured daily window, Cloudflare showed 26,000 total requests, of which 23,000 were allowed. The crawler panel included 4,870 allowed requests attributed to Google, 3,040 to Microsoft, 2,390 to OpenAI, and 263 to Anthropic.

The important distinction is worth repeating: A crawler request is not a citation. It only proves that a bot requested a resource. Still, the activity shows that Google's search visibility decline did not make the site invisible to the broader crawler ecosystem.

Then real people started emailing

The most useful signal did not come from a chart. It came from the inbox.

Founders and marketers started asking to add their tools, correct existing listings, or be included in alternatives pages. The directory is currently averaging about five requests per week. Some were automated submissions, while others were clearly written by people who had found a specific page and wanted to be part of it.

Redacted inbox showing tool submissions, listing update requests, and editorial suggestions received by the directory
A redacted selection of tool submissions and update requests received during September 2026.

GA4 also began recording referral traffic from AI assistants. In the channel report, AI-assistant referrals eventually exceeded the site's Google and Bing organic traffic combined. That comparison needs context because the organic baseline had become extremely small. I am also not using raw total GA4 sessions in this study because an unrelated scraper inflated the property's overall traffic.

What this experiment actually suggests

The evidence supports a narrower conclusion than the headline might tempt us to make:

  • A page does not always need meaningful Google visibility before another system can crawl or cite it.
  • Google search exposure and AI search exposure can move in different directions.
  • Bing's AI citation data can grow even when Google query coverage contracts.
  • AI discovery can produce real commercial interest, even when traditional organic traffic is weak.
  • Crawl access, grounded citations, referral traffic, and business actions need separate measurement.

That is useful for businesses because it means AI visibility does not have to wait until every SEO target is won. You can measure and improve how AI systems discover and describe your brand while continuing to invest in traditional search.

How this compares with broader research

Broader studies point to a more nuanced version of the same answer. AirOps found that 55.8% of ChatGPT-cited pages ranked in Google's top 20 for at least one original or fan-out query, while pages ranking first were cited 3.5 times more often than pages outside the top 20. Ahrefs found much lower exact-page overlap for ChatGPT: 10% of short-tail citations matched Google's top 10, compared with 65% for Perplexity.

Rankability reached a compatible conclusion across a separate dataset: strong traditional rankings increased the chance of AI inclusion, but 55.2% of its top-10 AI citations did not rank in the traditional top 10. These studies use different methods, query sets, and AI platforms, so their percentages should not be combined. The consistent point is that ranking can improve the odds, but it is not a universal prerequisite for citation.

Our experiment adds a different kind of evidence. It follows one new domain over time and compares its shrinking Google visibility with Bing AI Performance, crawler activity, AI-assistant referrals, and submission requests. It is a longitudinal case study, not a cross-engine ranking-correlation study.

What this experiment does not prove

The direct evidence in this article comes from Bing Webmaster Tools AI Performance, Google Search Console, Cloudflare crawler logs, GA4 referral channels, and the directory inbox. It does not measure a controlled cohort of citation results across ChatGPT, Claude, Gemini, Perplexity, and Google AI Overviews. References to AI citation behavior beyond Bing are context from external research or questions being tracked next, not findings from this domain experiment.

One domain cannot settle the relationship between Google and every AI model. This study does not prove that:

  • Google rankings never influence which sources AI systems choose.
  • Every crawler request led to storage, retrieval, or citation.
  • AI-written content was the reason the pages earned citations.
  • The same result would happen in another industry or on an older domain.
  • Publishing thousands of pages is a safe or efficient growth strategy.
  • The directory's submission requests came from one identifiable discovery channel.

The experiment shows divergence. It does not prove that the channels are completely independent.

Why I would not do this on a primary business domain

I was willing to burn this domain. That was part of the test.

Google is still the largest search channel for most businesses. Publishing thousands of pages at once can weaken topical focus, create thin or repetitive content, consume crawl resources, and expose an established domain to unnecessary risk. Even accurate AI-assisted pages can fail if they do not add enough value or if the domain has no authority supporting them.

If this had been the main AI Peekaboo website, I would not have run the experiment this way. A real business should start with a small number of commercially important pages, measure exact citations, and expand only when the evidence supports it.

The questions I am tracking next

The experiment is still running. These are some of the decision-focused prompts I am monitoring without naming the directory or AI Peekaboo inside the prompt:

  • Do websites need to rank on Google before ChatGPT will cite them?
  • Can a site with almost no Google traffic still appear in Perplexity?
  • Should a new SaaS company invest in SEO or AI search optimization first?
  • Can AI visibility grow while Google impressions fall?
  • What matters more for AI citations: rankings, links, or crawl access?
  • Is publishing thousands of AI-generated pages safe for an established domain?

For each prompt, I am keeping brand mentions separate from grounded URL citations. I am also recording the model, run date, exact cited page, referral traffic, and any measurable action that follows.

Methodology and limitations

  • Launch date: July 8, 2026.
  • Pages: 4,953 AI-assisted pages on a new marketing tools directory.
  • Content: alternatives, comparisons, category pages, tool profiles, and team-specific recommendations.
  • Quality control: a few dozen pages manually checked before scaling the approach.
  • Data sources: Google Search Console, Bing Webmaster Tools AI Performance, Cloudflare crawler analytics, GA4, and the directory inbox.
  • Google interpretation: indexed and crawlable, with no evidence in the collected data of a confirmed manual penalty.
  • AI interpretation: Bing citation metrics and crawler activity are reported separately. Neither is treated as proof of a citation by every model.
  • Identity: the directory name, domain, URLs, analytics identifiers, and sender details are withheld while the test continues.
  • Main limitation: this is one uncontrolled real-world domain, not a randomized experiment.

I will update the study as the prompt cohort, grounded citations, referral traffic, and submission behavior change.

Frequently Asked Questions

Do pages need to rank on Google before AI assistants can cite them?

Not always. This experiment recorded growing Bing AI citations and continued AI crawler activity after Google visibility had fallen sharply. It does not prove that Google rankings never influence citation selection.

Does an AI crawler request mean a page was cited?

No. A crawler request shows that a bot requested a resource. A grounded citation requires an exact page URL connected to an AI answer. Crawler access, citations, and referral visits must be measured separately.

Did Google penalize the directory?

We did not find evidence of deindexing or a confirmed manual penalty. Average position stayed broadly similar while the number of distinct queries collapsed, which is more consistent with an initial new-site evaluation period ending.

Should a business publish thousands of AI-generated pages for AI visibility?

No. This was a controlled experiment on a domain we were willing to risk. Repeating it on an established business website could damage Google visibility, dilute site quality, and create a large maintenance burden.

Why is the directory name being kept private?

The experiment is still running. Publicly naming the directory could change its traffic, links, submissions, crawler behavior, and brand demand, making later observations harder to compare with the earlier baseline.