How Google's AI Decides Who Shows Up (and What We Did to Show Up)
How does Google's AI choose which sites to cite?
Not by traditional ranking: industry analyses show only about a third of sources cited in AI Overviews rank in the top ten organic results. The AI must be able to read the page, find extractable structure (questions, direct answers, lists, FAQ in the HTML), trust signals (author, credentials, sources) and a passage that literally answers the question.
For twenty years there was one rule: be on Google's first page. AI broke that rule without warning. Today, when someone asks a question and Google answers with an AI-generated summary, the sources that appear there are chosen by a different logic. Industry analyses show that only a third of the pages cited in AI Overviews rank in the top ten organic results, and almost half are not even in the top fifty. In other words: you can be cited without ranking, and rank without being cited. This week Google took one more step in that direction, and we spent the week auditing our own blog to understand the game from the inside. This article is what we learned.
What changed this week
Google launched a "preferred sources" button that any site can embed on its own pages. The reader clicks, and from then on Google shows that site more often in news, AI Overviews and AI Mode. More than 600,000 sources have already been marked by users since the feature appeared inside search. Read what this means: Google is admitting that AI source selection needs a direct human signal, because traditional ranking is not doing the job.
What Google says matters (and what it does not say)
Google's official guidance for appearing in AI features is deliberately vague: people-first content, well structured, with demonstrable experience and authority (the famous E-E-A-T), and no special markup. There is no "AI Overview schema" or magic tag. What does exist, watching what AI actually cites, are four doors your page must pass through in sequence:
- Access. The AI must be able to read the page. It sounds obvious, and it is where most sites fail without knowing.
- Structure. Question-shaped headings, direct answers in the first paragraphs, lists, tables, FAQs visible in the HTML. AI extracts passages, it does not read articles; a passage that does not stand on its own is not cited.
- Trust. Who wrote it, with what credentials, citing which sources. AI discounts claims without an author and numbers without an origin.
- Answer. The user's question must have a literal, complete and short answer somewhere on the page. If the answer is diluted across three paragraphs, the AI finds another page where it sits in one.

What we found auditing our own blog
We ran a full GEO audit (Generative Engine Optimization) on the Data Lover blog. The initial score was 51 out of 100, and the main reason was embarrassingly simple: our robots.txt explicitly invited the ChatGPT, Claude and Perplexity crawlers to read the site, and a default rule in Cloudflare, our protection layer, was blocking exactly those crawlers with a 403 error. All the content work was invisible to half the AIs. One click in the dashboard fixed it. The first door was shut and we did not know.
The rest of the list is what we recommend to any company, because we got almost all of it wrong before fixing it:
- Hidden FAQs. Our FAQ showed the first answer and the other four only appeared on click. For an AI reading the HTML, 80% of the FAQ did not exist. Now every answer goes into the code, collapsed only visually.
- Numbers without sources. Articles on salaries and energy use said "surveys show" without linking any survey. We replaced that with named, linked sources: Robert Half, State of Data, Epoch AI, the journal Science.
- Author without an entity. The byline was loose text. We created an author page with credentials, photo and profiles, and connected it all with structured data (Person, BlogPosting, FAQPage).
- Quick answer at the top. Every article now opens with a question and a 40-to-60-word answer, the exact format AI summaries clip.
- llms.txt. A text file at the site root (an open standard) that introduces the site and its main pages to AI models, the way robots.txt does for search engines.
Checklist for your company
- Test access today: ask ChatGPT or Perplexity to summarize a page of your site. If it cannot, nothing else matters.
- One question, one answer: for each important page, write the question a customer would ask and answer it in one paragraph at the top.
- Show who is speaking: author with name, role and their own page; company with address, phone and consistent profiles.
- Cite what you claim: every number with a linked source.
- Data only you have: your own research, real cases, actual prices. It is the only content AI cannot find elsewhere and must cite you to use.
- Enable the preferred sources button and ask your customers to click it: it is a direct human signal to Google.
Traditional search rewarded those who ranked. AI search rewards those who answer, with proof and with a name. It is more work, but a more honest game: this time, the small page with the right answer beats the big page with the vague one. We started playing this week; we will report the score in a month.
Frequently asked questions
Not by traditional ranking: industry analyses show only about a third of sources cited in AI Overviews rank in the top ten organic results. The AI must be able to read the page, find extractable structure (questions, direct answers, lists, FAQ in the HTML), trust signals (author, credentials, sources) and a passage that literally answers the question.
A button sites can embed on their own pages; when a reader clicks it, Google shows that site more often in news, AI Overviews and AI Mode. More than 600,000 sources have already been marked by users since the feature reached search.
No. Google states there is no schema or tag specific to AI features. Structured data (Article, FAQPage, Person) helps AI understand the page, but what matters is people-first, well-structured content with demonstrable experience and authority.
Optimizing content to be read, understood and cited by generative engines such as ChatGPT, Claude, Perplexity and Google's AI Overviews. It includes allowing AI crawlers, structuring direct answers, exposing FAQs in the HTML, citing sources, having an identifiable author and publishing an llms.txt file.
Blocking AI crawlers without knowing: default protection rules (such as Cloudflare's) can return a 403 error to the ChatGPT, Claude and Perplexity crawlers even when robots.txt allows them. That is exactly what we found auditing our blog, and one click in the dashboard fixed it.

Data and AI executive with 20+ years building technology that moves businesses. Microsoft Certified Trainer, with executive education at MIT Sloan. At Data Lover, he trains professionals and leads enterprise AI projects.
See profile and all articles →

