Nobody outside the companies operating AI search systems knows the exact proprietary formula they use to choose every source. What we can say with confidence is that a page generally needs to be accessible, relevant to the question and useful enough to support the answer. Source quality, clarity, authority and corroboration are also sensible factors to think about — but anyone claiming to know the exact weighting is guessing.
This subject attracts confident diagrams.
You have probably seen versions like:
AI engines rank sources using 37% authority, 22% freshness, 18% citations...
Unless that comes directly from the system operator, treat it cautiously.
The major AI/search platforms do not publish their full source-selection algorithms.
We can work from what they do document.
And from ordinary information-retrieval logic.
The system has to find the page first
This is the first filter.
If the page is:
blocked;
behind login;
unavailable to the crawler;
returning an error;
not publicly accessible;
it becomes much harder or impossible to use as a web source.
OpenAI specifically tells publishers to allow OAI-SearchBot if they want content discovered and cited in ChatGPT Search.
Google says AI Overviews and AI Mode depend on pages being indexed and eligible for normal Search.
Different platform.
Same basic lesson.
Access comes first.
Relevance to the question is fundamental
A page can be authoritative and still irrelevant.
An excellent accounting website is not a useful source for plumbing advice.
Within the same subject, specificity matters too.
A page about:
"web design"
may be less useful for a question about:
"who should own the domain after a website project?"
than a focused domain-ownership guide.
That is why building distinct pages around real questions can be useful.
The distinction is real subject depth.
Not trivial keyword variation.
Clear factual content is easier to retrieve and use
If the important answer is hidden inside abstract marketing copy, the page is a poor information source.
Clear content gives systems something concrete to work with.
Definitions.
Prices.
Processes.
Locations.
Comparisons.
Dates.
Evidence.
Caveats.
That does not mean every article must become a database record.
Natural prose can still be precise.
Authority is not one metric
People sometimes talk about "authority" as though it is a single score.
In reality, credibility can come from many signals.
Who published the content?
Do they have relevant experience?
Is the site known in the subject?
Do other sources refer to it?
Is the information consistent with credible evidence?
Does the page show first-hand knowledge?
Is it maintained?
The exact signals vary by system.
The principle is simple.
A source with reasons to be trusted is more useful than an anonymous unsupported claim.
Corroboration matters in answer generation
AI systems often need to answer questions where no single page is sufficient.
They may retrieve multiple sources.
A factual claim supported across credible sources is easier to treat as reliable than a surprising claim appearing only on one unknown page.
This does not mean originality is bad.
Original data and first-hand expertise can be extremely valuable.
But if you make a factual claim that contradicts every authoritative source, expect systems — and people — to question it.
Original information can make a source more valuable
There is another side to this.
If every page repeats the same paragraph, none adds much.
Original sources can provide:
first-hand experience;
research;
data;
case studies;
local knowledge;
expert judgement;
specific examples.
Google's current AI-search guidance explicitly encourages unique, non-commodity content.
That is a strong clue about the direction of useful publishing.
Citations do not necessarily mean "highest ranking page"
AI answer systems can perform multiple searches, retrieve different sources and synthesize them.
Google says AI Mode and AI Overviews may use query fan-out across related subtopics and data sources.
ChatGPT Search can also rewrite a user's request into targeted search queries when using search providers.
So the source journey can be more complicated than:
query → rank #1 → cite #1.
That creates opportunities for specialist pages.
A smaller site may have the best page for one sub-question even if it is not the biggest domain in the overall topic.
This is why topic depth matters
Imagine a website with one generic SEO page.
Compare it with a site that clearly covers:
indexing;
Search Console;
local SEO;
sitemaps;
location pages;
AI Search;
content quality.
The second site provides more subject context.
That does not guarantee AI citations.
But it creates a much richer source base.
The pages can support each other through internal linking and consistent expertise.
Freshness matters when the fact can change
A definition of DNS does not need rewriting every month.
A guide explaining current Google AI Search requirements may.
A page describing today's platform limits may.
Source quality includes being current when currency matters.
That is why I include re-verification notes in articles that depend on platform rules.
A stale confident answer is worse than an older evergreen explanation that remains true.
Page experience still matters after the citation
Even if an AI system could technically read a terrible page, the commercial outcome still matters.
The user clicks.
What do they see?
Popups.
Ads.
Slow loading.
No author.
No contact information.
Confusing navigation.
A citation is not the end of the journey.
The website still needs to earn trust when the human arrives.
Brand and entity clarity can help the wider picture
Make it clear who the business is.
Consistent name.
Relevant About information.
Contact details.
Location.
Services.
Author attribution where appropriate.
Structured data that matches visible content.
This helps machines and people connect pages to a real organisation or person.
Do not invent entity markup as a substitute for having a credible business.
Links and mentions still have value
The web is built on references.
Relevant links and mentions can help search systems discover content and understand reputation.
That remains true even as interfaces become generative.
Do not reduce this to buying backlinks.
A genuine local mention, industry citation or useful resource reference can be more meaningful than 500 low-quality directory links.
Gavin's take
I would rather create a page worth citing than write sentences engineered to look quotable. Clear definitions and concise answers help, but the durable advantage is still having something useful behind them: evidence, experience, original detail, a strong explanation or a credible business position.
What about writing specifically to be quoted?
You can make important facts easy to understand.
That is good.
But writing every sentence as a supposedly "quotable AI chunk" can destroy the article.
The reader still matters.
I would rather write:
clear answer;
useful explanation;
specific examples;
supporting facts;
than engineer unnatural 40-word fragments in the hope a model copies them.
What I'd do
I would build the site so that a human or machine can quickly establish who the business is, what it knows, what it offers and why the information is credible. Then I would strengthen the pages that already show signs of demand rather than trying to reverse-engineer an undocumented citation formula.
No platform gives you a guaranteed citation formula
This is the point I would remember.
OpenAI says ranking in ChatGPT Search is based on multiple factors and cannot guarantee top placement.
Google says meeting requirements does not guarantee crawling, indexing or serving.
The operators themselves acknowledge uncertainty.
So an SEO provider claiming:
We know exactly how to get cited by all AI engines
should be able to produce extraordinary evidence.
Gavin’s perspective
Built by Gavin's approach
I want the site to become a source worth using.
That means:
technically accessible;
specific;
well organised;
transparent;
first-hand where possible;
careful with current facts;
and internally connected.
The aim is not to reverse-engineer an imaginary universal LLM score.
It is to make the public information unusually useful.
That strategy still works if the discovery interface changes again next year.
The bottom line
AI search engines choose sources using proprietary systems.
We cannot responsibly give exact formulas.
But strong foundations are clear:
be accessible;
be relevant;
publish clear information;
demonstrate real authority;
provide original value;
keep changing facts current;
and make the site credible when the human clicks through.
You cannot force a citation.
You can become a much better candidate for one.
Want your website to become a useful source rather than another generic marketing page?
See how I approach website content and structure or tell me the subjects your business genuinely knows well.
That is where I would build from.
---