How to Get Cited by AI Search Engines: A Step-by-Step Process
A nine-step process for earning citations in Google AI Overviews, ChatGPT search and Perplexity: access, questions, baseline, content, corroboration and measurement.
- Read time
- 17 min read
- Sections
- 20
- FAQs answered
- 15
- Topic
- AI Search
To get cited by AI search engines, make sure they can reach your pages, find the questions people ask about your topic, write clear sourced answers to them, earn mentions from other credible sites, and then test your results on a fixed set of questions every month. This guide is that process, step by step.
It is deliberately a process, not a theory piece. If you want the background on what GEO is and why it works, read the generative engine optimization guide first. If you want the answer-box side, read the AEO guide. Here we assume you have decided to try, and we walk through the work in order, with the checks that tell you whether each step is done.
Before you start: set honest expectations
No one can guarantee a citation. The engines do not publish how they choose sources, their behaviour changes, and the same question can return different sources on different days. What you can do is remove obstacles, improve the quality and clarity of your pages, and measure over time. Work on the assumption that you are improving your odds, not buying a result.
Also decide what a win means for your business. For some firms it is a citation in a buying-stage answer. For others it is a brand mention, or a visit that turns into an enquiry. Write it down now. It will shape which questions you choose in step two.
Step 1: Check that the engines can reach and use your pages
This step has the best return on time because the problems are binary. Either the page can be fetched and shown or it cannot.
Google AI Overviews and AI Mode
Full detail is in our Google AI Overviews guide.
Google says a page must be indexed and eligible to be shown in Search with a snippet to appear in its generative AI features, and that your site must be included in generative AI features in Search Console. Its guide, Optimizing your website for generative AI features on Google Search, adds that this work is still SEO. In Search Console, check that your site has not been excluded from these features. Confirm each important page is indexed using the URL Inspection tool in Search Console, returns HTTP 200, is not set to noindex, and is not limited by nosnippet or a very low max-snippet. Google states that more restrictive preview settings limit how your content can feature in its AI experiences.
ChatGPT search
Full detail is in our ChatGPT search guide.
OpenAI's documentation says OAI-SearchBot is used to surface websites in ChatGPT's search features and that sites opted out of it will not be shown in ChatGPT search answers, though they may still appear as navigational links. It recommends allowing the bot in robots.txt and allowing requests from its published IP ranges. OpenAI also says it can take about 24 hours after a robots.txt update for its systems to adjust.
Perplexity
Full detail is in our Perplexity guide.
Perplexity documents PerplexityBot as designed to surface and link websites in Perplexity search results and says it is not used to crawl content for AI foundation models. It recommends allowing the bot in robots.txt and permitting its published IP ranges. It also notes that Perplexity-User, which fetches pages when a person asks a question, generally ignores robots.txt rules because a user requested the fetch.
A robots.txt example
This example allows search-focused bots while still letting you decide separately about training crawlers. Adjust it to your own policy and test it before publishing.
User-agent: Googlebot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
# Training crawler: your choice, independent of search visibility
User-agent: GPTBot
Disallow: /
Then check the layers robots.txt does not cover. A CDN or web application firewall can block these crawlers with a bot-protection rule even when robots.txt allows them. Check your firewall logs for blocked requests from the user agents above. Sources: Google, OpenAI and Perplexity.
Done when: every key page is indexed, returns 200, shows a snippet, and your logs show the search bots fetching it successfully.
Step 2: Choose the questions you want to be cited for
You cannot be cited for everything. Choose 20 to 50 questions that match your business goal.
- Brainstorm from customers. Pull the real questions from sales calls, support tickets and reviews.
- Cover the buying journey. Include early questions ("what is..."), middle questions ("how does... compare to...") and late questions ("how much does... cost", "is... worth it for..."). Late-stage questions often lead to enquiries.
- Add your brand and category. Questions like "best [your category] in [your city]" and "[your brand] reviews" show how engines describe you.
- Write them as people speak. AI questions are longer and more conversational than classic keywords.
- Tag each one with the page on your site that should own the answer. If no page exists, that is a content gap. If two pages could own it, merge or differentiate them so they do not compete.
Done when: you have a spreadsheet of questions, each tagged with an owning URL or flagged as a gap.
Step 3: Run a baseline on each engine
Before you change anything, record where you stand. For each question in a fresh session, ask it in Google (check for AI Overviews and AI Mode), ChatGPT with search turned on, Perplexity, and Microsoft Copilot if your market uses it. Use the same wording each time and log the results in a simple table.
| Column | What to record |
|---|---|
| Question and date | Exact wording, the day you ran it |
| Engine and mode | For example Google AI Overview, ChatGPT with search, Perplexity |
| Brand mentioned? | Yes or no, and how it was described |
| Your URL cited? | Yes or no, and which page |
| Who else is cited | The domains that appear, in order |
| Source types | Competitor sites, publishers, forums, directories, video |
| Notes | Anything inaccurate said about you |
Answers vary from run to run, so do not draw conclusions from one result. Treat the baseline as a first reading, and repeat it monthly. Where you can, run each question twice and note differences.
Done when: you can say, for your top questions, who is being cited today and what kinds of pages they are.
Step 4: Study the sources that win
For the ten questions that matter most, open the pages being cited and read them as an editor would. You are looking for patterns, not for something to copy.
- What type of page is it? A guide, a comparison, a tool page, a forum thread, a news piece?
- How does it open? Does it answer in the first lines?
- What does it include that you do not? A table, original numbers, a worked example, a named expert?
- How current is it? Dates, versions and prices.
- Who else mentions it? Is it widely referenced?
Write down what the winning pages are missing as well. A gap is your opening: a question they answer vaguely, a case they ignore, a number that is out of date. Original information beats a rewrite of what already exists. Google itself says it wants unique, non-commodity content that satisfies the person.
Done when: every priority question has a note saying what you will add that the current sources do not.
Step 5: Build or upgrade the owning page
Now write. Work through this list for each page.
- State the answer first. The opening of each section gives the direct answer in a few sentences, then supports it.
- Make each passage self-contained. Define terms where they first appear and avoid vague pronouns.
- Add attributed specifics. Figures, dates and names, each with its source linked in the sentence. Remove unsourced numbers.
- Use the right shape. A table for comparisons, numbered steps for processes, short lists for options.
- Show real experience. Include examples from work you have done, named only if the client agrees, and be clear about limits.
- Name the author and the business. Give readers and systems something to trust. Use a real person's name only if they wrote or reviewed it.
- Date it honestly. Show the publish date, and show an updated date only when you made a real update.
- Keep entity names consistent. Use the same product, brand and place names throughout.
The GEO research paper, Aggarwal et al., arXiv 2311.09735, found in its benchmark that adding citations, relevant quotations and statistics tended to raise visibility while keyword stuffing did not. That is directional evidence, not a rule for any live engine, but it points the same way as good editing.
Done when: each owning page reads cleanly with every number sourced and every section opening with an answer.
Step 6: Support the page with structure and markup
Structure helps people and systems read the page.
- One
h1, logicalh2andh3headings, and anchor ids so a section can be linked directly. - Real HTML tables and lists.
- Descriptive internal links between related pages, so the topic is clearly connected on your site.
- Structured data only where it matches what is visible. Google says there is no special markup needed for AI Overviews or AI Mode, and that markup must match the visible text.
If you want help choosing which markup is worth adding, our schema services page explains our approach, and our article on FAQ rich results covers one change worth knowing.
Step 7: Build corroboration beyond your own site
When an engine has to name a recommendation, it tends to prefer things that several independent sources describe consistently. That is an inference from observed behaviour, not a published rule, but it matches how credibility works in general. Work on the sources the engines already cite for your questions.
- Industry publications and associations: contribute genuine expertise, comment on news or supply data.
- Directories and profiles: make sure listings are complete, accurate and consistent in name, address and description.
- Reviews: ask real customers on the platforms your buyers use. Never buy or fake reviews.
- Original research: publish something others can cite, with your method clear.
- Video and community: if the baseline shows video or forum sources cited, a clear, helpful presence there can matter.
For link-focused work, see our link building services, which are built around earned, relevant placements.
Step 8: Fix what the engines get wrong about you
Your baseline may show wrong prices, old addresses or claims you never made. Engines learn from the web, so the fix is to correct the sources. Update your own pages first, then the directories and profiles that carry the wrong detail, and publish a clear page for the fact in question, for example a current pricing or services page. Re-check in a month. Some engines also offer feedback options. Use them for serious errors.
Step 9: Measure monthly and decide what to change
Re-run the same question set each month in the same way. Add three more views to your dashboard.
- Search Console: open the Generative AI performance report for impressions in AI Overviews and AI Mode by page, country and device. It has no click data, so pair it with the standard performance report for clicks on the same pages.
- Analytics: sessions whose referrer is an AI product, with their conversion rate.
- Server logs: fetches by OAI-SearchBot, PerplexityBot and Googlebot on your key pages, with response codes.
Read results as patterns across many questions, not as single wins or losses.
A worked example of reading a baseline
This is an invented example to show how to interpret results, not data from a real client. Imagine a Manchester commercial cleaning company tracks 30 questions across three engines. After the first run it finds the following pattern.
| Question group | What the baseline showed | Interpretation | Action |
|---|---|---|---|
| "What is included in office cleaning?" | Competitors and a trade association cited. The company is absent. | The company has no clear page for this question | Create one owning page with a defined checklist and sourced standards |
| "How much does office cleaning cost in Manchester?" | Directories cited. A rival publishes ranges. | Price transparency is a gap | Publish honest ranges and what changes them, dated |
| "[Company name] reviews" | One engine shows an outdated address | Inconsistent profile data | Correct the listings and the website footer |
| "Best commercial cleaners in Manchester" | A local business list is cited, and the company is missing | Third-party corroboration is low | Earn legitimate inclusion and gather reviews |
Four patterns, four different fixes. That is the point of a baseline: it tells you which of the steps above needs the most work instead of making you guess.
A page brief template
Give each owning page a one-page brief before writing. It keeps writers, editors and specialists aligned.
| Field | What to fill in |
|---|---|
| Questions owned | The main question plus three to six close variants this page will answer |
| Reader | Who is asking and what they need to decide |
| Direct answer | Two or three sentences, drafted first |
| Required facts | Each fact, its source, its date and the link |
| Original contribution | What this page adds that cited pages do not |
| Format | Table, steps or list, based on the question shape |
| Internal links | Which related pages link in and out |
| Owner and reviewer | Who writes it and who checks the facts |
| Review date | When it will be checked again |
In-house or agency?
You can run the process yourself if you have a technical person for the access work and an editor for the content. Many firms find the baseline and monthly panel easy to do in-house, while content and off-site work benefit from outside help. If you hire an agency, ask for the baseline, the question list, the page briefs and a monthly report. If they cannot show their method, do not pay for a secret one. For a plain look at how we work, see our AI SEO services page.
Diagnosing why you are not being cited
| Symptom | Likely cause | What to try |
|---|---|---|
| Page not in Google at all | Not indexed, blocked or noindexed | Fix indexing, request recrawl |
| Cited by Google but not ChatGPT | OAI-SearchBot blocked or not yet crawled | Check robots.txt, firewall and logs, then wait for recrawl |
| Competitors cited, you are not | Their page answers more directly or more specifically | Rewrite answer-first, add sourced specifics and original data |
| You are mentioned but not linked | Engine used your brand name from other sources | Strengthen the owning page and its clarity |
| Wrong facts about you | Outdated or inconsistent sources | Correct your own pages and third-party profiles |
| Cited, but no visits | Answer satisfied the query in full | Offer something more to click for: tools, templates, data |
| Results change every run | Normal variation | Judge by patterns across many questions |
Special cases: new sites, large sites and script-heavy sites
New sites
A new site has little authority and few mentions, so it is rarely the first source an engine reaches for on broad questions. Pick narrow, specific questions where you can be the clearest answer, build real mentions slowly and make sure everything technical is clean from day one.
Large sites
On big sites, the usual problem is duplication: many pages answering near-identical questions. Audit for overlap, merge pages that compete, and keep one strong owning page per question with clear internal links from related pages.
Sites that depend on JavaScript
Google documents that it can process JavaScript, but you should not assume every crawler does the same. Check what your important pages contain in the raw HTML response, without scripts running. If the main text only appears after scripts load, test each engine you care about, and consider server-side rendering for key content. This is a sensible precaution rather than a confirmed requirement for each bot, so verify against each provider's documentation.
Multi-country sites
If you serve both the UK and India, or any two markets, keep the pages for each market distinct and accurate: local pricing and currency, local regulations, local contact details. Do not duplicate one page across markets with only the currency changed.
Common mistakes in this process
- Starting at step five, with new wording, when step one has not been checked.
- Changing too many things at once, so you cannot tell what worked.
- Treating one engine's result as typical.
- Publishing many similar pages for slightly different questions instead of one strong page.
- Buying mentions in bulk from low-quality sites.
- Leaving the prompt panel unrun, so improvement is a feeling instead of a record.
When to stop and reassess
Not every question is worth chasing. After two or three months, review the panel and ask honestly where the effort is paying off.
- Keep going where you have moved from absent to cited, or where wrong facts have been corrected.
- Change approach where competitors with far more authority dominate and you have nothing original to add. Choose a narrower question instead.
- Drop the question if it does not lead to enquiries, even when you could win it. Visibility that never reaches a buyer is not worth the work.
- Escalate a fix where a technical block remains, because that stops everything else.
Treat the programme as a series of small, measured experiments, and keep notes on what each one taught you. That record is worth more than any single citation.
A 90-day schedule
| Days | Focus | Output |
|---|---|---|
| 1 to 14 | Access checks and fixes, question list, baseline | Fixed access, spreadsheet of questions and first results |
| 15 to 45 | Source study, then upgrade the top five owning pages | Five pages rewritten with sourced specifics |
| 46 to 75 | Corroboration work, inaccuracy fixes, internal linking | New mentions and corrected profiles |
| 76 to 90 | Re-run the panel, compare, decide next sprint | A short report of what changed and why |
Where we can help
We are a digital marketing agency in Manchester, UK and Mumbai, India. Our AI SEO services follow this same process, including the access audit and the monthly prompt panel, and for ChatGPT-specific work see our ChatGPT SEO services. If you would like us to run the baseline for you, contact us and we will share what we find, including where you are already doing well.
Your questions, answered in plain English
Make sure the engines can reach your pages, choose the questions you want to be cited for, write clear sourced answers on one owning page per question, earn credible mentions elsewhere, and test monthly on a fixed question set. Citation is never guaranteed.
Allow OAI-SearchBot in robots.txt and your firewall, keep pages fast and indexable, and publish clear, specific answers. OpenAI documents that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers.
Allow PerplexityBot and its published IP ranges, make sure pages are accessible and clearly written, and give direct, sourced answers. Perplexity documents PerplexityBot as the crawler that surfaces and links sites in its results.
Your page must be indexed and eligible to show with a snippet, and your site must be included in generative AI features in Search Console. Google says no special markup is needed and that this work is still SEO. Focus on helpful, original content and standard SEO best practices.
For search visibility, allow Googlebot, OAI-SearchBot and PerplexityBot. Training crawlers such as GPTBot are a separate decision, and OpenAI states the settings are independent. Also check CDN and firewall rules.
No. Firewalls, CDNs and bot-protection tools can block crawlers even when robots.txt allows them. Check server and firewall logs for blocked requests from the search bots you want.
A set of 20 to 50 questions is enough to see patterns without becoming unmanageable. Include early, middle and late buying-stage questions and a few brand and category queries.
These systems generate answers dynamically, and retrieval, wording and sources can vary between runs and over time. Judge results by patterns across many questions and repeated runs, not a single result.
Clear definitions, comparison tables, step-by-step guides, original data and accurate pricing or service pages fit the questions people ask. Thin, generic or unsourced pages rarely do. This is practitioner experience, not a published rule.
Authority and independent mentions appear to help engines trust a source, but no engine publishes how it weighs them. Earn relevant, real mentions and treat them as supporting evidence, not a guaranteed lever.
Not as a GEO requirement. Google says no special markup is needed for AI Overviews or AI Mode. Use structured data where it matches visible content and serves a clear purpose.
Access fixes can show effects within days to a few weeks, depending on crawl frequency. OpenAI says robots.txt changes can take about 24 hours to be processed for search. Content and authority improvements usually take months.
Correct your own pages first, then update third-party profiles and directories, and publish a clear page for the fact. Re-check after a month, and use the engine's feedback options for serious errors.
Some tools sample AI answers at scale, but coverage and accuracy vary. A manual prompt panel, Search Console, AI referral traffic and server logs together give a reliable picture without relying on one tool.
It can be if you lack time or technical access, but be cautious of guarantees and of any provider who cannot explain their method. Ask for a baseline, a clear process and reporting on enquiries or revenue.
Still curious? Send us your question and a strategist will get back to you.
Found this useful?
Get recommended by AI search
We check whether Google, ChatGPT and Perplexity can reach your pages, then fix what blocks them. No guarantees, just a clear plan.