AI SEO Checklist: 25 Checks for 2026
AI SEO checklist for 2026: 25 checks across AI bot access (GPTBot, Google-Extended), entity consistency, answer-shaped content, schema and measurement.
This AI SEO checklist is a 25-point list of the checks that decide whether ChatGPT, Gemini, Perplexity and Google's AI Overviews can reach your site, trust it and cite it. It is written for a Singapore business that already has a website and wants to know, in one afternoon, what is missing. Each check names the tool and what counts as a pass, and where a fact comes from a platform's own documentation the link sits in the same sentence. We run this exact list at the start of every engagement for our AI SEO in Singapore, so nothing below is theoretical.
Quick take: Google's own guidance says there are no extra requirements to appear in AI Overviews beyond ordinary SEO. In practice three things fail most often: AI bots blocked by a firewall rule nobody remembers setting, a company name and address that differ from one profile to the next, and answers buried in images or scripts instead of text. Do groups 1 and 2 first.
How to use this AI SEO checklist
Score each check pass or fail, in order. The groups are sequenced by dependency: a system that cannot fetch your page (group 1) never evaluates your entity (group 2), and a page with no clear answer (group 3) gains nothing from perfect markup (group 4). Group 5 exists because most businesses cannot say, three months later, whether any of this worked. If the vocabulary is new, start with what GEO is and how it differs from classic SEO.
The 25 checks at a glance
| Group | Checks | Main tool | Pass looks like |
|---|---|---|---|
| 1. Crawlability and access for AI bots | 1-6 | robots.txt, CDN/WAF logs, URL Inspection | Named AI agents allowed on purpose; Googlebot never blocked; page indexed and snippet-eligible |
| 2. Entity and brand consistency | 7-11 | Organization schema, Google Business Profile, directories | One legal name, address and phone string everywhere; sameAs links resolve |
| 3. Answer-shaped content | 12-17 | Your own pages | One-sentence answer under each question heading; dated, sourced facts; comparison tables |
| 4. Structured data | 18-21 | Rich Results Test, Schema.org validator | Zero errors; markup matches visible text |
| 5. Measurement | 22-25 | Search Console, GA4, server logs | AI referrers isolated in GA4; monthly prompt panel recorded |
Group 1: Crawlability and access for AI bots (checks 1-6)
1. Decide which AI crawlers you allow, by name
OpenAI runs three agents, and blocking the wrong one is the most common self-inflicted wound on this list. Per OpenAI's crawler documentation, OAI-SearchBot decides whether you can appear in ChatGPT search answers, GPTBot collects content for training, and ChatGPT-User fetches a page when a person asks about it and may not obey robots.txt. The settings are independent: blocking GPTBot alone does not remove you from ChatGPT answers; blocking OAI-SearchBot does.
2. Never block Googlebot, and know what Google-Extended controls
Google-Extended is a control token, not a crawler. Google's crawler reference says it governs whether crawled content may train Gemini models or ground Gemini answers, and that it does not affect inclusion in Google Search or act as a ranking signal. AI Overviews are governed by the ordinary Googlebot rule: Google's AI features guide names robots.txt directives for Googlebot as the control for Search, AI features included. Disallow Google-Extended and you leave Gemini grounding; disallow Googlebot and you leave everything.
3. Check the other named bots: Anthropic and Perplexity
Anthropic's crawler help article lists ClaudeBot for training, Claude-SearchBot for search indexing, and Claude-User for fetches a person triggers. Perplexity's bot documentation lists PerplexityBot, which surfaces sites in results and is not used for training, and Perplexity-User, which generally ignores robots.txt. The pattern is the same at every AI company: a search bot you probably want, a training bot you may not, and a user-triggered fetcher robots.txt cannot reliably stop.
4. Check the CDN and firewall, not just robots.txt
On 1 July 2025 Cloudflare announced in its Content Independence Day post that it was changing the default to block AI crawlers unless they pay for content. Behind Cloudflare or any managed firewall, a bot can be blocked at the edge while robots.txt says Allow. Do not read the file; read the access log. Filter for the user agents in checks 1-3 and confirm HTTP 200. A 403 or a challenge page is a fail.
5. Confirm the page is indexed and eligible for a snippet
Google's AI features guide states that a page must be indexed and eligible to show in Search with a snippet to be a supporting link in AI Overviews or AI Mode - no additional technical requirement. Two settings quietly break that: a noindex left over from staging, and a nosnippet or max-snippet directive added years ago; Google's snippet controls page lists them, and the AI features guide names the same controls as the way to limit AI Overview text. Run URL Inspection on each money page.
6. Put important content in text, not images or scripts
Google lists making important content available in textual form as a practice that still matters for AI features. A price inside a JPEG, hours in a slider, a service list rendered only by a client-side script - none exist for a system reading HTML. Google's JavaScript SEO basics explains that Googlebot renders JavaScript in a deferred second stage; other fetchers may not render it at all. View the page source and search for the sentence a customer needs.
# Example robots.txt: citations yes, training no User-agent: OAI-SearchBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Claude-SearchBot Allow: / User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: * Allow: / Sitemap: https://www.example.sg/sitemap.xml
The block is a template, not a recommendation: allowing training crawlers is a business decision. What is not a matter of taste is the last group - never disallow Googlebot, and never leave a blanket Disallow that catches the search bots.
Group 2: Entity and brand consistency (checks 7-11)
7. One legal name, one address, one phone string - everywhere
AI systems build a picture of a company from every place it is mentioned, and they weight agreement. One name on ACRA, another on a directory and a third on LinkedIn reads as three weak entities rather than one strong one. Fix the string once - legal name, address with unit and postal code, one +65 number - and copy it character for character into every profile. Then open ten listings and count the variants.
8. Organization structured data on the homepage
Google's Organization structured data documentation lists the properties it reads, including name, url, logo, address, telephone and sameAs. Put one Organization block on the homepage carrying exactly the string from check 7. Schema that gives one phone number while the footer shows another is worse than no schema: you have published a contradiction in machine-readable form.
9. sameAs links you control, and that resolve
The sameAs property is the machine-readable version of "this LinkedIn page, this Facebook page and this Google Business Profile are all us". Every URL in it must return HTTP 200 and carry the same name and address. Dead profile links are common because the schema was written once and never revisited. Re-check quarterly.
10. A Google Business Profile that matches the website
Google's guidelines for representing your business require the profile name to reflect the real-world name and the address to be real premises. Match the primary category to what the site says you do, keep hours current, and point the website field at the canonical https URL, not a tracking link. A mismatch is visible to Gemini and AI Overviews immediately, because both draw on the same Google data.
11. Third-party corroboration you did not write
Your own site saying you exist is weak evidence; an independent source saying it is strong. Directory listings from recognised bodies, supplier pages that name you, event pages, press mentions and a Wikidata item all count. Wikidata's notability policy requires an item to be supported by serious, publicly available references - it records what is already documented elsewhere. If you do not qualify yet, earn the documented mentions first.
Group 3: Answer-shaped content (checks 12-17)
12. A one-sentence answer directly under every question heading
AI systems assemble answers from passages, not pages. The passage most likely to be lifted answers the heading's question in its first sentence, in plain words, with the number or the name in it. Rewrite each section so the first sentence could stand alone as the answer, then let the rest supply the evidence.
13. Headings phrased the way people ask
"Pricing" is a label; "How much does a website cost in Singapore?" is what a person types into ChatGPT. Google's People Also Ask boxes are a free list of your buyers' actual questions in their words. Use them verbatim as headings where they fit, and answer each under it.
14. Facts with a source and a date in the sentence
A system choosing between two pages prefers the one whose claim is attributable. "Studies show shoppers abandon slow sites" is not; "Google's threshold for Largest Contentful Paint is 2.5 seconds" with a link is. Every number should carry a source in the same sentence, or a plain statement that it is your own measurement and when you took it. If you cannot source a number, describe the mechanism instead.
15. Comparison tables for anything with more than two options
Tables are the most reliably extracted structure on the web because their meaning does not depend on the prose around them. Platform comparisons, price bands by scope, feature lists - each becomes a table with a caption that says what is being compared, short headers, and units in the header rather than every cell.
16. A named author, a role, and a visible last-updated date
Google's Article structured data guide lists author name and dateModified among the properties it reads; both should also be visible on the page. Anonymous, undated pages read as content of unknown provenance. A named specialist with a role and a fresh date reads as accountable. The byline on this page is the pattern.
17. At least one thing the model cannot get anywhere else
If every sentence on your page could be regenerated from general knowledge, there is no reason to cite you rather than answer directly. Original information is the exception: your own measurements, your own price band, a named trade-off from real projects, an honest statement of who should not buy. That is information only the vendor can supply.
Group 4: Structured data that matches the page (checks 18-21)
18. Zero errors in both validators
Run every template through Google's Rich Results Test, which reports what Google can use, and the Schema.org validator, which reports whether the markup is well-formed regardless of Google's features. A warning is acceptable; an error means the block is discarded.
19. FAQPage markup only where questions and answers are visible
Google's FAQPage documentation requires the full question and answer to be visible on the page, and restricts the rich result to well-known government and health sites. The markup will not earn a visual rich result, but it still gives AI systems a clean question-and-answer pair to extract. Use it on a genuine FAQ, never as a wrapper for marketing copy.
20. Article or BlogPosting on every editorial page, same author
Each guide should carry Article or BlogPosting markup whose headline, author and dates match what is printed. Consistency is the point: fifty posts by one named specialist with one job title build an entity; fifty posts by "admin" build nothing.
21. No markup that contradicts what a visitor can see
Google's AI features guide lists keeping structured data matched to the visible text as a continuing best practice. A rating no visitor can see, an address that differs from the footer, a price that has since changed - each is the fastest route to being treated as unreliable. Diff markup against page every time either changes.
Group 5: Measurement (checks 22-25)
22. Know what Search Console can and cannot show
Google's AI features guide states that clicks and impressions from AI Overviews and AI Mode are included in the Performance report inside the "Web" search type, with no separate filter - so Search Console shows total Google visibility, not AI-specific visibility. Track branded queries separately: a rise in searches for your company name after AI answers begin citing you is one of the few signals that survives the aggregation. Our guide to earning citations in Google's AI Overviews covers what can be attributed.
23. A GA4 channel group that isolates AI referrers
Traffic from a link inside a ChatGPT or Perplexity answer arrives in GA4 as a referral from chatgpt.com, perplexity.ai, copilot.microsoft.com or gemini.google.com. GA4's custom channel groups let you define a channel from a list of referring domains, so those sessions stop hiding inside the generic Referral bucket. Set it up before you need it; GA4 does not reprocess history for a new group.
24. A monthly read of the server log for AI user agents
The access log is the only record of which AI bots visited, which pages they fetched, and what status they got. Once a month, count hits by the user-agent strings in checks 1-3 and by response code. A bot that fetched 400 pages last month and 12 this month is telling you something changed - a firewall rule, a rate limit or a robots.txt edit.
25. A fixed prompt panel, recorded the same way every month
Pick ten prompts a real buyer would type - "best [your service] in Singapore", "how much does [your service] cost", your brand name plus "reviews" - and ask them in ChatGPT, Gemini and Perplexity on the same day each month, logged out, from a Singapore connection. Record whether you were mentioned, linked, and which competitor was. It is a small, drifting sample, but it is the only direct measurement there is, and a twelve-month series beats any single reading. Our note on how ChatGPT chooses which Singapore businesses to recommend explains why the panel should mix branded and unbranded prompts.
What this checklist does not do
An AI SEO checklist is a hygiene tool. Passing all 25 checks makes you eligible and legible; it does not make you the answer. A competitor with more independent mentions, more original information and a longer track record will be cited over you at identical technical scores, and no markup closes that gap. It does nothing for a business with a handful of pages or no proven offer - there is nothing to cite. And if you need enquiries in the next 30 days, this is the wrong list.
AI SEO checklist FAQ
01Is there an AI tool for SEO?
+
Yes, many - but none replaces the checks above. Tools that use AI to draft content, cluster keywords or summarise audits speed things up; tools that claim to "optimise for ChatGPT" with a score are guessing, because no assistant publishes a ranking algorithm. Use AI to work faster through this AI SEO checklist, and use the platform documentation linked in each check to decide what is true.
02What AI skills are needed for SEO?
+
Reading vendor documentation, writing answer-first content, and basic log analysis. The first tells you what OpenAI, Google, Anthropic and Perplexity actually do rather than what a LinkedIn post says; the second earns citations; the third finds a bot blocked at the firewall. Prompt-writing matters less than any of these three.
03Can ChatGPT do an SEO audit?
+
Partly. ChatGPT can review a page you paste in for clarity and missing answers, and it can explain a robots.txt file. It cannot crawl your site at scale, read your server logs, verify Search Console data or see how a CDN treats a bot, so it cannot complete groups 1, 4 or 5 of this list. Treat it as a second reader, not an auditor.
04What is an AI checklist?
+
An AI checklist is a fixed list of checks that determine whether AI systems can access, understand and cite a website. This one has 25 checks in five groups - access for AI bots, entity consistency, answer-shaped content, structured data and measurement - scored pass or fail, in order, because later groups depend on earlier ones.
05Should I block GPTBot?
+
Only if you have decided your content should not train OpenAI's models - and if so, allow OAI-SearchBot at the same time. OpenAI's documentation at https://developers.openai.com/api/docs/bots says the two are independent: GPTBot governs training, OAI-SearchBot governs appearance in ChatGPT search. Blocking both removes you from ChatGPT answers.
06Do I need an llms.txt file?
+
No. The llms.txt proposal at https://llmstxt.org/ is a community specification for a Markdown summary of a site for language models, and adding one is harmless. But Google's AI features guide states that you do not need new machine-readable files or AI text files to appear in AI Overviews or AI Mode, and no major assistant has documented reading one. Do the 25 checks first.
07How do I know if AI search is sending me customers?
+
Three places. GA4, with a custom channel group for chatgpt.com, perplexity.ai, copilot.microsoft.com and gemini.google.com referrals; Search Console branded-query trends, since AI Overview clicks are folded into the Web search type; and your own monthly prompt panel. Ask new enquiries how they found you as well - "I asked ChatGPT" is now a common answer in Singapore and appears in no report.
Get the 25 checks run on your site
Most Singapore businesses fail several of these checks on the first pass, and nearly every failure is cheap to fix once it is named. If you would rather have the list run for you, with log evidence and fixes in priority order, talk to us - or see how our AI search visibility work turns the checklist into a monthly programme.
