Technical SEO

Is My Website Indexable? Checks for Google and AI Search

Check whether Google and AI search crawlers can read and index a page: status, noindex, robots.txt, raw HTML, canonicals and Search Console statuses.

An indexability checklist covering status, noindex, robots.txt, raw HTML, canonical and sitemap.
Illustration by UsefulShelf.

A page is indexable when a search engine can fetch it, read its content, and is allowed to keep it. Fail any one of those and the page cannot appear in results, however good it is. AI search tools add a second set of crawlers with their own rules, and several of them read only the HTML your server sends.

This guide walks through the checks that answer “is my website indexable?” for Google and for AI search crawlers, explains the JavaScript problem behind many blank pages, and shows how to read Search Console’s answer. Sources were checked on October 1, 2026.

Crawl, render, index

Google’s JavaScript SEO documentation describes three phases: crawling, rendering, and indexing. Googlebot fetches the URL, queues the page to run its JavaScript, and then decides whether to index what it found. The rendering queue “may stay on this queue for a few seconds, but it can take longer than that.”

Each phase has its own way to fail:

  • Crawl: robots.txt blocks the URL, the server returns an error, or nothing links to the page.
  • Render: the content only appears after JavaScript runs, and the crawler does not run it.
  • Index: a noindex rule, a canonical pointing elsewhere, or content Google decides not to keep.

For a quick first look, search Google for site:example.com/your-page with your own URL. A result means Google has the page indexed. No result is not proof of a problem, because the site: operator may not list every indexed URL. Use the checks below, then Search Console’s URL Inspection, for a reliable answer.

Six checks you can run yourself

All six use a browser and the page’s public URL. Run them while logged out, because crawlers never have your session.

Indexability checklist
CheckHowA problem looks like
Status codeOpen the URL in a private window, or check the Network tab.A 404, 500, login wall, or a redirect to another page.
noindexSearch the page source for noindex, and check the X-Robots-Tag response header.<meta name="robots" content="noindex"> left over from staging.
robots.txtOpen /robots.txt and find the group that matches each crawler.Disallow: / under User-agent: * or a named search crawler.
Raw HTMLView the page source (not the inspector) and look for your main text.An empty <div id="root"></div> and a script tag.
CanonicalFind <link rel="canonical"> in the source.Every page pointing to the homepage, or to a staging domain.
SitemapOpen /sitemap.xml and confirm the URL is listed.The page is missing, or the file lists redirecting URLs.

Two details catch people out. A noindex rule only works if crawlers can fetch the page, so never block the same URL in robots.txt; our robots.txt examples explain why. And a sitemap is a hint, not a request Google must follow. Our guide to XML sitemaps covers which URLs belong in one.

The JavaScript problem

Many apps built with Vite, Lovable, Bolt, or Base44 send the browser an almost empty HTML file and build the page with JavaScript. A person sees a finished page. A crawler that does not run JavaScript sees this:

<!doctype html>
<html>
  <head>
    <title>Vite + React</title>
  </head>
  <body>
    <div id="root"></div>
    <script type="module" src="/assets/index.js"></script>
  </body>
</html>

Googlebot does render JavaScript, but the page waits in a queue first, and Google’s own documentation notes that “not all bots can run JavaScript.” It recommends server-side rendering or pre-rendering because it “makes your website faster for users and crawlers.”

The fix is to send real HTML: server-side rendering, static generation at build time, or a pre-rendering step for public pages. At minimum, each public page should arrive with its own title, meta description, H1, and a few paragraphs of text explaining what the product does. Our meta tags guide covers the head.

ChatGPT search and Claude search use their own crawlers, OAI-SearchBot and Claude-SearchBot, separate from the crawlers those companies use for model training. OpenAI states that sites which opt out of OAI-SearchBot will not appear in ChatGPT search answers. Check that your robots.txt allows the search crawlers you want, even if it blocks training crawlers.

Raw HTML matters even more here. Vercel’s December 2024 analysis of crawler traffic found that OpenAI’s crawlers (OAI-SearchBot, ChatGPT-User and GPTBot), Anthropic’s ClaudeBot and PerplexityBot did not execute JavaScript. Claude-SearchBot was not in that study, so assume the same until shown otherwise: a client-rendered page that Google eventually indexes can still be blank to AI crawlers.

Being readable is a requirement, not a guarantee. No AI provider promises to cite or recommend a page because its crawler can read it.

Read Search Console’s answer

Only Google can say whether Google indexed a page. In Google Search Console, paste the URL into URL Inspection. It reports whether the URL is on Google, and Test live URL shows the HTML Googlebot rendered. The Page indexing report groups every known URL by status. Google’s report documentation defines the common ones:

Page indexing statuses
StatusMeaningWhat to do
Discovered, currently not indexedGoogle found the URL but has not crawled it yet.Link to it from pages that are already indexed, and keep the server fast and error-free.
Crawled, currently not indexedGoogle crawled the page and chose not to index it for now. It “may or may not be indexed in the future.”Make the page more useful and distinct from your other pages. Google says there is no need to resubmit it.
Excluded by ‘noindex’ tagGoogle found a noindex rule.Remove the rule if you want the page indexed.
Blocked by robots.txtrobots.txt stops Googlebot from fetching the URL.Remove the matching Disallow rule.
Alternate page with proper canonical tagThe page points to another canonical URL that is indexed.Nothing, if that canonical is the one you meant.
Duplicate, Google chose different canonical than userGoogle prefers another URL as the canonical.Check which URL Google chose, and make the two pages clearly different or point them to one canonical.

Expect gaps. Google says it does not guarantee that every page will be indexed, and advises against expecting 100% coverage. A page that passes every technical check can still wait, or be left out because it adds little beyond pages Google already has.

What our checker reports

The UsefulShelf visibility checker fetches one public page as UsefulShelfBot, without running JavaScript, and reports what a non-rendering crawler sees:

  • The HTTP status, and the number of visible words in the raw HTML. It warns below 150 words and flags an empty #root, #app, or #__app element.
  • The title, meta description, H1, and canonical link, and any noindex in the robots meta tag or X-Robots-Tag header.
  • Whether robots.txt allows Googlebot, Bingbot, OAI-SearchBot, and Claude-SearchBot, with the training crawlers GPTBot and ClaudeBot shown separately.
  • Whether /sitemap.xml exists, whether Open Graph tags are present, and a platform-specific fix for apps built on Lovable, Bolt, Replit, Next.js, or Base44.

The score is our own summary of those findings, not a Google metric. The checker cannot tell you whether a page is indexed, how it ranks, or whether an AI tool will mention it. Use it to find technical blockers, then confirm with Search Console.

Check a page now

Enter a public URL to see the raw HTML, metadata, and crawler access a search or AI crawler gets, with a fix for each problem.

Open the visibility checker →

Have a public app? A reviewed UsefulShelf listing gives it a server-rendered page with structured data, a category, and a direct website link. See how listings support discovery and the listing plans.