Is ChatGPT blocked from your site? Free AI crawler checker
Enter a domain. We read its robots.txt the way OpenAI, Anthropic and Perplexity crawlers do, look for Cloudflare and Shopify blocks, and check whether the homepage has anything to read without JavaScript. No login, no email gate.
See who AI recommends in your category
Hypercitations shows which brands ChatGPT, Claude, Google AI Overviews and Perplexity name for a category and which sources those answers cite, then tells you what to change.
No credit card to sign up.
Search bots and training bots are different things
Every AI vendor runs more than one crawler, and they do different jobs. A search bot builds the index that an assistant searches when someone asks it a question. A user-fetch bot loads a page live because a person in the chat just asked about it. A training bot collects pages for the next model. OpenAI runs all three: OAI-SearchBot for search, ChatGPT-User for live fetches, and GPTBot for training. Anthropic mirrors that with Claude-SearchBot, Claude-User and ClaudeBot. Perplexity has PerplexityBot and Perplexity-User and says it does not train on crawled content at all.
The verdicts at the top of the result keep these apart on purpose. "AI search & answers" covers the search and user-fetch bots, because those are the ones that decide whether an assistant can find and cite your pages. "AI training" covers GPTBot, ClaudeBot, CCBot and the two robots.txt-only tokens, Google-Extended and Applebot-Extended. You can block one group and allow the other, and many sites should.
Why testing GPTBot alone gives the wrong answer
GPTBot is OpenAI’s training crawler. Blocking it in robots.txt is a legitimate choice, and a lot of sites made it in 2023 when it was the only OpenAI bot anyone had heard of. It has no effect on whether ChatGPT can cite you today. ChatGPT search reads the web with OAI-SearchBot and fetches individual pages with ChatGPT-User, and neither of those looks at a GPTBot rule.
The reverse mistake is more expensive. A robots.txt that blocks every "AI" user agent it can find, including OAI-SearchBot, removes the site from ChatGPT search entirely while GPTBot rules get all the attention. This checker evaluates every token separately and says which rule matched, so a blanket block shows up as exactly that.
Cloudflare blocks AI bots by default
Since July 2025 new Cloudflare zones block known AI crawlers at the edge unless the owner turns it off, and the older "Block AI bots" toggle and the AI Crawl Control panel do the same on existing zones. Those blocks happen before a request reaches your server, so nothing in your robots.txt can override them. Cloudflare also offers a managed robots.txt that prepends a block of rules and a Content-Signal line to whatever file you serve.
The checker looks for both. If the homepage is served through Cloudflare you get a note to check AI Crawl Control, and if the managed robots.txt block is present it is reported on its own line with the tokens it disallows. The managed block targets training crawlers, not search bots, so a site can have it on and still be open to ChatGPT search. A Content-Signal line on its own is a statement of preference: Google, OpenAI, Anthropic and Perplexity do not document reading it.
Shopify: edit robots.txt.liquid, not robots.txt
Shopify generates robots.txt for every store from a template, and the default lets every AI crawler through. It disallows checkout, cart, account and search pages, plus filtered collection URLs, and none of that stops a bot from reading a product page. If a Shopify store shows up as blocked, someone customised the template or an app did.
To change it, open the theme code editor, add a template called robots.txt.liquid under templates, and put your rules there. The template renders the default rules with a for loop over robots.default_groups; keep that loop and add your own groups after it so the checkout and cart rules stay in place. The checker tests a real product path and a real collection path on Shopify stores, taken from the homepage links, because a rule that blocks /products/ would never show up on a test of the homepage alone.
Content that only renders with JavaScript
Search engines render JavaScript. AI search and answer bots mostly do not. OAI-SearchBot, ChatGPT-User, Claude and Perplexity fetch the raw HTML and work with what is in it. If your storefront is a single-page app that ships an empty root element and fills it in from an API call, a bot sees the empty element, and there is nothing for an answer to quote.
The readable-without-JavaScript check counts the words in the raw HTML of your homepage after removing scripts, styles and navigation. It warns only when the count is low and the page looks like a client-rendered shell: an empty root, framework bundles and not much else. A short but server-rendered page does not trigger it. The same check reports whether JSON-LD is present and which types it declares, because an assistant that finds a Product object with a name, price and availability has an easier time than one that has to guess.
Verify it yourself with curl
Everything the checker does you can do from a terminal. Fetch robots.txt, then fetch your homepage with an AI user agent and compare the status code with a normal request. A 200 for the browser and a 403 or a challenge page for the bot means something between the internet and your server is treating AI user agents differently. Real crawlers are verified by IP, so a firewall may treat the genuine bot differently from your spoofed request in either direction, which is why the live part of this checker is labelled indicative.
curl -s https://yourstore.com/robots.txt
curl -s -o /dev/null -w "%{http_code}\n" https://yourstore.com/
curl -s -o /dev/null -w "%{http_code}\n" \
-A "OAI-SearchBot/1.4; +https://openai.com/searchbot" https://yourstore.com/The crawlers this checker tests
Verified against each vendor’s crawler documentation on 28 September 2026. Token-only entries never send a request; they only tell the vendor how content fetched by another crawler may be used.
| Token | Vendor | Purpose | Obeys robots.txt | Notes |
|---|---|---|---|---|
OAI-SearchBot | OpenAI | search | yes | docs |
ChatGPT-User | OpenAI | user fetch | may not | docs |
GPTBot | OpenAI | training | yes | docs |
Claude-SearchBot | Anthropic | search | yes | docs |
Claude-User | Anthropic | user fetch | yes | docs |
ClaudeBot | Anthropic | training | yes | docs |
PerplexityBot | Perplexity | search | yes | docs |
Perplexity-User | Perplexity | user fetch | may not | docs |
Googlebot | search | yes | docs | |
Google-Extended | training (token only) | yes | docs | |
Applebot-Extended | Apple | training (token only) | yes | docs |
CCBot | Common Crawl | training | yes | docs |
Bingbot | Microsoft | search | yes | docs |
Questions
Does blocking GPTBot remove my site from ChatGPT?
No. GPTBot collects pages for training. ChatGPT search uses OAI-SearchBot to find pages and ChatGPT-User to fetch them live. Block those two and ChatGPT cannot cite you; block GPTBot alone and nothing changes in search.
What does "indicative" mean on the live test?
The checker fetches your homepage with a browser user agent, a plain bot user agent, and four AI user agents, then compares the responses. It cannot come from the vendors’ published IP ranges, so a firewall that verifies IPs may treat the real crawler differently from our request, in either direction. Treat a difference as a lead, not a verdict.
It says robots.txt is failing but I can open it in my browser.
The checker gives robots.txt five seconds and follows up to three redirects. If the file answers with a server error or times out, crawlers following RFC 9309 treat the entire site as disallowed until it loads again, which is what the result reflects. A file that returns an HTML page instead of text is reported as missing.
Do you keep my domain or the results?
Results are cached for one hour per site so repeat checks are fast, and each connection gets ten fresh checks an hour. Nothing is tied to an email address and there is no account.
Is this an SEO tool?
No. It answers one question: can the crawlers behind AI answers reach and read your pages. Whether those answers then cite you depends on what the pages say and on the other sources the assistant has, which is what Hypercitations itself is about.
Why are Gemini and AI Overviews not listed separately?
Google fetches everything with Googlebot and uses the Google-Extended token only to decide whether fetched content may train or ground Gemini. Google says Google-Extended does not affect Search inclusion. AI Overviews draw on the Search index, so they follow the Googlebot row.