Skip to content

Citelift / glossary

Glossary

AI crawler

A bot an AI company runs to fetch pages for training, for a search index, or on a user's behalf at answer time. Each vendor names its bots and says which robots.txt rules they respect.

AI companies run several bots for several jobs, and their roles should be distinguished when editing robots.txt. ChatGPT's documentation describes GPTBot for training data, OAI-SearchBot for the index behind ChatGPT search, and ChatGPT-User for pages a user asks ChatGPT to open during a chat. Anthropic documents ClaudeBot and Claude-User along the same lines. Perplexity documents PerplexityBot for its index and Perplexity-User for on-demand fetches. Google-Extended controls training and grounding in specified Gemini products; it does not affect Google Search.

Training and search controls can be separate. Blocking a search crawler can restrict its documented retrieval path; user-directed fetches and navigational links may follow different rules. Each vendor page states which token to use and whether the user-triggered bots obey disallow rules at all. Citelift's robots.txt tool writes the file from the vendors' own agent names, and this site's robots.txt allows every one of them, because being read is the point.

Sources checked 20 September 2026: ChatGPT's crawler documentation, Google crawler controls and Perplexity crawler documentation. After checking access, use a fixed AI visibility panel to record what the assistants actually return.

Citelift is listed on the Shopify App Store: Citelift on the Shopify App Store.

Run the check after reading about AI crawler