Short answer
Choose the platform first. Check the public URL, crawler rules, and firewall behavior. Search crawlers and training crawlers can have separate controls. Access does not guarantee indexing, selection, or citation.
What you will learn
You will check whether a public page is accessible to relevant search crawlers. You will separate search access from model training access.
Ask the person who manages your site to help with server or firewall settings. This lesson does not require you to permit training.
Choose the platform first
There is no single AI crawler rule for every product. Decide where you want your public content to be available.
| Platform or feature | Access question |
|---|---|
| Google's generated Search features | Can Google index the page, and is the page eligible for snippets? |
| ChatGPT search | Can OAI-SearchBot access the content? |
| Perplexity search | Can PerplexityBot access the content? |
Access is one prerequisite. It does not guarantee that an answer will select or cite your page.
Separate search from training
OpenAI documents OAI-SearchBot for search and GPTBot for potential training use. Their robots.txt controls are independent.
Perplexity documents PerplexityBot for search. Its documentation says that crawler does not collect content for foundation model training.
User-requested page visits can use other agents. Do not assume one robots.txt rule controls every type of retrieval.
Read the current documentation before changing a rule. Google-Extended is separate from Googlebot and does not control inclusion in Google Search.
Check the public page
- Open the exact URL without signing in.
- Check that the page returns its intended content.
- Read the site's robots.txt file.
- Check rules for the relevant crawler.
- Check page-level indexing and snippet rules where applicable.
- Ask your site manager to check firewall challenges and crawler logs.
A page that opens in your browser may still block a crawler. A firewall can reject a request before robots.txt matters.
Use documented crawler verification when configuring a firewall. A matching user-agent name alone does not prove the request's identity.
Work through a search-only example
An owner wants a public service page available to ChatGPT search. The owner prefers to block GPTBot.
The site manager checks OAI-SearchBot separately. The manager also checks the firewall and the page's response.
The result is an access decision. It is not proof that ChatGPT indexed the URL or recommended the business.
Check whether the content is available
Important facts should appear in accessible page content. A login screen or browser challenge can prevent retrieval.
Do not assume every crawler runs JavaScript like a browser. If the main answer appears only after interaction, ask your developer to check retrieval.
Google can render JavaScript, but rendering and indexing have limits. Check the actual page through the platform's available tools.
Check Google Search inclusion settings
Google also documents a Search generative AI inclusion control. Check the setting available for your property before diagnosing missing AI exposure.
A Search exclusion and a training preference have different purposes. Read the current setting description before changing it.
Keep a record
Write down the platform, URL, rule, test date, and result. Save the reason for any access change.
Follow the indexing guide for Google URL checks. Then use AI visibility measurement to observe answer outcomes.