Pay in rupiah

Guide · AI search and assistants

How to check AI search access to your site

Check the relevant crawler and public page. Keep search access and model training choices separate.

7 min readUpdated

Baca dalam Bahasa Indonesia

Short answer

Choose the platform first. Check the public URL, crawler rules, and firewall behavior. Search crawlers and training crawlers can have separate controls. Access does not guarantee indexing, selection, or citation.

What you will learn

You will check whether a public page is accessible to relevant search crawlers. You will separate search access from model training access.

Ask the person who manages your site to help with server or firewall settings. This lesson does not require you to permit training.

Choose the platform first

There is no single AI crawler rule for every product. Decide where you want your public content to be available.

Platform or featureAccess question
Google's generated Search featuresCan Google index the page, and is the page eligible for snippets?
ChatGPT searchCan OAI-SearchBot access the content?
Perplexity searchCan PerplexityBot access the content?

Access is one prerequisite. It does not guarantee that an answer will select or cite your page.

Separate search from training

OpenAI documents OAI-SearchBot for search and GPTBot for potential training use. Their robots.txt controls are independent.

Perplexity documents PerplexityBot for search. Its documentation says that crawler does not collect content for foundation model training.

User-requested page visits can use other agents. Do not assume one robots.txt rule controls every type of retrieval.

Read the current documentation before changing a rule. Google-Extended is separate from Googlebot and does not control inclusion in Google Search.

Check the public page

  1. Open the exact URL without signing in.
  2. Check that the page returns its intended content.
  3. Read the site's robots.txt file.
  4. Check rules for the relevant crawler.
  5. Check page-level indexing and snippet rules where applicable.
  6. Ask your site manager to check firewall challenges and crawler logs.

A page that opens in your browser may still block a crawler. A firewall can reject a request before robots.txt matters.

Use documented crawler verification when configuring a firewall. A matching user-agent name alone does not prove the request's identity.

Work through a search-only example

An owner wants a public service page available to ChatGPT search. The owner prefers to block GPTBot.

The site manager checks OAI-SearchBot separately. The manager also checks the firewall and the page's response.

The result is an access decision. It is not proof that ChatGPT indexed the URL or recommended the business.

Check whether the content is available

Important facts should appear in accessible page content. A login screen or browser challenge can prevent retrieval.

Do not assume every crawler runs JavaScript like a browser. If the main answer appears only after interaction, ask your developer to check retrieval.

Google can render JavaScript, but rendering and indexing have limits. Check the actual page through the platform's available tools.

Check Google Search inclusion settings

Google also documents a Search generative AI inclusion control. Check the setting available for your property before diagnosing missing AI exposure.

A Search exclusion and a training preference have different purposes. Read the current setting description before changing it.

Keep a record

Write down the platform, URL, rule, test date, and result. Save the reason for any access change.

Follow the indexing guide for Google URL checks. Then use AI visibility measurement to observe answer outcomes.

Sources

Review your public pages

Find page issues, then confirm crawler access with the relevant platform documentation.