Decide separately whether you want search discovery and whether you permit training crawls. A blanket “block AI” setting may combine different decisions.
Separate search from training
An AI-related visit to your website does not always serve the same purpose. Before changing a security or crawler setting, identify the operator and what its documented agent does. The phrase “AI bot” alone does not give you enough information to make that decision.
OpenAI distinguishes OAI-SearchBot, which supports discovery in ChatGPT search, from GPTBot, which crawls material that may be used for model training. Its documentation says the two robots.txt settings are independent. Allowing search discovery does not require allowing GPTBot.
Reference: OpenAI: Search and training crawler controls
Recognize user-requested visits
OpenAI also describes ChatGPT-User as an agent for certain user-initiated actions, rather than automatic web crawling. Its documentation notes that robots.txt rules may not apply to those visits. It is not the control used to determine search eligibility.
This distinction matters when someone sends you a screenshot of an AI tool reading a page. That observation alone does not establish whether the page was crawled for search or collected for training. Ask which agent made the request before drawing a conclusion.
Reference: OpenAI: ChatGPT-User behavior
Write the decision before changing settings
For a small business, we recommend a short access policy that an owner and developer can both understand. Discuss each provider separately; do not assume one company’s bot names or controls apply to another.
- Which public pages should be discoverable through search?
- Which documented training crawlers should be permitted or declined?
- Who approves changes to crawler rules and hosting security?
- Who will recheck provider documentation when settings change?
Keep private information behind authentication
Robots.txt is a set of crawling instructions, not a password. Google warns that some crawlers may ignore it and that a blocked URL can still appear in search without its content. Sensitive information needs access controls, rather than a robots.txt entry.
Review customer files, internal documents, and staging pages separately from public marketing content. Do not paste private URLs into a public instructions file as a way of hiding them. Ask your developer to confirm how those resources are protected.
Reference: Google: What robots.txt can and cannot do
Verify access without promising visibility
After an approved change, ask your developer to review the live robots.txt file, hosting security rules, and server logs together. A rule in one place is not a complete picture of what reaches the site. Record the provider documentation and the date of the review.
Treat access as a technical permission, not a promise that an AI answer will mention or cite your business. A sensible review establishes what you allow, checks whether the controls match that choice, and leaves you with a record you can revisit.
How we can help
For a broader review of content clarity and crawler access, see our AI search optimization service.
Explore AI search optimization