There are several ways to block access to a website for AI bots. You can define specific directives in the robots.txt file, add restrictive meta tags, or close certain parts of a page using noindex. There is also another scenario – far less obvious:
The hosting provider itself blocks access, and the site owner may not even be aware of it. This is exactly what happened with the websites of several of our clients.
And this is the key point: if we had been working only with traditional SEO, we would most likely never have discovered the issue. The sites were working properly, indexing was fine, and nothing looked suspicious from a classic SEO perspective.
The problem only became visible because we were working with GEO (Generative Engine Optimization) and analyzing brand visibility in generative AI systems, not just rankings and traffic.
What actually happened
The explanation turned out to be quite simple. For security reasons, the hosting provider had installed a protection module on its servers, designed to block suspicious bots and prevent DDoS attacks. The issue was that this module was so aggressive that it also blocked AI bots. As a result, when generative AI systems attempted to access the site, they received an HTTP 415 error, making it impossible for them to read or use the site’s content.
From the site owner’s point of view, nothing was wrong. But for AI systems, the site effectively did not exist.
What does this mean in practice?
When AI bot requests are blocked at the hosting level, generative systems cannot access the site’s content. In this situation, the site stops being a source for generative answers, and direct brand or website promotion through ChatGPT, Perplexity, Claude or other generative platforms becomes impossible.
Of course, if a user asks a direct question, the brand or website may still appear in an AI-generated answer. This happens thanks to traditional SEO. ChatGPT and similar systems do not rely only on direct website access; they also learn from and reference information available in classic search results. Part of this learning process is based on Bing’s search results.
This means that by doing standard SEO work, it is still possible to appear in AI answers – but via an indirect path. The problem is that the information used by the AI in this case can be outdated or distorted and may not reflect the current reality shown on the website itself.
The most concerning part: there are no symptoms
Yes, there are literally no symptoms:
- the website works normally,
- pages are accessible to users,
- search engine bots continue to crawl and index the content,
- traffic looks fine, rankings are stable, and all SEO metrics appear healthy.
This case is especially important because it shows no external warning signs. Google Analytics will not reveal the problem. Search Console will not either. Formally, the site may even appear in ChatGPT answers, because SEO has done its job.
At that point, the decision becomes strategic: what matters more – overly aggressive protection or a full GEO strategy, rather than indirect visibility based solely on traditional SEO. This is not a technical issue. It is a question of marketing, brand positioning, and competition in new decision-making touchpoints.
A broader takeaway
This case is not about a specific hosting provider (although we do hope the company revisits its security module). And it is not about a specific AI system. What it highlights is something more fundamental:
everything starts with accessibility.
New areas of brand visibility are emerging outside of traditional search. Infrastructure-level decisions are beginning to directly influence marketing outcomes. In this context, blocking access for generative AI systems is no longer a neutral technical choice – it is a strategic risk for the brand.