Your robots.txt audit can be confidently wrong
A robots.txt checker is one of those tools that can look finished after an afternoon.
Fetch the file. Match a user agent. Detect Disallow: /. Print PASS or FAIL.
That version is easy to build.
It is also easy to make wrong.
While hardening our own visibility tooling, we found three different failure classes in the same control: crawler-role confusion, incomplete protocol parsing, and incorrect handling of fetch failures.
The result changed how we design technical audits.
A crawler name is not a crawler role.
Our first mistake happened before any parsing logic ran.
We grouped crawlers by brand and assumed that a crawler associated with an AI product had one obvious purpose.
It does not.
OpenAI currently documents separate roles for:
- OAI-SearchBot, used for search;
- GPTBot, used for crawling content that may be used for model training;
- ChatGPT-User, used for certain user-triggered actions rather than automatic web crawling.
Those controls are independent.
A company might allow search discovery while opting out of training. If an audit treats the training crawler as a search crawler, it can tell that company its search visibility is broken when the configuration is actually intentional.
So the first rule became: do not infer crawler purpose from the crawler name.
Use the operator's current documentation as the source of truth.
The protocol is path-aware.
The second class of failure came from simplifying RFC 9309 too aggressively.
A naive implementation can look for a site-wide deny and then ask whether any Allow rule exists.
That breaks quickly.
Consider:
User-agent: ExampleBot
Disallow: /
Allow: /blog/
The answer is not “allowed” or “blocked”.
The answer depends on the URL path.
RFC 9309 says the most specific matching rule wins. If equally specific Allow and Disallow rules match, Allow should win.
So an audit of / and an audit of /blog/article can legitimately produce different answers from the same robots.txt file.
If your checker does not know which path it is evaluating, it cannot implement that rule correctly.
Multiple matching groups are one policy.
There is another easy parser bug: stop at the first matching user-agent group.
RFC 9309 requires matching groups for the same product token to be combined.
That means a clean-looking .find() is structurally wrong for the job.
This is a recurring engineering smell: using an API that returns one result for a domain rule that can legally have many results.
A 404 and a 503 do not mean the same thing.
Fetch handling also needs protocol semantics.
Under RFC 9309:
- a 4xx response means robots.txt is unavailable, and the crawler may access resources;
- a 5xx or network failure means robots.txt is unreachable, and the crawler must initially assume complete disallow.
If your monitoring system stores both as fetch_ok = false, it has already lost the information needed to make the correct decision.
The HTTP status is not diagnostics around the policy. It is part of the policy.
Green tests can lock in the wrong model.
The most useful failure was that our tests initially agreed with the broken implementation.
The expected output had been written from the same simplified model as the parser.
So every test passed.
That kind of test suite proves that the code matches itself.
It does not prove that the code matches the standard.
We replaced example-driven confidence with a standards-derived case matrix:
- repeated matching groups;
- root deny plus narrower allow;
- equal-specificity rules;
- wildcards and end anchors;
- multiple target paths;
- 404 and 503 behavior;
- crawler-role separation.
A test is much stronger when its expected result comes from an authority outside the implementation.
The pattern applies to more than crawlers.
This is not really a robots.txt story.
It is a story about implementing external contracts.
The same failure mode shows up in:
- OAuth flows;
- webhook signature validation;
- structured data;
- caching semantics;
- payment transitions;
- email authentication.
You can have excellent code around an incorrect model of the protocol.
And PASS still has a narrow meaning.
Even a correct robots.txt PASS does not mean “this site is visible in AI search.”
It means the crawler and path being evaluated are not excluded by that control under the rules you tested.
Actual discovery depends on retrieval, entity understanding, source quality, freshness, and whether the engine associates the company with the category in the first place.
We now keep those layers separate.
Sources
The practical rule
For any customer-facing protocol check:
- classify actors from current vendor documentation;
- derive expected behavior from the standard;
- test negative and edge cases, not only known production examples;
- state exactly what a PASS proves.
That last line matters.
A precise PASS is useful.
A broad PASS built on the wrong model is worse than no answer.