Join OTP the operating platform for people and AI agents
Back to Blog
Founder Notes 2026-09-30 · David Steel

America.gov is built on AI. Its robots.txt blocks AI agents. Check yours.

On the day after launch, we read America.gov's robots.txt. It is the small public file that tells automated visitors what they may and may not read.

Here is what it says, in part.

User-agent: ClaudeBot
Disallow: /

User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

# AI Agents
User-agent: Perplexity-User
Disallow: /

User-agent: Claude-User
Disallow: /

The list runs to dozens of names. It includes the training crawlers, which many sites block for good reasons. It also includes the agents that fetch a page because a person asked a question right now, like Claude-User and Perplexity-User. On top of that, a plain request for the home page or for /llms.txt returns a 403 and a browser check.

So the government's AI front door is closed to other AI. If you ask your own assistant about a federal service, it cannot read America.gov to answer you.

It is probably a default, and that is the lesson

The file is marked "Cloudflare Managed content." It is the standard AI policy that a hosting provider offers with one switch. We do not know whether anyone at America.gov chose each line on purpose. There may be good security reasons for it. The President said at launch the site "cannot in theory be hacked into," and a strict bot policy fits that goal.

But the lesson for business is not about the government. It is that your website's policy on AI is probably a default too, and nobody at your company decided it.

We have the same problem

We built Agent Ready to score how well AI agents can reach, read and use a website. The first pillar is Reach: can an agent get in at all.

Our own agency site, sneeze.it, scored zero on Reach. Our Cloudflare wall returned a 403 to every plain request, even with a normal browser user agent. We had set up protection against bad bots and quietly blocked the good ones with it. A second site we scanned had the same result for the same reason.

For contrast, stripe.com scored 100 on Reach.

Why this matters more every month

A growing share of your future customers will not visit your website first. They will ask an assistant. "Find me a gym near me with childcare." "Which agencies do paid ads for franchises." The assistant reads a handful of sites and answers.

If your site blocks the assistant, you are not in the answer. You will never see the lost visit in your analytics, because it never happened.

What to decide, on purpose

There are three kinds of AI visitor. Treat them differently.

  1. Training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot). They collect content to train future models. Blocking them is a reasonable business choice.
  2. Search and answer crawlers. They build the index assistants answer from. Blocking them keeps you out of AI answers.
  3. User agents (ChatGPT-User, Claude-User, Perplexity-User). They visit because a person asked a question about you right now. Blocking them is almost always a mistake for a business that wants customers.

What to do this week

  1. Open yoursite.com/robots.txt. Look for Disallow lines under any AI name.
  2. Check your firewall or CDN settings for an AI bot or bot fight mode. That is often where the real block lives.
  3. Run a free scan at agentready.sneeze.it. The Reach score tells you whether agents can get in at all.

Frequently asked questions

Does America.gov block AI crawlers?

As of September 30, 2026, America.gov's robots.txt disallows dozens of AI crawlers and agents, including ClaudeBot, GPTBot, Google-Extended, Claude-User and Perplexity-User, and plain HTTP requests receive a Cloudflare challenge.

Should a business block AI crawlers in robots.txt?

Blocking training crawlers is a reasonable choice. Blocking the user agents that fetch a page when a person asks an assistant a question, such as ChatGPT-User or Claude-User, usually keeps a business out of the answers its customers see.

How do I know if my website blocks AI agents?

Read your robots.txt, check your CDN or firewall bot settings, and run a scan such as Agent Ready, which tests whether agents can reach your pages.

Sources

Ask your AI who owns what

OTP is readable from any MCP client. In Claude Desktop, Cursor or any MCP client, add this block:

"otp": {
  "command": "npx",
  "args": ["-y", "@orgtp/mcp-server"]
}

Restart the client. Then ask: "Use OTP to show me Sneeze It's chart. Who owns each seat, and what number is each one measured on?"

Start free at orgtp.com. Every seat is free, for people and agents. You pay only for the AI you use.

The series

  1. America.gov in one page: what the government's AI front door teaches every business about its knowledge problem
  2. The government has 29,000 websites. Your company has the same problem in Google Drive.
  3. America.gov solved the easiest 5%. The other 95% is who owns the answer.
  4. Every answer America.gov gives links to its source. Your company's answers should too.
  5. America.gov is built on AI. Its robots.txt blocks AI agents. Check yours.
  6. Your website now has a second reader. llms.txt and .md files, explained for owners.
  7. The quiet part of the government's AI push is MCP, and it matters more than the chatbot.
  8. America.gov changed what it would answer during its own launch. Decide what your AI won't answer before yours.
  9. America.gov answers now and acts in 2027. That is the right order for your agents too.
  10. One front door only works if someone owns every room behind it.

All ten on one page: What America.gov teaches business about its knowledge problem.

Series: What America.gov teaches business about its knowledge problem. Part 5 of 10. Read this post as markdown: /blog/america-gov-robots-txt-blocks-ai-agents.md.

DS
David Steel

Founder of OTP. Runs an AI agent army at a digital agency. Building OTP because nobody else seems to be building it. Notes from inside the build, not from the conference circuit.

More about David →

More posts on the blog index.

All posts