> ## Documentation Index
> Fetch the complete documentation index at: https://docs.fullreach.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Crawlability

> Whether this brand's site lets the AI crawlers in, what a plain request to its homepage gets, and whether ChatGPT can read it.

An AI platform can only cite a page that its crawler may read. The Crawlability page reads this brand's site the way a crawler would, and says what stands in the way. It says what the crawlers are allowed to fetch and what a plain request gets back. It does not say whether a crawler came: [Crawler Visits](/reference/crawler-visits) says that, from the site's own logs.

## When the site is read

FullReach AI reads the robots.txt and the homepage of this brand's site, and of each monitored competitor's site. It reads them:

* when the project is created,
* with each daily collection,
* when you press **Check again**. An organization can check again once a minute. That read skips the ChatGPT check.

The page states the day of the last read.

## The answer at the top

The page asks whether anything on the site blocks AI. "No" means that the robots.txt blocks no AI crawler and the homepage checks found no problem. Otherwise the page lists what it found, with what each finding costs.

## robots.txt

The audit reads the site's robots.txt as the crawlers' standard, RFC 9309, reads it. The most specific rule wins, and on a tie, Allow wins. Rules past the first 500 KB of the file bind no crawler, so the audit ignores them too. A site with no robots.txt allows every crawler.

Each crawler gets one verdict:

| Mark | Verdict                             | Meaning                                                            |
| ---- | ----------------------------------- | ------------------------------------------------------------------ |
| ✓    | Allowed                             | The crawler may read the site's pages                              |
| ◐    | Blocked by default, with exceptions | The crawler may read the homepage, but an ordinary page is refused |
| ✕    | Blocked                             | The crawler may not read the homepage                              |
| –    | Not read                            | The robots.txt could not be read, so there is no verdict           |

Hover over a mark to see the rule that decided it, with its line and its group. A rule for the crawler by name wins over the `*` group.

**By crawler** groups the crawlers:

* **Behind the platforms you track**: the crawlers that feed ChatGPT, Gemini, Perplexity, Copilot and Google's AI features. These decide what the tracked platforms can read.
* **Other AI crawlers**: the rest. This group is folded, unless this brand's site blocks one of them.

Within each group, the crawlers sort by purpose: search, asked by a person, training and other. [AI crawlers](/reference/crawlers) lists every crawler with its token, operator, purpose and the platforms it feeds.

When this brand's site blocks a crawler, the page says so: "You block" the crawler, what that costs, and how many tracked competitors allow it. A crawler that feeds a tracked platform also becomes a Crawler [Action](/reference/actions#the-kinds-of-action), first on the worklist.

**Test a URL** checks one path of the site against the stored robots.txt. **robots.txt** shows the stored file with its line numbers, and says when it changed. **Versions** keeps the last 30 versions.

## The homepage

The audit fetches this brand's homepage once, as a plain request with no crawler's name. A firewall refuses a forged crawler name from our address, and that would prove nothing about the real crawler.

| Check                                      | What it reads                                                                                                                                                          |
| ------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Does the homepage open to a plain request? | The HTTP status, and whether Cloudflare serves the site                                                                                                                |
| Can they read the page text?               | The visible text. Under 800 characters reads as "Almost nothing". Such a page builds its text in the browser. A crawler that runs no scripts then reads almost nothing |
| Can ChatGPT open it?                       | The ChatGPT check below. Only this brand's site gets it                                                                                                                |
| Main headline they read                    | The page's first H1                                                                                                                                                    |
| Does the page opt out of AI answers?       | The robots meta tag and header: `noindex`, `nosnippet`, `max-snippet`, `noarchive`, `nocache`. "Partly" means only some text carries `data-nosnippet`                  |

Each finding links to the operator's own page about the rule.

## Can ChatGPT open it?

Once a day, FullReach AI asks ChatGPT to read this brand's homepage. It picks a sentence from the middle of the page, never the title, the headline or the description. It gives ChatGPT the start of the sentence and asks for the rest, word for word. Each try uses a fresh address, so an old copy of the page cannot answer.

| Result           | Meaning                                              |
| ---------------- | ---------------------------------------------------- |
| Yes              | ChatGPT quoted the sentence. The page states the day |
| Not confirmed    | ChatGPT missed the sentence in 2 tries               |
| No text to check | The page holds no sentence to quote                  |

"Not confirmed" names no cause. The page can be slow, or ChatGPT can answer from a search result instead of the page. When Cloudflare serves the site, the page adds a note about Cloudflare's AI crawler controls.

## Alerts

When this brand's robots.txt starts to block an AI crawler, an Alert mail says so the same day. It compares each read with the one before, and it sends at most one mail a project a day. The homepage checks and the ChatGPT check send no mail.

## Where it appears

| Surface                       | Form                                                                    |
| ----------------------------- | ----------------------------------------------------------------------- |
| The Crawlability page         | The verdicts, the homepage checks and the stored robots.txt             |
| The sidebar                   | A badge that counts the blocked crawlers and the homepage problems      |
| A banner on the dashboard     | "robots.txt blocks N AI crawlers"                                       |
| [Actions](/reference/actions) | A Crawler Action for each blocked crawler that feeds a tracked platform |
| The weekly Digest             | A standing issue while a crawler is blocked                             |
| An Alert mail                 | The day a crawler becomes blocked                                       |
| MCP, `get_crawlability`       | The verdicts, their reasons and the homepage findings                   |

Every tier has the crawler audit.
