Report · scanned 16 August 2026

6 of 274 SEO agency sites pass all nine AI-readability checks.

We ran the nine technical checks an AI assistant needs to read a page on 274 European agencies that sell SEO or GEO services. Most are mostly readable: the median score is 83 out of 100. Very few are fully readable, and the most common gap is the cheapest to close.

By . Method and every figure below; no agency is named.

6
pass all nine checks
61%
fail two or more
45%
have no machine-readable copy
13%
are not a resolvable entity

Pass rate per check

The nine checks, the weight each carries in the score, and how many of the 274 sites passed. The last row is emerging and optional; it is reported for the trend and does not count toward the score.

robots.txt reachable
medium weight
97% 267 of 274

Crawlers get no guidance at all and some will not proceed.

AI crawlers allowed
high weight
95% 260 of 274

GPTBot, ClaudeBot or PerplexityBot is told to leave. The site cannot be cited by an assistant that never read it.

Sitemap declared
medium weight
96% 262 of 274

Pages are discovered by luck rather than by list.

Structured data on the homepage
high weight
94% 257 of 274

The page is prose to a machine, with nothing it can extract as a fact.

Organization schema
high weight
87% 239 of 274

The company is a document, not an entity. Assistants recommend entities.

Machine-readable copy (llms.txt or markdown)
medium weight
55% 150 of 274

An assistant has to parse the design to find the words.

Title and meta description
low weight
95% 259 of 274

The two lines every index and every answer engine reads first are missing.

Single h1
low weight
81% 223 of 274

The page has no single stated subject.

Content-Signal line in robots.txt
low weight
5% 15 of 274

Emerging and optional. Included for the trend, not the score.

Three findings

1. Almost nobody fails the basics. Almost nobody clears the bar.

Robots, sitemap, title and structured data each pass on more than 93% of sites. Those are the checks a good CMS does for you. The checks a person has to decide to do are where sites drop: 166 of 274 fail at least two, 64 fail at least three, and 6 pass all nine.

2. The most common gap is a text file.

Machine-readable copy (llms.txt or markdown) is the weakest real check: 124 sites, 45%, have neither an /llms.txt nor a markdown copy of the page. It is also the most common single biggest gap where the scan named one (95 sites). It takes an afternoon.

3. One site in eight is not an entity.

35 sites have structured data but no Organization schema, so an assistant can extract facts from the page and still has nothing to attach them to. Assistants recommend entities. A site that is a document, however well written, is not in the running.

Checks passed, out of nine

No site passed fewer than three.

3
1
4
6
5
21
6
36
7
102
8
102
9
6

Site language

The list skews Italian because the search did.

Italian 151
English 36
French 18
German 16
Dutch 13
Norwegian 9
Polish 7
Danish 7
Other 17

Method

  • Sample. 274 agencies in Europe that advertise SEO or GEO services, 160 in Italy and 114 elsewhere, narrowed from a first-party list of 1,758 reachable agency sites in 12 languages to those that market GEO and whose sites were live in August 2026. 2 further sites were excluded because the scan did not complete.
  • Fetch. Four public URLs per domain: the homepage, /robots.txt, /sitemap.xml and /llms.txt. No login, no crawl beyond the homepage, a 25-second deadline per site.
  • Checks. The nine listed above, each a pass or fail. They are the same checks the free AI-readiness check runs, and the code is in the open-source repository under lib/audit/agent-readiness.ts.
  • Score. Weighted share of checks passed: high for crawler access, structured data and Organization schema; medium for robots, sitemap and machine-readable copy; low for title, h1 and content signals. Content signals are excluded from the score.
  • Date. All sites scanned on 16 August 2026. Aggregates computed from the result file on 7 September 2026.

Questions about the data

Which sites were scanned?

274 agencies in Europe that advertise SEO or GEO services, 160 in Italy and 114 elsewhere, narrowed from a first-party list of 1,758 reachable agency sites in 12 languages to those that market GEO and whose sites were live. Two more were scanned and excluded because the scan did not complete. Agencies were chosen because they sell readability to their clients, so their own sites are the fairest test. No agency is named here.

What does the scan read?

Four public URLs per domain: the homepage, /robots.txt, /sitemap.xml and /llms.txt. It never goes deeper than the homepage and never logs in. The same nine checks run on the free check page on this site, and the code that runs them is in the open-source repository.

Why does the homepage say six pass, when the median score is 83?

Because the two measure different things. The score weights the checks, so a site that passes everything except the low-weight ones scores in the eighties. Passing all nine is the stricter bar, and six sites clear it. Most sites are mostly readable; very few are fully readable.

Can I get the raw data?

Not the per-site rows: they name the agencies and were gathered for a scan, not for publication. The aggregate counts on this page are the complete set of figures we quote anywhere, and the method is reproducible on any list of domains with the free check.

Will this be repeated?

Yes, on the same list, so the pass rates can be compared. The date on this page is the date of the scan it reports.

Be the site the assistant names.

Add a domain and the first draft is written while you watch. Every publish is your decision, a click on the draft or a rule you set and can hold, and you can read the source or self-host it free.

Add a domain, it sets up your workspace