Methodology, limitations & disclaimer

We believe a free tool earns trust by being transparent about what it can and cannot see.

What we check

SEO fundamentals (custom engine). We fetch your page's raw HTML exactly like a crawler (no JavaScript execution) and check: title and meta description presence and length, H1 and heading hierarchy, canonical URL, indexability (meta robots), robots.txt, XML sitemap, Open Graph and Twitter card tags, schema.org JSON-LD structured data, image alt-text coverage, language declaration, mobile viewport, HTTPS and mixed content. The SEO score is a weighted average of these checks; checks that could not run are excluded from the score and flagged, never guessed.

Performance (Google data). Performance, Core Web Vitals and the Lighthouse SEO / Accessibility / Best-Practices scores come directly from the Google PageSpeed Insights API, the same engine behind Chrome DevTools. Where Google has real-user field data (CrUX) we show it and label it as such; otherwise we show lab data and label it as lab. Lighthouse results naturally vary a few points between runs.

Accessibility. Two layers: our static checks (lang attribute, alt attributes, form labels, empty links/buttons, landmarks, skip link — heuristics on raw HTML) plus the Lighthouse accessibility audit (including real color-contrast measurement in a rendered browser). The form-label check skips controls that are both aria-hidden="true" and tabindex="-1": those are not in the accessibility tree, so a missing label on them is not a barrier — a spam honeypot is the usual case. A control that is aria-hidden but still reachable by keyboard is not skipped, because that is a real defect rather than an exemption.

AI-Readiness (0–100). Our own composite score measuring how much of your site an AI assistant can actually read. Points: content readable without JavaScript (30), robots.txt access for GPTBot / ClaudeBot / Google-Extended / PerplexityBot / CCBot (15), structured data (15), llms.txt (10), semantic HTML structure (10), title/description/social meta (10), sitemap (5), canonical (5). llms.txt is an emerging convention (llmstxt.org), not yet a formal standard — we treat it as a bonus signal.

Overall score. 25% mobile performance + 25% AI-Readiness + 20% SEO (average of the Lighthouse SEO score and our custom SEO score) + 20% Lighthouse accessibility + 10% best practices. If the Lighthouse portion fails, we show a clearly-labeled partial score instead.

AI Answer Visibility scan — a different question

The audit above asks whether an AI crawler can read your site. The AI visibility scan asks whether an answer engine names you when a buyer describes their problem. A perfectly readable site that nobody has ever written about scores 100 on the first and near zero on the second.

Which engine. Every number in that scan is measured on Google Gemini, and the UI prints which mode it ran in next to the score. It is not a measurement of ChatGPT, Claude, Copilot, Perplexity or Google's AI Overviews — those use different models, indexes and retrieval, and we do not extrapolate to them.

With or without live web search. Grounding with Google Search is not part of Gemini's free API tier, so the scan runs in one of two modes and always tells you which. Grounded (needs a paid Gemini tier) asks the question with a live Google search behind it and reports the sources the engine cited. Ungrounded — the free-tier default — asks the same question against the model's own trained knowledge. That still answers "does the model name you, and who does it name instead", but there are no citations to report, so the citation item is excluded from the score and its weight removed, rather than counted as a zero you did not earn.

How a scan runs. Three steps. (1) We read your homepage and work out your brand name — from schema.org Organization if you publish one, otherwise og:site_name, the page title, or the domain; the scan shows which, and you can correct it. We also look up whether a Wikidata item lists your exact domain as its official website, and check robots.txt for the AI crawlers. (2) We generate six questions from your own homepage copy, phrased the way a customer who has never heard of you would ask. None of them may contain your brand name — a question that names you would answer itself. You see all six and can edit them. (3) Each question is asked separately — with a live Google search behind it on a paid Gemini tier, or against the model's own knowledge on the free one — and we measure the answer.

The score (0–100). A weighted average of five measured items: named in the answer (35), your domain cited among the engine's sources (25), how prominently you are named among the brands in each answer (10), Wikidata entity match on your domain (15), and AI-crawler access in robots.txt (15). Anything that could not be measured is excluded and its weight removed from the denominator — never scored as a zero, because "we could not check" and "the engine never named you" mean opposite things. The scan prints how much weight the score was actually built from, and which items were dropped. One floor on that rule: if no question could be answered at all, there is no score. We will not publish a confident number built only on a Wikidata lookup and your robots.txt.

What you also get. The full text of every answer, which competitors got named instead, and — when the scan ran grounded — the source domains the engine actually cited (the pages you would need to be on). Because grounding cannot be combined with structured output, the competitor roster is requested as a delimited line at the end of the answer; when the model omits it, the table says how many answers it could be extracted from rather than guessing.

What it is not. A sample, not a constant — answer engines are non-deterministic, and the same question can produce a different answer an hour later. Six questions are a sample of buyer intent, not a keyword universe. "Cited" means your domain appeared in the answer's grounding metadata, not that a human read your page. A Wikidata non-match means no match, not that you are unknown. And none of it predicts traffic or revenue. If your brand name is a common word, the scan flags itself as low-confidence rather than reporting a mention rate it cannot stand behind.

Data sent to Google. Running the scan sends your site's public homepage copy and the generated questions to the Gemini API. On the free tier Google may use that data to improve its products. Nothing personal is included.

Quota. The free scan is capped at 20 runs a day across all visitors, and 2 per hour per address, because the free Gemini allowance is finite. When the quota is exhausted the scan says it could not be measured — it never substitutes a cheaper, weaker number.

Monitor, €99/month. Re-runs the same questions on the same engine each month and reports what moved. Billing is not switched on yet: the form there records your interest and you will hear from a human, with a payment link, before anything is charged.

What an automated audit cannot check

We would rather tell you this upfront than let a number mislead you. This tool cannot verify:

· Content quality and relevance — whether your copy convinces anyone, targets the right keywords, or answers real search intent.
· Rankings and traffic — we do not know your Google positions, impressions or visitors.
· Backlinks and authority — no link-graph data.
· Full WCAG conformance — automated tools catch only roughly a third of accessibility barriers. Keyboard flows, screen-reader experience, cognitive load and meaningful alt text need human testing.
· Legal compliance of any kind — GDPR, cookie consent, EAA/ADA, imprint requirements. We may detect technical artifacts, but compliance is a legal judgment.
· Security — this is not a penetration test or vulnerability scan.
· Pages behind logins, paywalls or geo-blocks, and sites that block our fetcher (we audit one public page per run, plus robots.txt / sitemap / llms.txt).
· JavaScript-rendered content — deliberately: we measure what crawlers without JS see. Your human visitors may see more; that gap is exactly what the AI-Readiness score expresses.
· Conversion, UX and business outcomes — a 95/100 site can still sell nothing.
· How specific AI models actually use your content — we measure technical readability signals, not inclusion in any AI index or answer.

Disclaimer

This automated audit is provided free of charge, for informational purposes only. Results are generated by software without human review and reflect a point-in-time analysis of publicly accessible parts of the submitted URL. The audit does not constitute legal, regulatory, accessibility, security or professional advice, and it does not certify or guarantee compliance with WCAG, the European Accessibility Act (EAA), the ADA, GDPR or any other standard, directive or law. Automated tools detect only a subset of potential issues; a complete assessment requires manual expert review, testing with real users and assistive technologies, and for compliance questions, qualified legal counsel. Scores are estimates derived from third-party data (including Google PageSpeed Insights / Lighthouse) and heuristics, may vary between runs, and may be incomplete if a target site blocks automated access. Use of this tool is at your own risk; no warranty of any kind is given. We are not affiliated with Google, OpenAI, Anthropic or Perplexity; product names are used solely to describe crawler behavior.

Privacy

We store anonymous, aggregated usage counters (audits per day, audited domain names, category scores). If you request the PDF report, your email address and the audited URL are stored in Netlify Blobs so the report can be delivered and occasional product updates sent; unsubscribe anytime by replying. We do not sell data. Audited pages are fetched transiently and not archived.

The same applies to the AI visibility scan: if you ask for the monthly report, your email address and the scanned URL are stored the same way, tagged so we know which of the two you asked for. Running that scan also sends the scanned site's public homepage copy and the generated questions to Google's Gemini API — see the AI Answer Visibility section above.

Who publishes this

SiteReadable is an automated tool, published at sitereadable.com. It fetches a page's raw HTML the way a crawler does, runs the checks described above, and reports what it found. Nobody reviews your site by hand as part of a free audit, and no result here is a human opinion.

The references it checks against are public and named on this page, so any result can be verified independently: Google's PageSpeed Insights API and the Lighthouse audits behind it for performance, accessibility and best practices; schema.org for structured data; the Open Graph and Twitter card conventions for social metadata; the Robots Exclusion Protocol for robots.txt, including the published user-agent strings of the AI crawlers; the sitemaps.org XML format; and llmstxt.org for llms.txt, which is an emerging convention rather than a standard and is scored as a bonus signal only.

Scoring is deliberately conservative. The SEO score is a weighted average of the individual checks, and AI-Readiness is an additive 0–100 over the eight weighted items listed above. Our static accessibility layer produces no number at all, because automated testing reaches only part of what accessibility means. Any check that could not run is excluded from its score and flagged as such — never guessed, never counted as a zero, because “we could not check this” and “this failed” are not the same finding.

Questions about a result, a disagreement with one, or a bug: hello@sitereadable.com. We also post at LinkedIn and X.