The Frilly Readability Score (FRS) measures how legibly a website communicates to three readers at once: the people who visit it, the search engines that index it, and the AI assistants that increasingly read it on a visitor's behalf. An audit reads a site's pages, scores them across three lanes, and returns a single 0 to 100 number, a score for each lane, and plain-language findings on what to improve first.
The FRS is a readiness measure, not a forecast. It reports how legible a site is today and what to improve first. It does not predict or guarantee that any engine will feature a site, because no honest instrument can.
Discovery is shifting from a page of links to a direct answer. Search engines increasingly resolve a query on the results page, and a growing share of people ask an AI assistant for a recommendation outright. Local businesses feel this from two directions at once: traffic that once arrived from search declines, and the machines that now summarize the web misread or skip sites that are hard to parse.
Large enterprises are already funding the work to stay legible through this shift. Most local businesses cannot, lacking a content team or a budget for it. The FRS is built to close that gap with a clear, honest diagnosis and a concrete path to improve.
The FRS is not a repackaged SEO audit. It is the product of six months of research and revision as the field changed.
In the spring of 2026, Frilly ran nine live FRS audits across companies ranging from Schema App and Linear to NRC and Ensodata, attended sessions from leading SEO and AEO agencies, sought honest technical review from trusted practitioners, and studied the foundational research on how LLM retrieval actually works. That process forced several redesigns.
Frilly's early versions (v0.1 through v0.5) made three mistakes. Naming them is part of why the current score can be trusted.
They conflated internal retrieval with external AI behavior. Frilly's Smart Chat embeds a customer's content into a controlled index, where structured formats and clean markup genuinely improve retrieval, because the index, the chunking, and the embeddings are all controlled. The early score projected that logic onto ChatGPT, Claude, Perplexity, and Google's AI Overviews, which run their own crawling, indexing, retrieval, and ranking pipelines that a content producer cannot directly influence through formatting.
They over-weighted schema. Early versions assigned more than half of the total score to structured-data dimensions. The retrieval literature is consistent that these systems prioritize semantic relevance and clarity over markup. Schema is valuable for a specific reason: it helps search engines and answer engines understand who an entity is, what it offers, and how its pages connect. It is a comprehension layer, not a discovery lever.
They measured the old paradigm. As Choudary's Reshuffle describes, incumbents tend to apply a new technology to build a faster version of the old system. Scoring schema completeness and meta-tag presence measures the inputs of the old system, not what modern AI selects on, which is relationships, context, and meaning.
The rebuilt FRS corrects all three. Section 8 lists the research basis in full.
The audit scores three lanes, each serving a different reader and each labeled with what it does and does not do. The order is a dependency chain: clear content is the prerequisite for useful markup, which is the prerequisite for meaningful AI preparation.
Evaluates how well the prose on each page communicates what the business is, who it serves, and why it is credible. This is the only lane a site owner controls with words alone, and it is the highest-leverage lane, because every machine consumer selects on clarity and meaning.
Evaluates the structured-data entity graph, the machine-readable labels (schema) that tell search engines like Google and Bing, and answer engines like Perplexity and ChatGPT, who an entity is, what it offers, and how its pages connect, plus crawl and index hygiene.
Evaluates machine-readable files written for AI consumption: an llms.txt site summary, clean per-page markdown, and an intentional AI-crawler policy.
Content and technical dimensions (Lanes 1 and 2) are scored in two parts.
Baseline (0 to 70): does this serve a human reader? The bulk of the score, earned with ordinary good writing and sound structure. The scorer deducts only where a real reader would genuinely be impeded, never for stylistic taste. First-person voice and a distinctive tagline are never penalized.
Advanced (0 to 30): machine-extraction headroom. Framed as opportunity only; its absence is never treated as a flaw. It is earned through placement and structure, such as front-loaded definitions and parallel comparison structures, not by contorting the prose.
Lane 3 uses a different split that fits deterministic files: a 0 to 60 presence floor (the file or policy exists and is well-formed) plus a 0 to 40 quality bonus (structural richness, with a light quality check on the llms.txt summary).
Each lane rolls its dimensions into a 0 to 100 lane score. The overall FRS is the average of the three lane scores, and each lane score is the average of its own dimensions. Because the lanes hold six, five, and three dimensions, this weights each lane equally rather than each dimension: a single Lane-3 dimension moves the total roughly twice as much as a single Lane-1 dimension. The final number is rounded.
Separately from the number, the audit computes a designation by gating, not by banding the score. A lane "passes" only when every one of its dimensions scores at least 60; a partial pass requires a majority at 50 or above, so one weak dimension cannot be averaged away. Thresholds are founder-tunable. The designation is then a sequential ladder: Unclear if Lane 1 has not passed, Clear once Lane 1 passes, Discoverable once Lanes 1 and 2 pass, and Prepared once all three pass.
The designation is a diagnostic of which layers have cleared, not a tier of the FRS number, and the two are deliberately distinct. A site can be "Clear" while its overall number is well short of 100, because the number averages all three lanes and the designation gates on content first. The designation guides the order of fixes; it is used internally and is not shown on the customer's report.
Content-lane scoring is performed by an AI evaluator that reads each page as a first-time visitor would. The same scorer runs the free audit, the full customer audit, and the pre-ship check, so a free score is a real score, measured on a smaller sample. Before any Frilly page about the score ships, its copy is run through the live FRS, so the material that sells the audit is held to the standard the audit measures.
There is no reliable way today to measure how often AI systems mention or cite a specific brand across real user conversations. Frilly is explicit about that. What exists is partial.
Prompt-based inference runs test prompts and observes whether a brand appears. It measures behavior on synthetic queries, not real conversations. Useful as direction, not ground truth.
Search analytics track referral traffic from AI-generated search features. That measures clicks, not mentions.
Schema validation verifies that structured data is implemented correctly. That measures input quality, not output presence.
Frilly reports what is measurable and labels what is inferential. It will not sell a dashboard that pretends to measure something no one can reliably measure yet.
The build rests on the following research, each finding mapped to the design decision it drove. This is the basis for Frilly's authority to publish a score.
This is a living document. As the evidence changes, it is updated, with a note on what changed and why. One recent example: an earlier internal position held that AI systems do not consume schema directly. Current sourced evidence shows that both search engines and answer engines parse structured data to understand a brand, with a delivery caveat for crawlers that do not run JavaScript. Frilly updated the position accordingly, and this document reflects it. The scorer prompts are the source of truth and are Founder-calibrated; when they change, this methodology is re-derived to match.
Since the current rubric went live in July 2026, its dimensions, thresholds, and splits have not changed; the work since has been reliability, delivery, and presentation, not scoring. Every audit is also stamped with a deterministic scorer version, recording the prompt hashes, thresholds, and model policy in effect, so if the rubric ever does change, that change is attributable on a customer's next audit. The score's provenance is versioned.