Frilly LogoCONCIERGE

The Frilly Readability Score: Methodology

1. What the score measures

The Frilly Readability Score (FRS) measures how legibly a website communicates to three readers at once: the people who visit it, the search engines that index it, and the AI assistants that increasingly read it on a visitor's behalf. An audit reads a site's pages, scores them across three lanes, and returns a single 0 to 100 number, a score for each lane, and plain-language findings on what to improve first.

The FRS is a readiness measure, not a forecast. It reports how legible a site is today and what to improve first. It does not predict or guarantee that any engine will feature a site, because no honest instrument can.

2. The problem it addresses

Discovery is shifting from a page of links to a direct answer. Search engines increasingly resolve a query on the results page, and a growing share of people ask an AI assistant for a recommendation outright. Local businesses feel this from two directions at once: traffic that once arrived from search declines, and the machines that now summarize the web misread or skip sites that are hard to parse.

Large enterprises are already funding the work to stay legible through this shift. Most local businesses cannot, lacking a content team or a budget for it. The FRS is built to close that gap with a clear, honest diagnosis and a concrete path to improve.

3. Why the score is credible

The FRS is not a repackaged SEO audit. It is the product of six months of research and revision as the field changed.

In the spring of 2026, Frilly ran nine live FRS audits across companies ranging from Schema App and Linear to NRC and Ensodata, attended sessions from leading SEO and AEO agencies, sought honest technical review from trusted practitioners, and studied the foundational research on how LLM retrieval actually works. That process forced several redesigns.

Frilly's early versions (v0.1 through v0.5) made three mistakes. Naming them is part of why the current score can be trusted.

They conflated internal retrieval with external AI behavior. Frilly's Smart Chat embeds a customer's content into a controlled index, where structured formats and clean markup genuinely improve retrieval, because the index, the chunking, and the embeddings are all controlled. The early score projected that logic onto ChatGPT, Claude, Perplexity, and Google's AI Overviews, which run their own crawling, indexing, retrieval, and ranking pipelines that a content producer cannot directly influence through formatting.

They over-weighted schema. Early versions assigned more than half of the total score to structured-data dimensions. The retrieval literature is consistent that these systems prioritize semantic relevance and clarity over markup. Schema is valuable for a specific reason: it helps search engines and answer engines understand who an entity is, what it offers, and how its pages connect. It is a comprehension layer, not a discovery lever.

They measured the old paradigm. As Choudary's Reshuffle describes, incumbents tend to apply a new technology to build a faster version of the old system. Scoring schema completeness and meta-tag presence measures the inputs of the old system, not what modern AI selects on, which is relationships, context, and meaning.

The rebuilt FRS corrects all three. Section 8 lists the research basis in full.

4. The three lanes

The audit scores three lanes, each serving a different reader and each labeled with what it does and does not do. The order is a dependency chain: clear content is the prerequisite for useful markup, which is the prerequisite for meaningful AI preparation.

Lane 1 — Content Clarity (primary)

Evaluates how well the prose on each page communicates what the business is, who it serves, and why it is credible. This is the only lane a site owner controls with words alone, and it is the highest-leverage lane, because every machine consumer selects on clarity and meaning.

  • Dimensions: Semantic Clarity, Content Structure, Information Density, Terminology Consistency, Voice and Tone, Definition Placement.
  • Bound: clear writing improves how every system understands a site; it is not a guarantee of any particular placement.

Lane 2 — Technical Readability (search and answer engines)

Evaluates the structured-data entity graph, the machine-readable labels (schema) that tell search engines like Google and Bing, and answer engines like Perplexity and ChatGPT, who an entity is, what it offers, and how its pages connect, plus crawl and index hygiene.

  • Dimensions: Entity Identity, Entity Graph Connectivity, Page-Type Differentiation, Rich-Result Eligibility, Crawl/Index Hygiene.
  • Delivery detail: many AI-native crawlers (GPTBot, ClaudeBot, PerplexityBot) do not execute JavaScript yet, so schema must be present in the raw HTML to reach them.
  • Bound: this lane improves comprehension; it does not by itself guarantee a citation.

Lane 3 — AI Forward-Compatibility (emerging discovery)

Evaluates machine-readable files written for AI consumption: an llms.txt site summary, clean per-page markdown, and an intentional AI-crawler policy.

  • Dimensions: llms.txt, AI-crawler policy, per-page accessibility.
  • Bound: no major LLM provider has confirmed it uses these files today. The lane is included because the cost is low, the potential future value is real, and organizing content this way is a useful exercise regardless. It is insurance, not optimization.

5. How scoring works

Content and technical dimensions (Lanes 1 and 2) are scored in two parts.

Baseline (0 to 70): does this serve a human reader? The bulk of the score, earned with ordinary good writing and sound structure. The scorer deducts only where a real reader would genuinely be impeded, never for stylistic taste. First-person voice and a distinctive tagline are never penalized.

Advanced (0 to 30): machine-extraction headroom. Framed as opportunity only; its absence is never treated as a flaw. It is earned through placement and structure, such as front-loaded definitions and parallel comparison structures, not by contorting the prose.

Lane 3 uses a different split that fits deterministic files: a 0 to 60 presence floor (the file or policy exists and is well-formed) plus a 0 to 40 quality bonus (structural richness, with a light quality check on the llms.txt summary).

Each lane rolls its dimensions into a 0 to 100 lane score. The overall FRS is the average of the three lane scores, and each lane score is the average of its own dimensions. Because the lanes hold six, five, and three dimensions, this weights each lane equally rather than each dimension: a single Lane-3 dimension moves the total roughly twice as much as a single Lane-1 dimension. The final number is rounded.

Separately from the number, the audit computes a designation by gating, not by banding the score. A lane "passes" only when every one of its dimensions scores at least 60; a partial pass requires a majority at 50 or above, so one weak dimension cannot be averaged away. Thresholds are founder-tunable. The designation is then a sequential ladder: Unclear if Lane 1 has not passed, Clear once Lane 1 passes, Discoverable once Lanes 1 and 2 pass, and Prepared once all three pass.

The designation is a diagnostic of which layers have cleared, not a tier of the FRS number, and the two are deliberately distinct. A site can be "Clear" while its overall number is well short of 100, because the number averages all three lanes and the designation gates on content first. The designation guides the order of fixes; it is used internally and is not shown on the customer's report.

Content-lane scoring is performed by an AI evaluator that reads each page as a first-time visitor would. The same scorer runs the free audit, the full customer audit, and the pre-ship check, so a free score is a real score, measured on a smaller sample. Before any Frilly page about the score ships, its copy is run through the live FRS, so the material that sells the audit is held to the standard the audit measures.

6. What the score measures, and what it does not

There is no reliable way today to measure how often AI systems mention or cite a specific brand across real user conversations. Frilly is explicit about that. What exists is partial.

Prompt-based inference runs test prompts and observes whether a brand appears. It measures behavior on synthetic queries, not real conversations. Useful as direction, not ground truth.

Search analytics track referral traffic from AI-generated search features. That measures clicks, not mentions.

Schema validation verifies that structured data is implemented correctly. That measures input quality, not output presence.

Frilly reports what is measurable and labels what is inferential. It will not sell a dashboard that pretends to measure something no one can reliably measure yet.

7. What the score will not claim

  • No formatting, schema, or file guarantees a citation or mention by any AI system.
  • The FRS does not promise rankings or traffic.
  • It does not claim to measure AI visibility with precision; that measurement is inferential, not definitive.
  • Frilly does not write a customer's content for them. It coaches structure and clarity; the brand keeps its own voice.
  • Frilly does not position itself as an "AEO vendor." The FRS is an honest audit that serves three different systems with three different mechanics.

8. Research foundations

The build rests on the following research, each finding mapped to the design decision it drove. This is the basis for Frilly's authority to publish a score.

  1. Vaswani et al., "Attention Is All You Need" (2017) — semantic relationships between tokens matter more than surface formatting. Informs: Lane 1 primacy; weighting meaning over markup.
  2. Lewis et al., "Retrieval-Augmented Generation" (2020) — the retrieval pipeline that content must survive inside. Informs: Content Structure (sections must survive being extracted out of context).
  3. Karpukhin et al., "Dense Passage Retrieval" (2020) — dense semantic retrieval outperforms keyword matching. Informs: prioritizing semantic clarity over formatting.
  4. Reimers and Gurevych, "Sentence-BERT" (2019) — sentence embeddings and semantic similarity. Informs: rewarding clear, self-contained sentences and sections.
  5. Nogueira and Cho, "Passage Re-ranking with BERT" (2019) — the re-ranking step often matters more than initial retrieval. Informs: rewarding content quality over attempts to game retrieval.
  6. Liu et al., "Lost in the Middle" (2023) — models perform best with key information at the beginning or end of context. Informs: the Definition Placement dimension.
  7. Venkit et al., "Search Engines in an AI Era" (2024) — citation is probabilistic, not deterministic, across answer engines. Informs: the honest-framing rule; never promising citation.
  8. Choudary, Reshuffle (2025) — platform shifts reward recognizing a new paradigm over optimizing the old one. Informs: rejecting the relabeled-SEO-audit model.
  9. Lee, "I Rank on Page 1, What Gets Me Cited by AI?" (2026) — within a rank band, content features (comparison structure, query-term coverage, statistical density, objective tone) predict AI citation independent of domain. Informs: Content Structure and Information Density.
  10. Lee, "How AI Platforms Search: Fan-Out Query Behavior" (2026) — AI decomposes queries into stochastic internal searches; structural query types stay stable even though the exact searches are random. Informs: the honest measurement position, and writing content that serves many query shapes.
  11. Lee, "The SEO Floor" (2026) — Google rank is the dominant predictor of AI citation, and schema's predictive signal is likely a proxy for publisher quality rather than a direct causal mechanism. Informs: Lane 2 framed as indirect value; schema as a comprehension layer, not a discovery lever.

9. How the methodology evolves

This is a living document. As the evidence changes, it is updated, with a note on what changed and why. One recent example: an earlier internal position held that AI systems do not consume schema directly. Current sourced evidence shows that both search engines and answer engines parse structured data to understand a brand, with a delivery caveat for crawlers that do not run JavaScript. Frilly updated the position accordingly, and this document reflects it. The scorer prompts are the source of truth and are Founder-calibrated; when they change, this methodology is re-derived to match.

Since the current rubric went live in July 2026, its dimensions, thresholds, and splits have not changed; the work since has been reliability, delivery, and presentation, not scoring. Every audit is also stamped with a deterministic scorer version, recording the prompt hashes, thresholds, and model policy in effect, so if the rubric ever does change, that change is attributable on a customer's next audit. The score's provenance is versioned.

Frilly Inc.Madison, WI
DISCOVER:
STAY CONNECTED:
© 2026 Frilly