# Implementation Annex — companion to the Home Service Website Architecture spec v4.2 Source: https://fc-build-spec-a5aaf279.vercel.app/annex.html Publisher: Fortitude Creative. Last updated 2026-09-25. THE DOCUMENT SET - Wireframe v4, the architecture: /wireframe-v4.md - This document, the build spec: /annex.md (sections A-L) - AI Visibility Runbook: /ai-visibility.md (off-site protocols, citation panel) - Operational Runbook: /runbook.md (deployment and maintenance) - Both core documents in one fetch: /llms-full.txt - Everything condensed under 4,000 chars: /brief.txt v4 froze the architecture. This annex is the layer beneath it: the conventions, validation rules, and gates an engineer needs in order to start without asking what anything meant. Two items in the handed-down spec would not have survived contact with a validator. Both are corrected here and marked. --- ## A. CMS data model | Content type | Required reference fields | Cardinality | Validation | |---|---|---|---| | `IntersectionPage` | `parentService` -> ServicePillar; `parentCity` -> LocationHub; `featuredProject` -> Project | 1:1, 1:1, 1:n (min 1) | Reject save if any empty. Reject if a published `IntersectionPage` already exists for the same (service, city) pair. | | `ProblemPage` | `parentService` -> ServicePillar; `topicKey` | 1:1; unique | Reject if empty. Exactly one parent, never two. Reject if another published ProblemPage already holds this `topicKey`. | | `EditorialPost` | `parentService` -> ServicePillar; `topicKey` | 1:1; unique | Reject if empty, if the `topicKey` is taken, or if body under 600 words. See section O. | | `Project` | `parentService` -> ServicePillar; `parentCity` -> LocationHub; `technician` -> Person | 1:1 each | Reject if any empty. Reject if `gps` or `workOrderRef` missing. Reject publish before `technicianApprovedAt` is set. | | `ServicePillar` | none, queries inverse relationships | — | Slug must match the locked service registry in section B. | | `LocationHub` | none, queries inverse relationships | — | Requires NAP, `geo`, `openingHours`, GBP CID before publish. | | `AuthorityClaims` (singleton) | Not a reference type. Holds `yearEstablished`, `licenses[]`, `serviceArea[]`, `certifications[]`, `fleetSize`, `employeeCount`, `awards[]`, `bbbRating`. **See K.1 for the full field definitions.** | 1 per site | Reject save on any empty required field. License numbers pass per-state format validation. No LocationHub publishes while the singleton is incomplete. Lives on the section I config singleton, not beside it. | | `EstimatorPage` | `parentService` -> ServicePillar; `outputTiers[]` (min 2); `sampleScenarios[]` (min 2) | 1:1; 2:n; 2:n | Reject if outputTiers < 2, ctaText or disclaimer empty, or body under 300 words. No currency field exists on this type (M.1). | | `EquipmentModelPage` | `parentBrand` -> BrandPage; `installedByTechnicians` -> Person[] | 1:1; 1:n | Four gates: brand in `installedBrands[]`, model in `approvedModels[]`, within `totalModelCap`/`perBrandModelCap`, and at least one Project references it. Reject under 600 words (M.2). | | `SpecialOffer` (repeatable on the /specials/ singleton) | `linkedService` -> ServicePillar (optional) | 0:1 | Reject save when `validUntil` is in the past. Nightly job archives expired offers and removes them from the render (M.3). | | `Person` | `credentials[]`, `sameAs[]`, `knowsAbout[]` | 0:n each | Profile cannot publish with zero projects attached. `sameAs` is optional at the Person level and required (min 2) on Organization; see K.4. | **Inverse queries drive every generated block.** A ServicePillar renders its city list by querying IntersectionPages that reference it; a LocationHub renders its service index the same way. Editors write copy. They do not build navigation, and they cannot override a generated link's target or anchor text. --- ## B. URL and slug conventions ### Locked rules - **Pattern:** `/[service]/[city]/`, service first, always. - **Trailing slash required.** Non-slash variants 301 to the slash form. - Lowercase, hyphen-separated, ASCII only. No stop words, no dates, no IDs. - **Singular service terms.** One convention, no exceptions, so nobody has to remember which pillar was plural. - **No query parameters for city,** no client-side city switching, no JS that swaps city content on one URL. Each intersection is server-rendered and independently crawlable. This is the rule the entire strategy rests on. ### Locked service registry Every pillar slug, frozen. Adding a service means adding to this list, not inventing a slug at publish time. ``` /ac-repair/ /ac-installation/ /furnace-repair/ /furnace-installation/ /heat-pump-repair/ /heat-pump-installation/ /maintenance-plan/ /indoor-air-quality/ ``` **Correction to v4 (applied):** the v4 sitemap showed `/heat-pumps/` and `/maintenance-plans/`, which violate the singular rule this annex locks. Both are corrected in the registry above and in the v4 document. Catching it now costs an edit; catching it after launch costs a redirect map. ### City disambiguation, decided now rather than at expansion - Default slug is the bare city name: `/ac-repair/phoenix/`. - When two cities in the registry share a name, **both** take a state suffix: `/ac-repair/portland-or/` and `/ac-repair/portland-me/`. Never suffix only the newcomer, because that silently changes the meaning of the original URL and leaves the older page ranking for an ambiguous term. - The city registry includes planned expansion markets, not just live ones. A collision is detected against the plan, so the suffix is decided before either page exists. - The slug generator enforces this on save. It is not a naming guideline. --- ## C. JSON-LD entity graph ### Correction to the handed-down example (would have failed validation) The supplied graph typed each project as `HowTo` and gave it `performer` and `locationCreated`. `HowTo` has no `performer` property, so that reference is dropped on parse, and Google retired HowTo rich results for most surfaces, so the type buys nothing while inviting a mismatched-markup signal. A completed job is a record of work, not a set of instructions for the reader. **Corrected below:** projects are typed `Article` with `author`, `about`, and `contentLocation`, which are real properties that resolve. Reserve `HowTo` for problem-page content that genuinely walks a reader through steps. ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "@id": "https://example.com/#org", "name": "Example HVAC Inc.", "url": "https://example.com", "logo": "https://example.com/logo.png", "sameAs": ["", "", ""] }, { "@type": "LocalBusiness", "@id": "https://example.com/service-areas/phoenix/#localbusiness", "name": "Example HVAC Inc. - Phoenix", "parentOrganization": { "@id": "https://example.com/#org" }, "areaServed": { "@type": "City", "name": "Phoenix", "addressRegion": "AZ" }, "telephone": "+1-555-123-4567", "address": { "@type": "PostalAddress", "addressLocality": "Phoenix", "addressRegion": "AZ" }, "geo": { "@type": "GeoCoordinates", "latitude": 33.4484, "longitude": -112.0740 }, "openingHoursSpecification": [ "/* per day */" ], "sameAs": "" }, { "@type": "Service", "@id": "https://example.com/ac-repair/#service", "name": "AC Repair", "serviceType": "HVAC Repair", "provider": { "@id": "https://example.com/service-areas/phoenix/#localbusiness" } }, { "@type": "Person", "@id": "https://example.com/team/mike-rodriguez/#person", "name": "Mike Rodriguez", "worksFor": { "@id": "https://example.com/#org" }, "hasCredential": { "@type": "EducationalOccupationalCredential", "credentialCategory": "NATE Certification" } }, { "@type": "Article", "@id": "https://example.com/projects/capacitor-replacement-ahwatukee/#project", "headline": "AC Capacitor Replacement in Ahwatukee", "about": { "@id": "https://example.com/ac-repair/#service" }, "author": { "@id": "https://example.com/team/mike-rodriguez/#person" }, "publisher": { "@id": "https://example.com/#org" }, "contentLocation": { "@id": "https://example.com/service-areas/phoenix/#localbusiness" }, "datePublished": "2026-09-25", "image": ["", ""] } ] } ``` **Emission rule:** every page emits its own slice of this graph, and every `@id` it references must resolve to a node emitted somewhere on the site. Validate against the Rich Results test before launch, and again in CI on every template change. --- ## D. Anchor text conventions | Source | Target | Anchor | |---|---|---| | Service pillar | Intersection page | Exact: *AC Repair in Phoenix* | | Intersection page | Problem page | Symptom-rich partial: *AC blowing warm air* | | Problem page | Service pillar | Exact service term: *AC Repair* | | Project | Service pillar | Exact service term | | Project | Technician profile | Full name: *Mike Rodriguez* | | Intersection page | Location hub | City name: *Phoenix service area* | Templates generate all of these. Manual override is disabled in the editor. One generated link per target per page, so a page never carries the same anchor to the same URL three times. --- ## E. Template defaults - **Meta description:** required; publish blocked when empty. - **Canonical:** self-referential by default, override only alongside an explicit redirect. - **Phone:** above fold, `tel:` link, present in `contactPoint` schema, pulled from the singleton in section I. - **Breadcrumb:** rendered and marked up on every template except the homepage. - **Title tag:** required, and constrained. Target 30-60 characters; **warn above 60, fail above 70**, fail when empty, and fail when it duplicates another title anywhere on the site. Patterns are set per page type and generated, not typed: intersection pages use `{Service} in {City} | {Business}`, which also satisfies the citation-durability rule in K.7. *Why this is now explicit:* a competitor crawl found branch pages carrying title tags of 491 to 536 characters. Nothing in a CMS prevents that by default, the page still validates, and the title is the single most visible piece of text a search result has. This spec required a meta description from the start and never constrained the title. That was an omission. - **Hero image:** 100 KB or under, WebP, `preload` hint in ``, explicit width and height to hold layout. - **Video:** lite-YouTube facade. No third-party JS on load. - **Third-party widgets** (reviews, financing, chat): lazy-loaded through `IntersectionObserver`, never render-blocking. - **Fonts:** self-hosted, `font-display: swap`, subset. ### Page-type template rules - **LocationHub order is fixed:** H1, then the machine-readable About block (K.2), then NAP, hours and map, then the service index, then the project feed. The About block sits above NAP because it is the first prose an extractor reaches, and it should be the summary you wrote to be quoted rather than an address. Its data source is the AuthorityClaims singleton plus that hub's own city fields. - **The ProblemPage template does not import the FAQ block component at all.** Absence in the template is the first line of defence; the CI gate in section F is the one that actually enforces it (K.3). - **Intersection templates** render the H1 and opening sentence from the citation-durability pattern in K.7, both carrying business name and city. --- ## F. CI and QA gates ### The build fails when - Any template renders without a meta description, self-canonical, phone number, breadcrumb, or JSON-LD block. - Any `@id` reference in an emitted graph does not resolve to a node on the site. - Lighthouse mobile performance scores under 80 on any of the six template archetypes. - Cumulative Layout Shift exceeds 0.1 on simulated mobile. - A slug fails the section B validator. - **No-JS crawl gate (K.6):** an intersection page renders under 400 bytes of body text with JavaScript disabled, or emits its JSON-LD client-side. Test: `curl -s -L | wc -c` against the extracted body text. - **Date freshness (K.5):** an article-type page is missing `dateModified`, or carries one older than 365 days. - **Entity corroboration (K.4):** `Organization.sameAs` has fewer than 2 entries. Person profiles are not gated. - **FAQ scope (K.3):** `FAQPage` schema is detected on any page whose content type is `ProblemPage`. This is the enforcement of record for the FAQ rule; the template omitting the component is a convention, and conventions lose to a future import. - **Alt text (section H):** any rendered `` is missing its `alt` *attribute*. An explicitly empty `alt=""` passes only on an image also marked `role="presentation"` or `aria-hidden="true"`, because empty alt is the correct markup for a decorative image and forcing a description onto one makes the screen-reader experience worse, not better. - **License expiry (K.1):** any license in the AuthorityClaims singleton is past its `expires` date. - **Title tag:** missing, over 70 characters, or duplicated elsewhere on the site. - **Topic collision:** two published pages of the same content type share a `topicKey` (see below). - **G.5 redirect integrity:** a new 301 whose source URL did not exist in the previous crawl, or a source mapping to more than one target. Volume caps are per action type — CONSOLIDATE 100, manual addition 10 — and exceeding a cap triggers mandatory human review rather than automatic failure, because a large legitimate consolidation should be looked at, not blocked. - **Cost rule (M.0):** a dollar amount (`$` followed by a digit, or "N dollars") or a banned positioning word (affordable, cheap, cheapest, budget-friendly, bargain) appears in any rendered page body. - **Estimator no-JS content (M.1):** an estimator page returns under 300 words with JavaScript disabled, *or* is missing `` with 2 or more rows. Both conditions are checked, because word count alone passes a page of prose with no answer in it. - **Model exemption expiry (M.2):** a model page whose `exemptionExpiresAt` has passed still has zero inbound Project references. The nightly job unpublishes it, sets `robots: noindex, nofollow`, removes it from `/sitemap-model.xml`, and alerts content plus the SEO Lead. - **Review counter integrity (M.4):** `aggregateRating.reviewCount` is lower than the floored number shown in the header. ### The nightly crawl reports - Orphan pages, meaning fewer than 3 inbound internal links. - Core Web Vitals regressions against the previous run. - 4xx and 5xx responses. - Schema validation failures against the structured data spec. - Intersection pages published without a live featured project. - **Sitemap hygiene:** any URL in a sitemap that returns a non-200, or that carries `noindex`. A sitemap is a list of pages you are asking to have indexed; a `noindex` page in it is a contradiction you are sending to a crawler on purpose. - **Near-duplicate topics:** pages of the same type whose H1 and opening 200 words exceed a 0.8 similarity score. Reported, never gated — near-duplication is a judgement call, and a threshold that blocks publishing would be wrong as often as it was right. **Warning, never a failure:** `/llms.txt` or `/llms-full.txt` missing, or out of sync with the current architecture. Lint it, ship anyway. **Standing upgrade trigger:** if Google, OpenAI, Anthropic, or Perplexity formally documents programmatic use of `llms.txt` or an equivalent convention, this becomes a hard failure on all branches immediately, with no further deliberation required (K.8). *Declined twice, on the same grounds: failing production deploys on a missing `llms.txt`. Blocking a real release over a file that no retrieval system has documented using inverts the risk, and the second request restated the ask without answering that objection. The trigger above is the part worth keeping, and it fires the day the evidence exists.* ### Tiered content floors The 350-word absolute minimum from section 1 does not move: nothing publishes below it, ever. The tiers below sit *above* that floor and set the target for a market's priority. | Tier | IntersectionPage | LocationHub | |---|---|---| | tier1 | 1,500 words | 3,000 words | | tier2 | 800 words | 1,500 words | | tier3 (default) | 500 words | 750 words | - **Assignment:** `contentTier` on the LocationHub, defaulting to tier3. The SEO Lead owns the initial mapping across all locations as a one-time strategic pass; marketing ops then executes against the assigned tiers. - **Tier-1 eligibility:** a location needs **6 or more published Projects** with photos, geotagged to its service area, before tier-1 can be assigned. Tier-1 asks for 1,500 words per intersection page, and without the job history behind it that target is met with padding, which is the failure mode the whole content floor exists to prevent. - **Override:** tier-1 may be assigned without the project threshold for an aggressive new-market launch, with named approval and the justification written into the commit message. - **Scope of the count:** narrative copy only. Promo badge text, financing terms and other conversion module strings do not count toward it, or the floor becomes satisfiable with furniture. - **Enforcement:** CI warns below the tier floor. The deploy proceeds only when the commit message carries `--ack-short-copy={pageSlug}` for each flagged page, which makes shipping thin copy a deliberate, attributable act rather than a silent one. - **Tracking:** the nightly dashboard flags acknowledgements older than 30 days. - **At 90 days below floor:** the page is recommended for soft-unlink from global navigation, requiring the same named approval as the GBP tier-3 escalation in M.5. It stays published and indexable. **Hard word-count gates are rejected permanently**, and the directive's own reasoning is the right one: blocking a page that could earn reviews and links until it reaches a word count removes the mechanism that makes the content investment pay. Gates run on pull requests against the six template archetypes, not against every published page. Page-level problems are the nightly crawl's job; template-level problems must never reach production. --- ## G. Operational rules **Minimum viable launch footprint:** top 1 to 2 services x top 1 to 2 cities, each with a live Featured Project. Do not wait for grid coverage. An unpublished intersection costs nothing; a thin one damages the cluster it sits in. - **Project reuse:** a project may appear on its location hub (one of many) and on its service pillar (in a recent-work block). It may appear on **exactly one** intersection page. No reuse across intersections, because the same job in two cities is the first lie a reviewer would catch. - **GBP SOP stands as written in v3:** weekly posts, Q&A answered inside 24 hours, photo cadence matched to the project pipeline. - **Content supply governs page count.** The grid will land below its theoretical maximum. That is the expected outcome, not a shortfall to explain away. - **Minimum project count before a city opens:** 3 documented projects in that service area. One project is a page; three is a pattern a reviewer believes. A city with fewer stays unpublished until the trucks have been there. - **Before any city expansion:** the capture workflow is running with field operations, and ongoing project documentation has a named owner. Expansion that outruns capture produces exactly the thin grid this architecture exists to avoid. --- ## H. Accessibility - Descriptive `alt` text on every meaningful image. Before-and-after project photos describe the equipment and the condition, not "image1". - Accordion and FAQ components carry `aria-expanded` and are keyboard operable. - Captions on all video content. - Body text contrast at 4.5:1 minimum. - Visible focus states on every interactive element, including the sticky phone CTA. --- ## I. Brand governance These values live in one config singleton (CMS singleton or environment, one of the two, chosen once). Templates read from it. Editors cannot fork them, because a second phone number in the wild is a tracked-call attribution failure and a NAP inconsistency at the same time. | Key | Used by | |---|---| | `primaryPhone` | Header, every CTA, `contactPoint` schema, `tel:` links | | `napFormat` | Footer, location hubs, LocalBusiness schema, GBP parity checks | | `gbpCid` per location | `sameAs` on each LocalBusiness node | | `legalName` | Organization schema, footer, financing disclosures | | authority claim fields (K.1) | The About block, Organization and LocalBusiness schema, trust rows. Same singleton, same rules. | | `gbpQualityThresholds` | `minReviews`, `minRating`. Read by the nightly report in L.1. Operations tunes them without a deploy. | | `aiReferrerDomains[]` | The assistant referrer list in L.3. Analytics reads it; quarterly review updates it. | | `llmCrawlerRegex[]` | The user-agent patterns in L.4. CI compiles each one. | | `estimatorParamsAllowlist` | Querystring parameters the estimator may accept; everything else 301s to canonical (M.1). | | `installedBrands[]`, `approvedModels[]` | The brand registry and the demand-gated model list, maintained quarterly by the SEO Lead (M.2). | | `totalModelCap`, `perBrandModelCap` | 20 and 5. Enforced in CMS validation (M.2). | | `allowFinancingTerms` | Default false. Financing blocks state availability and link out; numeric credit terms require legal sign-off and a recorded decision (N.6). | | `contentTier` per LocationHub | tier1/tier2/tier3, default tier3. Sets the word-count target above the 350-word absolute floor (section F). | | `allowOfferAmounts` | Default false. Dollar amounts on specials are off per the cost rule; enabling it per client is a recorded decision (M.3). | | `reviewCountDisplay` | Header format, floor rule, and the schema-exactness requirement for the review counter (M.4). | | `kRequirementStatus` | Per-requirement `required` or `optional`. Changed only in a pull request that also changes this document, so config and prose cannot drift apart (L.6). | --- ## J. Zero-click event schema (closes the flagged gap) The handed-down spec left remarketing instrumentation "for sprint tickets." Leaving it undefined is how it ships as three inconsistent event names. Defined here instead. | Event | Fires when | Parameters | |---|---|---| | `search_bounceback` | Referrer carries a search query and viewport time stays under 7 seconds. **Not zero-click:** a true zero-click visitor never reaches the site and cannot be measured from it. This is a return-to-search, which is still worth remarketing to. | `page_type`, `page_slug`, `dwell_ms` | | `diagnostic_start` | Diagnostic tool receives its first input | `page_type`, `symptom`, `framing` (urgent or checkup) | | `diagnostic_complete` | Tool reaches a recommendation | `symptom`, `outcome`, `steps_completed` | | `emergency_cta_click` | Above-fold urgent CTA tapped | `page_type`, `page_slug`, `service`, `city` | | `checklist_download` | Secondary capture submitted | `page_slug`, `capture_type` (email or sms) | | `call_initiated` | `tel:` link activated | `page_type`, `service`, `city`, `position` (header, hero, footer) | Every event carries `service` and `city` where the page type has them, so remarketing audiences can be built per intersection rather than per site. Pixel placement is a template concern, not a page concern: it ships once in the base layout. --- ## K. AI retrieval optimization Everything above optimizes for search engines reading the site, including their AI features. This section covers the other half: being quotable when someone asks ChatGPT, Perplexity, or Claude for an electrician in their city. That traffic never appears as a ranking, and the page that earns it is built differently. ### Two reconciliations, so the spec does not contradict itself **One singleton, not two.** The authority claims below extend the existing config singleton in section I. They do not create a second one. A site with two sources of truth for business facts has none. **FAQ markup is not loosened.** Section C still bans `FAQPage` on a whole page. K.3 wraps only the discrete Q&A block on a comparison page, and only when the questions are real. That is the same rule applied, not an exception to it. ### K.1 Authority claims, added to the section I singleton Verifiable business facts every template can pull from. Editors cannot fork them. A LocationHub cannot publish while any required field is empty. | Field | Notes | |---|---| | `yearEstablished` | The year, not "over 20 years." A year survives being quoted a decade later; a relative claim rots. | | `serviceArea[]` | Cities and ZIP codes served. Feeds `areaServed` and the city registry in section B. | | `licenses[]` | Array of objects, one per license, so acquisitions with two licenses in one state and staggered renewals both fit: `{ state, type, number, expires }`. Example: `{ "state": "AZ", "type": "ROC", "number": "123456", "expires": "2027-06-30" }`. **Format validation:** a per-state regex map checks the number on save; a state absent from the map does not block publishing, the field saves marked *unverified format* and goes to the Compliance Lead. **Expiry is a gate, not a display field:** publishing with an expired license fails the build, and the nightly job flags any license inside 60 days of expiring. | | `certifications[]` | NATE-certified technician count, manufacturer authorizations. | | `fleetSize`, `employeeCount` | Capacity signals. Numbers, kept current. | | `awards[]` | Each with a date. An undated award is unquotable. | | `bbbRating` | Rating plus the profile URL for `sameAs`. | **Deployment coupling in the regex map, named rather than hidden.** The per-state format map lives in code, so a new state needs a DevOps change before its numbers validate. Moving the map into the CMS so a Compliance Lead could edit it is the wrong trade: a regex typed into a text field can be silently over-permissive, and `.*` passes every check while validating nothing. Validation you cannot trust is worse than none, because it reads as verified. So the map stays in code, the Compliance Lead owns its contents and files the DevOps request, and an unmapped state never blocks a launch: the license saves as *unverified format* and goes to manual review, because a whole market waiting on a regex ticket is a worse failure than a typo in a license number. **One field name, deliberately not the one that was asked for.** The review asked for `yearsInBusiness`. This spec stores `yearEstablished` instead. A count of years is correct on the day it is typed and wrong every day after, and nothing in the build will ever tell you it went stale. A year is a fact that stays true, and "years in business" is a one-line computation from it at render time. Store the fact, derive the phrasing. **Pending values, and why they need a shape.** A build almost always starts before the client sends the certificate. The wrong answer is a plausible-looking placeholder, because a plausible placeholder is one merge away from being published as a real licence number. The rule: - A pending entry is `{ state, type, number: "PENDING-VERIFICATION", expires: null, status: "pending" }`. It is obviously not a licence number, and the format validator recognises the sentinel rather than trying to parse it. - **The template omits the licence line entirely while status is pending.** It never renders "PENDING-VERIFICATION", and never renders a partial claim. Absent beats wrong. - The build does not fail, so work continues, but **no page making a licensure claim publishes** while the sentinel is in place. - **The placeholder itself expires.** Pending for more than 30 days escalates to the Compliance Lead, because a placeholder with no expiry is how "temporary" becomes permanent and a field nobody remembers stays empty for a year. **These are legal claims, not copy.** A license number or certification count published wrong is a regulatory problem, not an SEO problem. The singleton needs an owner who verifies each value against the source document, and a review whenever a license renews. Wire the fields; do not invent the values. ### K.2 Machine-readable About block Every LocationHub renders a 150 to 200 word structured summary in third person, directly below the H1, generated from the singleton plus that location's fields. It is factual, quotable prose written to be extracted, not marketing copy. > [Business Name] has served the [City/Region] area since [Year], providing residential heating, > ventilation, and air conditioning services. Licensed in [State] ([License > Type] #[Number]), the company employs [N] NATE-certified technicians and specializes in [primary > services]. [Business Name] is an authorized [Manufacturer] dealer and maintains > [credential/rating]. Third person throughout. "We have served" cannot be quoted by a model answering someone else's question; "[Business Name] has served" can. ### K.3 FAQPage schema on comparison pages only - Comparison and decision pages (repair vs replace, heat pump vs furnace, AC size guide) carry `FAQPage` wrapping **3 to 5** genuine Q&A pairs. - Each question is a real question someone asks, not a heading rewritten with a question mark. Each answer is self-contained and readable with no page context around it. - **Problem and symptom pages do not get FAQ schema.** They are single-intent diagnostic content; adding Q&A markup creates circular structure and dilutes the topical focus that makes them rank. ### K.4 Person schema, extended | Field | Requirement | |---|---| | `Organization.sameAs[]` | **Required, minimum 2.** The GBP profile URL counts as one. Populate the rest from the verified aggregator list the NAP audit produces. This is the entity the assistants are actually trying to identify. | | `Person.sameAs[]` | **Optional.** Downgraded from required: a residential technician's LinkedIn is thin corroboration, and gating every profile on it stalls the build for the least certain payoff in section K. The field stays available and is worth filling for a lead tech with a real professional footprint. | | `hasCredential[]` | Structured credentials: NATE certification, state license, manufacturer training. | | `knowsAbout[]` | References to the ServicePillar pages this person is genuinely expert in. | `sameAs` is what turns a name on a page into an entity a model can corroborate elsewhere. It carries its weight at the Organization level, where the GBP profile and the verified aggregator listings do real identification work. At the Person level it was demoted after review: requiring a findable professional profile for every residential technician stalls the build for the least certain payoff in section K. ### K.5 Article dating, enforced (freshness hygiene, not an AI signal) Stated honestly: this is a freshness signal for traditional crawlers and a forcing function against content decay. Its effect on LLM retrieval is assumed, not demonstrated. It stays required because content rot is real; it should stop being sold as an AI optimization. - Every article-type page (problem pages, comparison and decision pages) emits `datePublished` and `dateModified` in its JSON-LD. - **CI gate:** the build fails when `dateModified` is missing, or older than 365 days, on any article-type page. - `dateModified` changes only when the content actually changed. A build that touches the date without touching the substance is lying to the crawler, and the date-freshness gate becomes a ritual instead of a signal. ### K.6 Crawl accessibility gate - Every intersection page renders **400 bytes or more of body text with JavaScript disabled**. - JSON-LD is present in the raw HTML response, never injected client-side. - **Test method:** a plain `curl` fetch, or a headless browser in no-JS mode, during build validation. The build fails on any page that does not clear it. Most assistant fetchers do not execute JavaScript. A page that needs JS to show its content is, to them, an empty page. ### K.7 Citation durability The H1 and the first sentence of body copy on every intersection page carry the business name and the city: - **H1:** AC Repair in Phoenix | [Business Name] - **First sentence:** "[Business Name] provides 24/7 AC repair service throughout the Phoenix metropolitan area..." Models quote sentences and drop links. If the identifying information is inside the sentence, the attribution survives the citation being stripped. ### K.8 LLM accessibility files `/llms.txt` indexing the key pages in plain English, and `/llms-full.txt` carrying the full text for a single fetch. Regenerate whenever the architecture changes or a new service area launches. **Status:** both files are already live on this document site, which is also the working proof of the pattern. Honest caveat, carried over from the source spec: their utility for third-party crawlers is plausible and unproven at scale. Cheap to ship, so ship them; do not count on them. ### K.9 Crawler licensing stance Decide explicitly whether AI systems may cache and redistribute the content, rather than letting silence pick for you. For a local service business the answer is nearly always **permissive**: the entire goal is to be quoted to someone shopping for a contractor, and a restriction that keeps you out of an answer costs a lead to buy nothing. - **Permissive:** a `License:` line in the `/llms.txt` header. - **Restrictive:** documented in `robots.txt` or the emerging `ai-permissions.txt` convention. This is a client decision, made per client and recorded in the build, not a default anyone should inherit silently. ### K.10 Content supply, stated plainly > Content supply is an operational constraint external to this specification. CMS gates prevent > publishing intersection pages without Featured Projects; they cannot generate those projects. The > business operations layer must deliver project content at the volume the architecture assumes. ### K.11 Speakable markup **Required, enforced as a warning.** - **Scope:** LocationHub and IntersectionPage. - **Target:** the page's `WebPage` node. **Corrected from the directive**, which targets `LocalBusiness` and `Service`. Google has only ever documented `speakable` for `Article` and `WebPage` in news-publisher contexts, and never for `LocalBusiness` or `Service`, so emitting it on those two types is markup for a combination that has never been supported anywhere. If it is worth emitting at all, emit it on the type where it is at least defined. - **Value:** a CSS selector pointing at the introductory paragraph, the first two or three sentences. - **Constraint:** the selected text resolves to something non-empty and stays at or under 300 characters. - **Enforcement:** CI warns when the selector is missing, resolves empty, or exceeds the limit. It does not fail the build, because a selector that stops matching after a template edit is a content problem, not a broken deploy. - **Deprecation clause:** drop the requirement if Google formally deprecates `speakable` for LocalBusiness. Worth knowing now: `speakable` support has only ever been documented for news content, so treat this as cheap positioning rather than a mechanism with evidence behind it. ### Acceptance criteria - AI retrieval complete - Every LocationHub renders the About block, pulling from the singleton. - Every Person entity has at least one validated `sameAs`. - Every comparison page carries valid `FAQPage` schema over genuine Q&A pairs, and no symptom page carries any. - CI passes the no-JS render gate and the date-freshness gate. - No intersection page publishes without its Featured Local Project. ### Build order Slotted into the existing phases from v4, by dependency rather than sprint number: - **With the content model (v4 phase 2):** authority fields on the singleton, extended Person type, About block component. - **With the templates (v4 phase 3):** the no-JS and date-freshness CI gates, citation-durability H1 and opening sentence, FAQ schema on comparison templates. - **After launch:** the LLM accessibility files, regenerated on every architecture change. --- ## L. Off-site authority and measurement **Scope, stated plainly, replacing any percentage anyone quotes.** Section K covers on-site preparation for AI retrieval. It is necessary and not sufficient. When an assistant answers "best AC repair in Phoenix," it assembles that answer mostly from off-site surfaces: Google Business Profile, review platforms, aggregator listings, local listicles, forum threads. Section L gates publication on minimum off-site readiness and builds the loop that measures whether any of this works. The AI Visibility Runbook (https://fc-build-spec-a5aaf279.vercel.app/ai-visibility.md) holds the human protocols that generate that authority. ### L.0 The runbook is a build dependency The build fails when `/ai-visibility.md` or its `runbook.yaml` is missing, or when `runbook.yaml` has a null `owner` or `reviewCadence`. Off-site work is the half that quietly does not happen; tying it to the build makes its absence loud. ### L.1 GBP identity required, GBP quality measured (gate rejected) **Required fields on LocationHub:** `gbpCid` and `gbpPlaceId`. A location hub cannot publish without them, because they are the join between this site and the surface the assistants actually read. **What was asked for, and why it is not implemented.** The review asked that publishing be blocked when a location's Google rating is under 4.5 or its review count under 20. Three reasons that gate does not ship: - **It inverts cause and effect.** A new location has 8 reviews precisely because it has no web presence yet. Refusing to publish its page removes the thing that generates reviews. - **It hands a third party a veto over publishing.** A Places API outage, a quota exhaustion, or a billing lapse means nobody can publish anything. Build systems must not depend on a live external API to emit a page. - **It fires on noise.** A 4.49 average blocks the release; one new five-star review unblocks it. That is not a quality gate, it is a coin toss with a cron job. **Implemented instead:** the thresholds live in the section I singleton as `gbpQualityThresholds.minReviews` and `.minRating`, the nightly job reads the Places API and reports every location against them, and a location below threshold ships with a recorded acknowledgement naming who accepted it. Measured, visible, owned. Never a publish block. ### L.2 NAP verification freshness - **Required field on LocationHub:** `napLastVerifiedAt`. - **Creating a new LocationHub** is blocked when the timestamp is null or older than 90 days. - **Updating an existing one** warns but never blocks. Otherwise a stale audit stops you fixing a typo, and the fix people reach for is faking the timestamp. - The nightly job flags anything past 60 days so the audit is scheduled, not discovered. **Why 90 and not 30:** the review asked for 30 days while its own runbook sets the NAP audit at quarterly. Ninety matches the cadence that is actually staffed. A gate tuned tighter than the process behind it does not raise the standard, it teaches people to bypass the gate. ### L.3 AI referrer tagging Analytics tags traffic arriving from assistant platforms as `channel=ai_generated`, with the domain list in the section I singleton as `aiReferrerDomains[]`: ``` chatgpt.com perplexity.ai claude.ai gemini.google.com copilot.microsoft.com you.com ``` **One domain removed from the supplied list.** `bing.com` is not on it. Bing Copilot and ordinary Bing organic search share that referrer, so including it files a large volume of plain search traffic as AI-generated. The measurement exists to tell you whether AI is sending anyone; a list that cannot distinguish the two produces a number that always looks encouraging and means nothing. `copilot.microsoft.com` and `claude.ai` are added in its place. **Deployment test:** request the homepage with a spoofed `Referer` from each listed domain and assert the beacon fires with `channel=ai_generated`. The deployment fails if it does not. ### L.4 Assistant crawler logging Middleware matches incoming User-Agent headers against `llmCrawlerRegex[]` in the singleton and logs hits: `GPTBot`, `anthropic-ai`, `ClaudeBot`, `Claude-User`, `Google-Extended`, `PerplexityBot`, `CCBot`, `OAI-SearchBot`. - CI compiles every pattern; a syntax error fails the build. - An unmatched user agent exceeding 1,000 requests a day emits a warning naming it, so the list grows from observation rather than memory. - Quarterly review of the list sits in the runbook. This is the cheapest honest signal in section L: it tells you whether assistants fetch the site at all, which is a precondition for every other claim here. ### L.5 Citation panel data The measurement table exists before launch, and CI asserts its schema. Storage is vendor-agnostic: BigQuery, Postgres, or Supabase all qualify. The schema does not. ``` run_id STRING timestamp TIMESTAMP assistant STRING // chatgpt, perplexity, gemini, claude, copilot query STRING appeared BOOLEAN link_present BOOLEAN sentiment STRING // nullable: positive, neutral, negative, hallucination snapshot_url STRING // nullable: path to stored response or its hash ``` How the table is filled is the runbook's problem. That it exists, and that nothing ships before it does, is this spec's problem. ### L.6 Falsifiability, without letting the spec rewrite itself (automation rejected) The review asked for quarterly automated correlation between each K requirement and `appeared=true`, auto-demoting any requirement showing no positive correlation for two quarters. - **There is no variance to correlate against.** Every page implements every K requirement, by construction, because the CMS refuses to publish otherwise. A variable that never varies has no correlation with anything. The computation is not hard, it is undefined. - **The sample cannot carry the weight.** Ten queries across five assistants, monthly, is roughly 150 observations a quarter, confounded by ranking changes, review volume, competitor activity, and model updates nobody outside the labs can see. - **The consequence is silent.** A spec that edits its own requirements from a statistic produces a document nobody can trust to still say what it said last quarter. **Implemented instead:** the quarterly review is real and scheduled, and it asks the answerable question, which is not "does K.2 correlate with appearing" but *what did the assistants actually cite when we appeared, and when we did not.* That reads the sources, which is where the answer lives. A human proposes demotions with evidence, and a demotion is a normal documented change to this spec, made by a person who signs it. A requirement demoted this way moves to `optional` in `kRequirementStatus` in the singleton, and the change lands in this document in the same pull request. Config and prose never disagree. --- ## M. v4.1 surfaces: estimator, equipment pages, specials, review counter Three new page types and one global element. Each inherits the same envelope as everything else here: server-rendered, gated on real supply, measured. None of it changes the section L conclusion that off-site authority dominates AI citation on local commercial queries. These are conversion and traditional-SEO surfaces that must not become the thin-content vector the rest of this spec exists to prevent. ### M.0 The cost rule, binding on everything in this section **No dollar amounts anywhere in client content. Not a price, not a range, not an hourly or flat rate.** Also banned as positioning: affordable, cheap, cheapest, budget-friendly, bargain. What is allowed, and actively wanted: cost *intent* and cost *factors*. Pages may target "cost to replace AC in [city]" and should. They answer with what drives the number — labor complexity, equipment access, system age, efficiency tier, fuel type, emergency versus scheduled, repair versus replacement — and close with "we provide a detailed estimate after assessing the job." This is not a stylistic preference. Published numbers commoditize the offer and undercut the sales conversation before a technician has seen the home. **CI gate:** the build fails when a dollar amount (`$` followed by a digit, or "N dollars") or a banned positioning word appears in the rendered body of any page. This gate is what keeps the rule alive after the person who wrote it stops reviewing every page. ### M.1 Installation-intent estimator **Locked URLs:** ``` /ac-installation/estimate/ /furnace-installation/estimate/ /heat-pump-installation/estimate/ ``` **What it estimates: equipment, not price.** This is a sizing and tier recommender. It takes home size, system age, efficiency preference and fuel type, and returns a capacity range, an efficiency band, and an equipment tier recommendation, then hands off to a real quote. It never returns a number with a currency symbol in front of it (M.0). That is not a weakened version of the tool. A homeowner searching "cost to replace AC" wants to know what they are buying and what moves the number; the quote itself requires seeing the house, which is also the honest answer and the one that books the appointment. **Architecture: content page first, calculator second.** - At least 300 words of substantive text, server-rendered, visible with JavaScript disabled: why an estimate matters, the factors the tool weighs, a sample table of 2 to 3 common scenarios with their tier recommendations, the CTA, the disclaimer. - The interactive calculator loads as progressive enhancement, after first paint. - A `