# Playhouse Digital - full content > Marketing strategy and research for owner-managed services businesses. Agency thinking, without the agency: Mark Brownlie shows you exactly what your competitors are doing, where you blend in, and what to do about it - then builds the fixes. Research, strategy and execution, joined up by one expert, at a price an owner-managed business can start with. This is the full-content companion to /llms.txt: reference resources and every post's complete text inlined. ## Resources # Ad Assets Cheat Sheet Source: https://backstage.playhouse.digital/resources/ad-assets-cheat-sheet Updated: 2026-08-28 Six creative assets cover almost every placement across Google Ads and LinkedIn. Start here if you are planning a shoot, and jump to your platform if you already know what you are building. ## Six assets cover everything Build these six and you cover every campaign type on both platforms. Stop at four if the campaign is desktop-focused. | # | Asset | Ratio | Unlocks | Mobile only? | |---|-------|-------|---------|:---:| | 1 | Square image | 1:1 | All Google + all LinkedIn image placements | No | | 2 | Landscape image | 1.91:1 | All Google + LinkedIn Single Image | No | | 3 | Horizontal video | 16:9 | All video placements, both platforms | No | | 4 | Square video | 1:1 | Feed-friendly video, both platforms | No | | 5 | Vertical image | 4:5 | Demand Gen + PMax + LinkedIn mobile | Partial | | 6 | Vertical video | 9:16 | Shorts and Stories inventory | Yes | > [!TIP] > **One to four** give solid cross-platform coverage on every device. **Five and six** are mobile-first vertical formats. Skip them if you are chasing desktop conversions. Worth considering at the planning stage: if you shoot the 1.91:1 landscape as your master and keep the subject centred, cropping it to a square is straightforward. Deciding that before the shoot is cheaper than discovering it after. **One 4:5 export covers both platforms.** Export vertical images once at 960 x 1200. That is Google's recommended 4:5. LinkedIn recommends 720 x 900 with headroom to 1080 x 1350 on the same ratio family, so 960 x 1200 sits inside both ranges. No second LinkedIn-only export. ## Where each asset can run You have a file. This says where it goes. | Asset | Build at | Google Search | Google DG | Google DG Carousel | Google PMax | LI Single Image | LI Carousel | LI Video | |---|---|:---:|:---:|:---:|:---:|:---:|:---:|:---:| | **Sq (1:1) image** | 1200 x 1200 | **R** | Y | Y | **R** | Y | **R** | - | | **Land (1.91:1) image** | 1200 x 628 | Y | Y | Y | **R** | Y | - | - | | **Vert (4:5) image** | 960 x 1200 | - | Y | Y | Y | M | - | - | | **Port (9:16) image** | 1080 x 1920 | - | M | M | Y | - | - | - | | **Horiz (16:9) video** | 1920 x 1080 | - | Y | - | Y | - | - | Y | | **Sq (1:1) video** | 1080 x 1080 | - | Y | - | Y | - | - | Y | | **Vert (9:16) video** | 1080 x 1920 | - | M | - | M | - | - | M | | **Vert (4:5) video** | 1080 x 1350 | - | Y | - | - | - | - | Y | > **Key:** R = required · Y = supported (all devices) · M = mobile only · - = not available > **Columns:** DG = Demand Gen · LI = LinkedIn **Build at** is one file per row that every platform in that row accepts. Google and LinkedIn publish different recommended sizes for some ratios, so a single "recommended" figure would be wrong for one of them. Where LinkedIn's recommendation is the smaller of the two, the Google figure still sits inside LinkedIn's own stated maximum, so one export covers both and there is no second LinkedIn-only file. Two rows worth reading twice: - **9:16 images in PMax.** Google publishes no 9:16 spec for Performance Max, but the crop tool offers the ratio. Absence of a spec is not a prohibition. Use the Demand Gen dimensions if you build one. - **4:5 video in Demand Gen.** Google's Demand Gen page lists it at 1080 x 1350. PMax does not take it. LinkedIn's current single-image spec supports three layouts only: horizontal, square and vertical 4:5. The older 2:3 and 1:1.91 ratios are gone. ## Google Ads: images Performance Max and Demand Gen use the same dimensions, and Search shares the ratios it supports with PMax. The last three columns repeat where each ratio can run. | Ratio | Name | Recommended | Format | Max size | PMax | Demand Gen | Search | |-------|------|-----------|--------|----------|:---:|:---:|:---:| | 1:1 | Square | 1200 x 1200 | JPG, PNG | 5MB | **R** | Y | **R** | | 1.91:1 | Landscape | 1200 x 628 | JPG, PNG | 5MB | **R** | Y | Y | | 4:5 | Vertical | 960 x 1200 | JPG, PNG | 5MB | Y | Y | - | | 9:16 | Portrait | 1080 x 1920 | JPG, PNG | 5MB | Y | M | - | > **Key:** R = required · Y = supported (all devices) · M = mobile only · - = not available > [!IMPORTANT] > **Keep the load-bearing content in the centre 80%.** Google's rule: content "needs to be in the center 80% of the image (the safe area that won't be cut off despite device screen size)". Let artwork bleed to the edge, but keep the subject inside. Google separately disallows "images that are cropped in a way that makes the subject of the asset difficult to distinguish", so a clipped subject is a policy problem as well as a design one. **Google auto-crops.** The interface says so on both images and logos: "Selected images may be auto-cropped. You can always edit afterwards." That cuts both ways. You can supply fewer ratios and let Google derive the rest, and a crop you did not choose can cut your subject. The centre 80% rule is the defence. ## Google Ads: video Video runs through YouTube. YouTube's file rules apply rather than an ad-platform cap: 256GB maximum, format MPG (MPEG-2 or MPEG-4). Shoot 1080p. Google recommends Full HD and advises against SD. | Ratio | Name | Recommended | Performance Max | Demand Gen | |-------|------|-----------|:---:|:---:| | 16:9 | Horizontal | 1920 x 1080 | Y | Y | | 1:1 | Square | 1080 x 1080 | Y | Y | | 9:16 | Vertical | 1080 x 1920 | M | M | | 4:5 | Portrait | 1080 x 1350 | - | Y | ### Duration | | Performance Max | Demand Gen | |---|---|---| | Minimum | **10 seconds** | **5 seconds** | | Under 10 seconds | Not accepted | Accepted, but ineligible for YouTube In-stream | | Maximum | Not stated | Not stated | | Shorts eligibility band | 10 to 60 seconds | Not stated | > [!WARNING] > **A six-second Demand Gen video passes upload and loses inventory.** It clears the 5-second floor, then becomes ineligible for YouTube In-stream. Ten seconds is the practical floor in both campaign types: a hard rule in PMax, an inventory constraint in Demand Gen. ### YouTube Shorts There is no Shorts-specific minimum duration. The floor comes from the campaign type, not from Shorts. One asset covers both campaign types. A vertical of 10 to 60 seconds satisfies the PMax band and clears the Demand Gen In-stream threshold. Shorts is not a 9:16-only surface. Vertical is "best suited" for the format, and horizontal and square "are also supported". Prioritised, not required. ### Where Google hosts the video PMax offers a Google-managed House channel you can upload to during campaign setup, instead of linking your own YouTube channel. That matters for a client who does not want ads appearing on their own channel, and it removes the need for channel access before launch. ## Performance Max Landscape and square are both required. Vertical is optional. | Ratio | Status | Minimum | Rule of thumb | Maximum | |---|---|:---:|:---:|:---:| | 1.91:1 landscape | Required | 1 | 4 | 20 | | 1:1 square | Required | 1 | 4 | 20 | | 4:5 vertical | Optional | 1 | 4 | 20 | | 9:16 vertical | Optional | not published | not published | not published | ## Demand Gen At least one landscape or square image is required. One or the other, where PMax needs both. All four image ratios are supported and neither of the other two is mandatory. Carousel accepts the same four image ratios, every card must share one aspect ratio, and carousel takes images only. ### Channel selection Channels are chosen at ad group level, not campaign level. The interface offers **All Google channels**, or **Let me choose**, which limits the ad "to show only on the eligible channels of your choice". | Channel | What it covers | |---|---| | YouTube in-stream | Ads before, during or after YouTube videos | | YouTube in-feed | Ads in YouTube's watch next, home and search feeds | | YouTube Shorts | Ads on YouTube's short-form video platform | | Discover | Personalised content feeds on Google mobile experiences | | Google Mail | Ads within Gmail inboxes | | Google Display Network | Third-party sites and apps | | Maps | Ads on Google Maps | The three YouTube surfaces are independently selectable, so Shorts-only is achievable: choose "Let me choose", keep YouTube then Shorts, and clear everything else. Because the control sits on the ad group, a Shorts-only vertical and a broader multi-ratio set can live in the same campaign in separate ad groups. > [!WARNING] > **The Display Network reshapes creative.** The Display Network is the only channel carrying the warning "some ads may be modified to fit publisher formats". It is the one place Google says it will adapt your creative to a surface, so it is where you have least control over how the ad looks. **Supplying one ratio does not restrict where it runs.** Google's wording: "If you only have one video in the ad, that video will serve on 'all formats.'" A vertical-only ad group is not a Shorts-only ad group, and Google will put that 9:16 on the horizontal surfaces too. Narrowing placement is the channel setting above, never a property of the asset. Google recommends three videos per orientation and marks every video ratio as not required. A single-ratio build is inside the rules and well below what Google suggests. ## Google Ads: logos Google treats logos as separate assets from marketing images. Every Google logo is square except the Performance Max landscape logo. | Where it appears | Ratio | Min | Recommended / max | Format | Max file | |------------------|:---:|-----|-------------------|--------|----------| | Brand Profile (Merchant Center, Shopping, Business Profile) | 1:1 | 500 x 500 | up to 2000 x 2000 | BMP, JPEG, PNG | 5MB | | Performance Max, square | 1:1 | 128 x 128 | 1200 x 1200 | JPG, PNG | 5MB | | Performance Max, landscape | 4:1 | 512 x 128 | 1200 x 300 | JPG, PNG | 5MB | | Demand Gen | 1:1 | 144 x 144 (help) / 128 x 128 (UI) | 1200 x 1200 | JPG, PNG | 150KB (help) / 5120KB (UI) | | Search / business logo | 1:1 | ~128+ | 1200 x 1200 | JPG, PNG | 5MB | > [!WARNING] > **The Demand Gen logo has two different specs.** Google's help page publishes a 150KB cap and a 144 x 144 minimum. The campaign interface states 5120KB and 128 x 128. Unresolved, so treat it as a live disagreement rather than a settled figure. > > **Export at 1200 x 1200 and compress below 150KB.** That satisfies both positions, so you never have to settle it. A logo that passes PMax at 2MB may be refused under the help-page rule. "Only approved logos will appear in your ads." There is a review step between upload and serving. Merchant Center accepts a square logo only. LinkedIn has no separate logo asset, it pulls the company logo from the posting Page. ## What goes in a Google image PMax builds ads at auction time, combining your image with headlines, descriptions, logo and a call-to-action button that vary by surface. Google publishes no creative-content rule specific to Demand Gen, so the guidance below is written for PMax and inherited by Demand Gen. **What works:** - A styled product or lifestyle shot with a clear focal point, strong enough to stand without text. Editorial, not banner. - A single clear headline or value line, when it stands out and reads cleanly. - A tasteful brand logo. It can help, since the surface's own logo placement is not always prominent. - Subtle brand presence, like product packaging that naturally shows the logo. **What fails, or defeats itself:** - A baked-in call to action. Google renders its own button. Never draw "Shop now" or "Click here" into the image. - Full banner-style creative with headline, body copy, CTA and border. - A wall of copy that fights Google's composed text. - An oversized logo or brand watermark. Google recommends supplying at least one image without overlays for each aspect ratio. ### Text and overlays: the Search and PMax split Google's own documentation contradicts itself here. Three official pages take three positions. | Source | Position | |---|---| | Image assets format requirements (policy) | **Bans them.** Images with text or graphic overlay, brand logos included, are not allowed | | About image assets for Performance Max | **Permits them.** Overlays are allowed, and at least one image without overlays per ratio is recommended | | About image assets for Search campaigns | **Absolute prohibition.** "there are absolutely no digitally added text, logos, or graphic overlays, as these will be immediately disapproved" | What holds in practice: - **Search is the only campaign type where the stated rule is real and enforced.** Any digitally added text, logo or overlay gets disapproved. A Search image asset has to be a clean scene and nothing else. - **The general policy ban does not hold for PMax.** Brand logos inside PMax image assets run in live accounts without issue. - **PMax accepts image plus embedded text.** Where the policy page and the PMax page disagree, the PMax page governs. ## LinkedIn: specs **Images** (LinkedIn's published recommendations) | Ratio | Name | Recommended | Notes | |-------|------|-----------|-------| | 1.91:1 | Landscape | 1200 x 628 | Min 640 x 360 | | 1:1 | Square | 1200 x 1200 | Min 360 x 360 | | 4:5 | Vertical | 720 x 900 | Mobile only; export at the shared 960 x 1200 | **Video** | Ratio | Name | Recommended | Notes | |-------|------|-----------|-------| | 16:9 | Horizontal | 1920 x 1080 | Up to 1920 x 1080 | | 1:1 | Square | 1080 x 1080 | Min 360, max 1920 | | 4:5 | Vertical | 720 x 900 | Up to 1080 x 1350 | | 9:16 | Vertical | 720 x 1280 | Up to 1080 x 1920 | > Format: MP4 (H.264 or VP8). File size: 75KB to 500MB. Duration: 3s to 30min; LinkedIn recommends 15 to 30 seconds. Vertical 4:5 images are mobile only and do not deliver to desktop. 9:16 video is mobile only. Carousel recommends square images and scales them to 312 x 312 on display, with 2 to 10 cards. **Audience targeting:** Job Title must not be combined with Seniority or Job Function. LinkedIn rejects that combination outright, and it is a rejection rather than an audience-size collapse. Title combined with Skills is allowed, and is the better move when a skill alone is too broad. ## LinkedIn: ad copy and creative > [!IMPORTANT] > The first 150 characters of intro text are everything. Front-load a hook that reads as a finished thought before the mobile "see more" cut. If it does not land, nobody expands it. | Element | Target | Maximum | Notes | |---------|--------|---------|-------| | Intro text | First 150 chars = a complete sentence | 3,000 chars | Mobile truncates at ~150 chars. After the fold, write as much as you like | | Headline | 70 chars or fewer | 200 chars | 70 is where LinkedIn truncates. Under ~40 keeps it on one line | | Description | 100 chars | 300 chars | Optional and rarely shown. Skip it | - The headline sits below the image. Keep it punchy and direct. - No CTA in the copy. LinkedIn renders its own button (Download, Learn More, Sign Up, Attend) beside the headline. Do not spend headline or intro space on "Download now". - LinkedIn expects a professional tone, not corporate waffle. Specific outcomes, real numbers and peer-level language win. **Why five ads per ad group** LinkedIn shows a person at most one impression per ad per day. Five ads in one ad group means up to five impressions per person per day, one from each ad. The point is frequency and reach, not testing, though you get testing as a side benefit. Each ad only has to be technically different, so a single word change counts. Fewer than five leaves daily touchpoints on the table. **What belongs in a LinkedIn image** LinkedIn creative is image-led, and this is the opposite of the Google rule above. The headline and description fields sit below the image and are low-prominence, often needing a click to expand, so the image does the persuading. - Bake the headline or pull-quote into the image. On-brand type is encouraged here, not penalised. - No call to action in the image. LinkedIn renders its own button, so do not draw a "Sign up" button into the creative. - Logo is optional. The ad already shows your company logo and name as the post author, so a second logo inside the image is usually redundant. ## Official documentation **Google Ads** - [Demand Gen asset specs](https://support.google.com/google-ads/answer/13704860) - [Demand Gen video asset specs](https://support.google.com/google-ads/answer/17141078) - [Demand Gen specs and format requirements](https://support.google.com/google-ads/answer/17091672) - [Creative preferences in Demand Gen campaigns](https://support.google.com/google-ads/answer/14552284) - [About Demand Gen campaigns](https://support.google.com/google-ads/answer/13695777) - [YouTube Shorts ads asset specs](https://support.google.com/google-ads/answer/16041697) - [How to buy YouTube Shorts ads with Demand Gen](https://support.google.com/google-ads/answer/16040528) - [PMax image assets](https://support.google.com/google-ads/answer/14530211) - [PMax video assets](https://support.google.com/google-ads/answer/14528532) - [Best practices for image assets in PMax](https://support.google.com/google-ads/answer/15995550) - [Search image assets](https://support.google.com/google-ads/answer/9566341) - [Image assets format requirements (policy)](https://support.google.com/adspolicy/answer/10347108) - [Best practices guide for responsive display ads](https://support.google.com/google-ads/answer/9823397) - [Ad formats and sizes](https://support.google.com/google-ads/answer/13676244) - [Brand Profile logo requirements](https://support.google.com/brandprofile/answer/16785039) **LinkedIn Ads** - [Single Image Ads specs](https://www.linkedin.com/help/lms/answer/a426534) - [Carousel Ads specs](https://www.linkedin.com/help/lms/answer/a427022) - [Video Ads specs](https://www.linkedin.com/help/lms/answer/a424737) --- # Capture, Distil, Decide Source: https://backstage.playhouse.digital/resources/capture-distil-decide Updated: 2026-07-27 Every piece of marketing work rests on the same handful of facts about your business: who you are, who buys from you, what you're up against. Those facts exist. They scatter the moment a project ends, so the next one goes looking for them all over again. This article is about a way of working that stops that: capture what's known about your business once, keep it in one place that stays current, and have every project read from it and add to it. The whole of it fits in a single sentence. > [!IMPORTANT] The whole method, in one sentence > Capture your site into the lake, distil the lake into the foundation, and turn the foundation into a plan. Three layers, each made from the one before: - **The lake** - what we captured: every page, yours and your rivals', dated. - **The foundation** - what we know: how your brand looks and sounds, your offer, your audience, your proof, your competitor set - every fact sourced. - **The plan** - what to do: the judgment calls and next steps, the only layer that ever reaches you. The foundation isn't filled in once and left. It grows out of the work itself: the brand reading, the competitor reports, the deep dives each leave it knowing more than they found. That's what makes the second project cheaper than the first and the fifth sharper than the fourth. ![The method in three actions, left to right: the websites, yours and your rivals', are captured into the lake, what we captured; the lake is distilled into the foundation, what we know - identity, offer, audience, proof, competitor set, every fact sourced and dated; the foundation is decided into the plan, what to do. Under each layer, how it fails: lose the lake and you keep what you know but can never prove it; lose the foundation and you rebuild it from the lake; lose the plan and the knowledge survives but nothing ships - the only layer that reaches you](/diagrams/engine-three-layers.svg) ## The knowledge that goes missing Whenever someone does marketing work for you, the deliverable is visible and the groundwork behind it isn't. Every new piece of work - a report, a landing page, a campaign, a website - needs the same facts about your business to hand. Some are written down in an old proposal or a shared drive; some are in browser tabs and someone's memory of a call. Rarely is any of it in one place, current, and complete. So each project spends its first hours gathering the pieces it needs again, rebuilding whatever the last one didn't leave behind. Two things make that expensive. It's lossy - the piece that lived only in one person's head is gone the moment they move on. And it's siloed - the person writing your landing page and the person running your ads are each holding a different fragment, and no one is looking at all of it together. Nobody puts that cost on an invoice, but it sets the ceiling on how fast anything can ship, because the gathering always runs before the work. ![Two panels: without a foundation, every project gathers the same scattered facts, delivers, and lets them disperse again; with a foundation, projects stack on top of a captured base, each starting where the last finished and writing back what it learned](/diagrams/engine-ch1-two-loops.svg) ### What the fix has to do The right-hand side of that diagram is the promise: projects stacking on a base that already knows your business, each one starting where the last finished. A shared drive full of documents doesn't get you there, because nothing in it has to stay true, so it drifts. The fix has to do two jobs at once: hold what's known in a form every project actually reads, and take back what every project learns. The arrows pointing down into the base matter as much as the base itself. If you want to know where your own business stands, ask the last three people who did work for you what they worked from. Whatever comes back, and whatever doesn't, is your foundation as it exists today. Doing both jobs properly needs three distinct layers, one for each step in the opening sentence. The first holds the raw material. ## A page is not a scrape The first layer is **the lake**: every captured page, yours and your rivals', held in one place with a date on each. The name comes with a test attached: after capture, nothing downstream ever needs the live site again, because everything after it reads the lake. So what does "captured" have to mean for that to hold? When people say they've captured a website, they usually mean one of two things: a scrape, which rips the words out and throws away everything about how they were arranged, or screenshots, which keep the arrangement and throw away the machine-readable words. Each answers exactly one kind of question. The lake has to answer questions nobody has asked yet - so it needs both, plus one more thing neither provides. ### Three files In my system, capturing a page produces three files: - **What it says** - the words, as clean readable text - **What it is** - the structure: navigation, headings, links, metadata, the skeleton a rebuild needs - **What it looks like** - a full screenshot, because some truths only exist visually ![Flow diagram: one page, visited once, becomes three files - text, structure, screenshot - landing in the lake alongside competitor pages, with a per-page ledger of when it was captured and when it actually changed; brand capture, intelligence reports and the site build all read from the lake and none revisit the website](/diagrams/engine-ch2-envelope.svg) Together, the three files are the page as it stood on the day: what it said, how it was built, how it looked. Enough for a later project to read its words, rebuild its structure or study its design without going back to the site. Not everything a website does lives in those three files - how fast it loads, how a form behaves, whatever sits behind a login - so anything in that class is captured on purpose or not at all. Here are those three files for one real page, my own About page, as it was captured on 26 June 2026: ``` playhouse-digital__about-about-us.md 5.1 kilobytes what it says playhouse-digital__about-about-us.json 5.7 kilobytes what it is playhouse-digital__about-about-us.png 412.8 kilobytes what it looks like ``` The first opens straight into readable text, no markup to strip: ``` # About “Us” I bring value to supporting businesses of all sizes that market directly to consumers where success is measured by generating a lead such as a form fill or a phone call from their website... ``` The second holds the same page as structure and nothing else: its canonical address, its social metadata, one heading, three navigation blocks, twenty internal links, twenty-nine calls to action, and three structured-data declarations. No words, no design. The third is the screenshot, and at more than seventy times the size of either text file it carries what neither of them can. ### The ledger Each page in the lake also carries a small ledger: when it was first captured, when it was last captured, and when its content actually changed. Those are different facts. A page re-checked yesterday that hasn't changed since March tells you something a scrape never could: the evidence is current *and* stable. It's what lets the layers above stay trustworthy without endlessly re-reading everything - and it means every claim built on the lake can carry a date. "Their homepage said this, as of this day" is evidence. "I remember their homepage saying this" is not. > [!IMPORTANT] Any site. Yours - or your biggest rival's > The capture doesn't know whose site it's reading. The same machinery that puts *your* website in the lake puts *any* website there. Point it at your own site and you get the base everything in this article builds on. Point it at your biggest rival and you get their words, their structure, their look - dated, complete, and sitting in the lake one folder away from yours. So your pages and your competitors' pages live side by side, in the same three-file format. It sounds like a filing detail, and it decides every comparison downstream: messaging gaps, design benchmarks, positioning all read like-for-like captured evidence rather than one fresh site against one remembered one. That is what the test at the top of this chapter buys. Capture a site once and every later piece of work reads the capture instead of the site, which is why one pass over a site keeps paying out for weeks. The lake holds everything and knows nothing. Turning what it holds into knowledge starts with the most visible thing a website carries: the brand. ## Reading the brand: how it looks and how it sounds Your design language is already on your website: every choice of typeface, colour, spacing and shape. What it has probably never been is collected in one place. It lives as an accumulation of decisions, held together by whoever built the site and your eye for "that looks right." The same is true of your voice. How your business talks - the words it reaches for, the promises it makes, how formal or plain it is with a buyer - is just as real, and wherever it has been written down, it is usually in the copy itself rather than in a guide anyone could hand to a new writer. Both halves of the identity are sitting in the lake, waiting to be read. One is read from the styles; the other from the words. ### The look: read from the styles The design half rests on one fact: to draw a page at all, a browser has to compute the final, resolved value of every style on every element. That resolved value is the result, whatever anyone intended. So the design language isn't lost. It's fully specified, machine-readable, and sitting in plain sight on the rendered page. Reading those computed styles off a page yields thousands of raw values. Distilling them down produces a small set of **role-assigned tokens**: this typeface is the display face, this one the body, this colour is the accent, this is the surface everything sits on, these are the shapes. Alongside the tokens comes a design document - the brand's rules, written down for the first time. ![Flow diagram: a rendered page yields thousands of computed style values; distillation produces an identity file of role-assigned tokens, each graded exact, derived or defaulted, plus a readable design document; below, the same process pointed at a rival's site makes brand comparison a comparison of two files](/diagrams/engine-ch3-codifier.svg) And every token carries an honesty grade: **exact** (read directly off the page), **derived** (computed from what was read), or **defaulted** (nothing to read, filled from convention). The grade is what separates capturing a brand from guessing at one. Role by role, you can see how much of the file was read off your site and how much was filled in for you. ### The voice: read from the words The other half of the reading works on the text in those same files. Every page you've published is a sample of how your business speaks: what it leads with, what it calls your customers, whether it says "we help" or "solutions are delivered". Distilled the same way, that becomes a written tone of voice - the words your brand uses, the ones it never would, and how it talks to the person you're trying to win. And because the rivals' pages sit in the lake in the same format, the same reading runs across the whole market. How does everyone else in your space talk? Where do the voices blur into one interchangeable hum, and where would a different way of speaking stand out? Every line of that answer traces to something a competitor actually published, dated and on file. ### The trust layer nobody could measure In one study, people shown a page for 50 milliseconds rated its visual appeal much as they did when shown it for ten times as long ([Lindgaard et al., 2006](https://www.tandfonline.com/doi/abs/10.1080/01449290500330448)). That is too fast to be reading, and too fast to be reasoning. Whatever a visitor is judging that quickly, it isn't your copy. The look of the page has already landed. That has been difficult to act on, because nobody could say what was wrong. You could feel that a site was "off"; you couldn't say by how much, or exactly where. > [!IMPORTANT] A look you can measure > Once a brand is a file, that changes. It can be measured, compared, versioned and handed over. The part of your business that was hardest to pin down becomes the part that travels best. ### Point it at your rival The lake chapter said the capture doesn't know whose site it's reading. Neither does this step - and that's where it stops being an efficiency and becomes an advantage. Run the same reading over your biggest rival's site and you get *their* identity as a file: their type scale, their palette, their level of craft, and their voice - graded with the same honesty. "How does our brand hold up against theirs?" is usually answered with feelings. Now it's a comparison of two files - role by role, measurably, with the evidence attached. A brand nobody has written down is a brand nobody can hold to account. Yours is a file now, and so is theirs. Where both files go next is the second layer. ## The foundation: what we know Ask what any marketing project needs to know about your business and it comes down to the same five objects: the identity (how your brand looks and sounds), the offer, the audience, the proof, and the competitor set. In most businesses all five exist - scattered across formats and folders, each version written for one project and none maintained. The second layer, **the foundation**, is those five objects held in one structured place: current, maintained, and with every fact carrying its source and its date. "Their guarantee changed in March" is a foundation fact, and the captured page that proves it is sitting in the lake underneath. ### Built by the work, not before it Nobody sits down and writes the foundation. It comes out of the work. The brand reading in the last chapter deposits the identity. A competitor report deposits the competitor set and the unclaimed positions. A deep dive deposits what one rival is really doing. The foundation is the accumulation of every reading and every report that's ever run, which means it's never a form someone filled in optimistically on day one and forgot. Every fact in it is there because a piece of real work needed it, found it, and left it behind. That also means it starts faster than a documentation project would. The first capture opens it, the brand reading fills in the identity, and everything after that adds to it. The five objects are the same five in every business, whatever it sells, so the shape doesn't change from one to the next. ### Reading replaces re-learning The first project pays to start the foundation. Every project after it begins by reading what's already there. The second competitor report doesn't begin at zero; it begins where the first one finished, because the identity, audience and competitor set are already sitting there, current and sourced. What makes this practical now, rather than a nice theory, is that an AI can read the entire foundation before every task and write what it learned back afterwards. Reading everything known about a business before every single deliverable was never worth a person's time, which is why nobody did it. Handing that reading to a machine is what changed. A quick test for your own business: name the five objects, then see whether you could put your hands on a current version of each one in the next ten minutes. The ones you can't find are the ones your next supplier will rebuild from scratch and bill you for. > [!IMPORTANT] That's the foundation > One place where everything known about your business accumulates, sourced, dated and current. Nothing in it was written to impress anybody. Every line is there because a piece of real work needed it. The three layers earn their separate names because they fail in three different ways. Lose the lake and you still know everything; you just can't prove any of it, or date a claim. Lose the foundation and nothing is gone for good: it rebuilds from the lake, because it's derived, never primary. Lose the plan and the knowledge survives, but nothing ships. Holding the three apart by name is the working discipline behind everything else. Every piece of work can say which layer it reads and which it writes: the brand reading reads the lake and writes the foundation; a competitor report reads the foundation and feeds the plan. That one habit keeps evidence, knowledge and judgment from blurring into each other as the system grows. It's a question worth asking of anything you get sent: which of the three is this? Evidence, knowledge, or a decision. Work that can't answer is usually none of them. ## The branches that grow the foundation With a foundation started, intelligence work stops being a series of one-off projects and becomes a set of **branches**. Each branch reads what's already known, asks one question of the evidence, and writes its answer back. The competitor report asks where the messaging gaps and unclaimed positions are, across your whole rival set. The deep dive asks what one rival is doing, forensically. The design benchmark asks how your brand's craft measures up, file against file. The performance analysis asks where your spend is leaking. Each asks a different question; all of them ask it of the same evidence. ![Diagram: your foundation at the centre; four branches around it - competitor report, deep dive, design benchmark, performance analysis - each reading from the foundation in grey and writing findings back in gold; everything to the left culminates on the right in the plan: the assembled view of decisions, not data, and the reasoning behind them kept as long-term memory](/diagrams/engine-ch6-branches.svg) ### Every deliverable is also a deposit Reading is only half of what a branch does. **It writes back what it learned.** When a competitor report finishes, its findings about each rival file into the foundation as dated, sourced knowledge. Run the report again next quarter and it doesn't start from a blank page; it starts from everything the last run established, and spends its effort on what changed. Without the write-back, intelligence evaporates into PDFs. With it, every deliverable grows the foundation it was built from, and the fifth question you ask costs a fraction of the first, because everything before it has already paid in. ### When the intelligence was wrong The system has shipped wrong findings. One report read a routine lead time as a scarcity signal, and called a proof gap that turned out to be sitting one click from the homepage. Human review caught both, and the report was rebuilt from verified capture and re-issued. The fixed report was the small outcome. The lasting one was that the pipeline gained the checks that catch that class of error, and the corrections went back into the foundation with everything else. So when you read a finding, the useful question is what would have happened if it were wrong. Somebody noticing is a weaker answer than something checking. Every branch ends the same way: with answers about what's wrong, what's missing, what's unclaimed. Answers on their own don't run a business. ## The plan: decisions, not data **The plan** is what you actually receive: judgment calls, to-dos and next steps, benchmarked against your named rivals, assembled across every branch into one narrative with a single prioritised action list. > [!IMPORTANT] What you're actually getting > Nobody needs a stack of findings. What you need is knowing what to do on Monday. Getting there is assembly and judgment rather than more analysis: every branch already ends in answers, so the plan's job is to weigh them against each other and put them in order. ![Diagram: four intelligence branches - competitor report, deep dive, design benchmark, performance analysis - feed one gold document, the plan: judgment calls on what to claim and what to drop, to-dos in priority order, every line benchmarked against named rivals, the reasoning kept for the decision after; a set of decisions, not a stack of reports, ending in what to do on Monday](/diagrams/engine-the-plan.svg) Of the three layers, the plan is the only one that ever reaches you. The lake is evidence and the foundation is memory; both exist so the plan is right, and so the next one is sharper. It's a useful filter to apply to anything you're sent: if a page of it doesn't change a decision, it's trivia, however interesting the data behind it. ### The reasoning travels with the decision What makes the plan compound is that it keeps the decision and the reasoning together: the evidence behind it, and what was ruled out. Come back two quarters later and the question "why did we drop that channel?" has an answer on file, so the next decision starts from the last one instead of from scratch. Decisions with their reasoning attached become long-term memory; decisions on their own are just this quarter's opinions. A plan still doesn't move a business until something ships against it, and shipping has a discipline of its own. ## From diagnosis to treatment Every line in a plan splits the same way: a **diagnosis** - what the intelligence says is wrong, missing or unclaimed - and a **treatment** - the artifact that acts on it. The design read pairs with a rebuilt page. The competitor gap analysis pairs with positioning, which feeds generated landing pages. The search-term analysis pairs with restructured campaigns. > [!IMPORTANT] Diagnosis and treatment need each other > Diagnosis without treatment is trivia; treatment without diagnosis is guesswork with good production values. Capture, distil and decide stop at the plan: they are how knowledge gets made. Treatment is how work gets made. The pairing also means you can see why a piece of work is being proposed before you agree to it. The diagnosis names the problem and shows the evidence behind it; the treatment is the answer to that problem and nothing else. ![Two-column diagram: diagnoses on the left - the design read, the competitor gap analysis, the search-term analysis - each with an arrow to its treatment on the right - the rebuilt page, positioning feeding generated pages, restructured campaigns; a gold gate on the middle arrow marks that vague positioning is refused, and a return arrow along the bottom notes that what ships writes back](/diagrams/engine-ch7-diagnosis-treatment.svg) ### Rebuilding the page properly When the treatment is a website, new colours on the old structure won't do it. The reading that captured your brand becomes the way into rebuilding the page itself: the structure, the order the argument unfolds in, the journey a visitor takes from arriving to enquiring - designed for the audience the foundation identified, saying what the voice reading says will land, looking the way the identity file says it should look. > [!IMPORTANT] All of that puts a condition on the work > A page built this way has to know what it's arguing before anyone writes a line of it. ### Variants, scored against each other One more discipline when the treatment is a page: never generate one and hope. The system generates structurally different variants of the same page - different arguments, different structures, same locked positioning - and each one is scored independently, out of ten, on one dimension at a time. Here is a real scoreboard from a run on my own landing page. The same proposition was built three ways: one written from the facts alone with no brand styling, one with the same facts inside the brand, and one with every structural module in place. | Dimension (out of 10) | Facts only | On-brand | Full build | Winner | |---|---|---|---|---| | Comprehension | 6 | 8 | 9 | Full build | | Relevance | 5 | 5 | 9 | Full build | | Differentiation | 3 | 4 | 8 | Full build | | Trust | 9 | 9 | 4 | On-brand | | Craft | 8 | 8 | 5 | On-brand | | Design | 7 | 8 | 5 | On-brand | Read down the columns and the split is hard to miss. The version with the best message had the worst proof; the version with the best proof had the least to say. Neither of them shipped. The merged page took its hero and its differentiation from the full build, and its proof, its craft and its whole visual frame from the on-brand one. The facts-only version won nothing at all, which is still a result: it set the proof discipline the winner had to match. Lifting the winning copy also meant losing some of it. A claim about how many reports had been delivered went, because at the time the honest count was one, and so did a competitor price comparison with no source behind it. ### One identity, every surface None of that consistency is maintained by hand. Because the identity is data, every surface the work produces for you wears it automatically: the rebuilt pages, the intelligence reports, the documents. Your brand is defined in one place and read from there by everything, so nothing drifts out of sync, and a brand change ripples everywhere at once. You can check that claim against the page you're reading. Here are five lines out of the identity file this site is built from, and where each one is on screen right now: | Role | Value | Where it is on this page | |---|---|---| | `primary` | `#16181D` | the chapter headings | | `body` | `#565B62` | the text you're reading | | `accentAccessible` | `#8B6B1A` | the gold labels above each highlighted box | | `binding` | `#0B2545` | the deep panels heading tables in the reports | | heading typeface | Fraunces | every heading you've read so far | Twenty values in total: twelve colours, two typefaces, one corner radius, five spacing rules. Swap that one file and the same page renders in someone else's brand instead, with nothing else touched. Change a value inside it and every surface follows, which is the whole point of keeping a brand as data rather than as a PDF nobody opens. ![Hub and spoke diagram: the identity file - about twenty decisions - at the centre, with one gold arrow in labelled edit here, once, and arrows out to your website, intelligence reports, landing pages, and documents, all sitting on a slab of house craft that never varies by brand](/diagrams/engine-ch4-one-file.svg) That invites a fair worry. If one small file is the only thing standing between your surfaces and anyone else's, how much of what you get is actually yours? ## The gates What's shared is the process: the capture, the readings, the reports, the checks. What comes out of it is built from your lake and your foundation, so the insight and the output are yours alone. A template starts from the same answers every time. This starts from the same questions, and your website, your rivals and your market make the answers different every time. What makes that shared process safe to run at speed is the last piece of machinery: the gates. A system that can produce a report, a page or a whole site in minutes can produce a wrong one just as fast, so speed is only worth having if you can trust what comes out. That trust is built in as gates: automatic checks on every path to "shipped" that block rather than advise. Nothing reaches you that hasn't passed them. ![Flow diagram: generated work - pages, reports, sites - passes through a wall of three blocking gates: the readability floor, the completeness guard, and configuration traps; passing work ships, and any red gate stops the build entirely](/diagrams/engine-ch5-gates.svg) Two of the three gates check ordinary hygiene, and they should. A **readability floor**: every text-and-background pairing has to clear the AA contrast level of the Web Content Accessibility Guidelines. A **completeness guard**: every page has to carry its machine-readable mirror, its structured data and its place in the sitemap. Any competent build clears both of those. That second gate runs on this site, and what a machine reader actually does with those pieces once they exist is the subject of [Your Next Visitor Is a Machine](/resources/your-next-visitor-is-a-machine). What differs here is what happens when one fails: the build stops and names the gap, instead of logging a warning into a report nobody opens. The third is less ordinary. A site refuses to produce a production build while its enquiry form still points at an interim inbox, so the mistake that's easy to make and expensive to discover live is simply not buildable. Blocking is what makes them worth having. A gate that warns gets ignored on a busy day; these stop the line. ### A gate before anything is built The same discipline sits earlier than the build. On the diagnosis-to-treatment path, one arrow carries a gate of its own: the page generator **refuses to run on vague positioning**. If the positioning can't say plainly who the page is for, what's being claimed, and why it beats the alternatives, no page gets generated. > [!IMPORTANT] A beautiful page built on mush is still mush > It just fails more expensively, in front of paid traffic. That refusal encodes something most marketing production gets backwards. The bottleneck in good pages sits upstream of the writing and the design, in the thinking. So the thinking becomes a hard dependency - you can't skip the diagnosis to get to the treatment faster. The build gates stop bad work shipping; this one stops it starting. ### Green ticks lie Gates have a blind spot, and it's worth knowing about even if you never go near one. A check that silently fails to run reads exactly like a check that passed. A build can be green because everything works, or green because the thing that should have failed never ran at all. Two habits guard against that. **Count, don't just exit.** A photo gallery built to skip any broken entry can pass its checks with no photos in it at all. Only a count of how many images actually rendered proves the green is real. Wherever absence can pass silently, the gates count. **Measure the artifact.** One build here looked right on every page and passed every audit while quietly shipping every photo at several times the resolution it needed. Nothing caught it until someone listed the built files by size and asked why they were so big. Looking at what was actually produced, rather than at the ticks, is the habit that finds that class of problem. ### Gates that know their limits When a *rival's* captured brand fails the readability floor, the system warns instead of failing - their identity is what it is, and the capture's job is fidelity. The blocking floor applies to what ships to you. A verification system that can't tell "this is our standard" from "this is their reality" would fail at both jobs. For your own site, the question to ask whoever builds it is what happens when a check fails. If the answer is a warning in a report, the check is advisory, and advisory checks get skipped on the day everyone is busy. That's what the gates buy: a process that repeats for any business, at any speed, without the output ever being less than checked, and without it ever being anyone else's. ## The loops Run end to end, on your business, it goes like this. The **capture** comes first: every page of your existing site into the lake as three files, your rivals' alongside them. The **brand is read out of the capture** - your identity as a file, graded for honesty, down to the difference between the shade you display and the shade you can set text in. The **foundation grows from there** - identity, offer, audience, proof, competitor set, every fact carrying its source. The **branches** run, each asking its one question of the same evidence and writing the answer back, so the second run starts where the first finished. Their answers assemble into **the plan**: what to claim, what to fix, what to build, in order. And the **treatment** ships against it, through the gates, on the identity the capture read - with the files yours at the end of it. ![Flow diagram: your website is captured once into a lake of pages as three files each; the lake is distilled into the foundation - identity, offer, audience, proof, competitor set; intelligence branches read the foundation and write back what they learn; their answers assemble into the plan - decisions, not data; everything renders from the same foundation and ships as public proof, and the loop begins again](/diagrams/engine-master.svg) Each stage's output is the next stage's input. Nothing gets re-learned at any point along the way. ### The loops inside Inside that end-to-end run, each branch is a loop of its own. It reads the foundation, asks its question, and writes the answer back, so the next run of that branch starts where the last finished rather than from a blank page. The same write-back applies after something ships. Performance data, responses, your own reactions all return to the foundation, so the next diagnosis starts sharper than the last. Every treatment that goes out makes the one after it better informed. What leaves writes back. ### The system grew while it ran Some of the gates in the last chapter exist because a real build needed them. The trap that refuses a production build with the wrong form inbox was written the day that mistake became possible. Every engagement leaves the system with better defences than it found. > [!IMPORTANT] Why the checks keep changing > Processes that never change are processes nobody's checking. ### Your next step If you're wondering what any of this means for you on Monday, it starts with one small step: the capture. An afternoon, no meetings, nothing needed from you but the website you already have. Everything else compounds from there. --- # Your Next Visitor Is a Machine Source: https://backstage.playhouse.digital/resources/your-next-visitor-is-a-machine Updated: 2026-07-23 When a buyer asks an AI assistant "who should I use for this?", the assistant doesn't take anyone's word for it. It goes and reads - your site, your rivals' sites, whatever it can reach - and assembles an answer from what it found. The buyer never sees your homepage. They see the answer, with you in it or not. That reading is happening now, on your site, whether or not you've planned for it. A growing share of the "visitors" in your traffic aren't people at all; they're machines reading on a person's behalf. And they are the least forgiving audience you have: a human visitor squints past a broken layout, a machine reader takes what it can parse and moves on. This article covers what serving that reader means in practice: every technique in it is running on this site, one click away, and each chapter shows you where to look. The list beside this article is the chapter map. ## The reader you can't see The old shape of being found online: a person searches, sees a list of links, clicks yours, and lands on your homepage - where your design, your photos and your words get their chance. The new shape inserts a reader between you and the buyer. The person asks a question; an assistant reads the candidate sites; the person gets a written answer. Your site's chance to persuade happens inside that machine's reading, or never. ![Diagram: a buyer asks an assistant, which reads your site and two rivals' sites on their behalf and assembles the answer the buyer sees; one rival's site needs scripts the machine can't run, reads as blank, and is missing from the answer; the buyer then visits the winning site in person - today still to make the purchase themselves](/diagrams/machine-ch1-two-readers.svg) A large share of the web is invisible to that reader. Most modern sites lean on scripts: the page arrives half-empty and JavaScript fills in the content after it loads. Human browsers run those scripts without you noticing. The machine readers mostly don't: when Vercel and MERJ [tracked the major AI crawlers](https://vercel.com/blog/the-rise-of-the-ai-crawler) across months of live traffic in late 2024, the readers from OpenAI, Anthropic and Perplexity executed no JavaScript at all - of the big names, only Google's and Apple's readers render a page the way a browser does. Point one of the others at a script-dependent site and it reads the part that arrived in the first response: sometimes a header and a cookie notice, sometimes nothing. Absence is the whole penalty: the site simply never enters the answer, and nobody is told. The buyer asked a real question with money behind it. The assistant answered from what it could read. A business that would have been the right answer wasn't in the reading. A worked example of the difference: every page on this site arrives as complete text in the first response - no script needs to run before the words exist. That single property moves a site from the unreadable bucket to the readable one, and it was a build decision, not an optimisation bolted on later. The rest of this article is the layers on top: giving the reader a clean text version, a map of the site, and pages that say what they are - then an automatic check that keeps all of it true as the site grows. The human step hasn't disappeared - it has moved to the end of the chain. The buyer reads the answer, picks the winner, and then visits that site in person: to look them in the eye, and to buy. Today that final step still belongs to the person; the direction of travel is visible, though, with assistants that can complete a booking or a purchase already arriving, and buyers will hand that step over exactly as fast as they come to trust it. Which reframes what your website is for. The machine's reading decides whether you make the shortlist; the human's visit decides whether you win it - and everything about serving that human visitor well, the design, the navigation, the readability, is its own subject for its own article. Before the chapters walk through them one by one, here is the full set in one place - every machine-readable surface this site ships, live right now. Each link opens the artifact itself, in a new tab: | Surface | What it is | |---|---| | [This article, plain text](/resources/your-next-visitor-is-a-machine.md) | The `.md` twin of this page: the full content with the design stripped away (chapter 2) | | [llms.txt](/llms.txt) | The machine map: what this business is and where every page lives, one line each (chapter 3) | | [llms-full.txt](/llms-full.txt) | The whole site's text in a single file, for readers that digest rather than skim (chapter 3) | | [sitemap.xml](/sitemap.xml) | The classic crawler map, listing every indexable page (chapter 3) | | [rss.xml](/rss.xml) | The feed that announces new posts to anything subscribed (chapter 3) | | Structured data | Not a separate file: it lives inside each page's HTML, declaring what the page is (chapter 4) | > [!IMPORTANT] One reader between you and the buyer > You don't get to choose whether AI systems read your site on buyers' behalf - that's already happening. You only get to choose what they find: a site that hands over its content plainly, or one that reads as a blank page and quietly falls out of the answer the buyer sees. Next: the simplest of the machine surfaces - a plain-text twin of every page. ## The plain-text twin A web page is words wrapped in design. The wrapping is for people: layout, type, colour, photographs, the things the last article on this site is about. A machine reader wants the words and has to dig them out of the wrapping - navigation repeated on every page, cookie banners, decorative markup, the scaffolding that makes a page look right and read badly. So this site gives the machine its own version. Every page here has a plain-text twin at the same address with `.md` on the end: the full content, in reading order, headings marked, links intact - and nothing else. No layout, no scripts, no design to trip over. You're reading the styled version of this article now; [here is this article's plain-text twin](/resources/your-next-visitor-is-a-machine.md). Same words, built for a different reader. ![Diagram: the same page twice - the styled version with layout and images for human eyes, and its plain-text twin at the same address plus .md, just the words in order; a gold arrow marks the twin as the version the machine reader takes](/diagrams/machine-ch2-twin.svg) The twin does two jobs. The obvious one is reliability: a reader that gets clean text can't mis-parse the design, because there's no design to mis-parse. The quieter one is cost. Machine readers work within limits on how much they can take in per request; a styled page spends most of that allowance on wrapping. The twin spends nearly all of it on your actual content, which means more of what you wrote survives into what the machine knows about you. Every page on this site carries one - articles included - and each styled page announces its twin in its metadata, so a reader that lands on the styled version is told where the clean one lives. The signpost is one line, sitting in this very page's head right now: ```html ``` One line per page, generated automatically, and no machine reader ever has to guess whether a clean version exists or where to find it. What should you take from this if your site is a normal site? Not that you need this exact mechanism tomorrow. The transferable point is the principle: the words are the product, and there should be a version of your site where the words are available without a fight. If your content only exists inside a page builder's output, wrapped in six layers of styling markup and assembled by scripts, then the words you paid to have written are, for this new kind of reader, barely there. > [!IMPORTANT] Words a machine can carry away > A machine reader takes your words and leaves your design behind. A clean text version of every page - full content, no wrapping - is the difference between a reader that quotes you accurately and one that reconstructs you from fragments. Next: a reader that can read every page still needs to know which pages exist. The site should hand over the map. ## Handing over the map Reading one page well solves half the problem. The other half: which pages exist, and which matter? A machine reader arriving at a site it's never seen has to crawl - follow links, guess importance, hope the navigation reflects the structure. Crawling is expensive and patchy, and readers doing it at scale spend as little time per site as they can get away with. A site that's easy to map gets read more completely than one that has to be explored. The web's answer is a small convention called `llms.txt`: one file at a fixed address that tells an AI reader what the site is and where its content lives. Ours is [live at /llms.txt](/llms.txt). It opens with a plain-language summary of what this business does, then lists every page with a one-line description and a pointer to its plain-text twin. A reader that fetches that one file knows the whole shape of the site - what exists, what each thing is, where the clean version lives - without crawling anything. ![Diagram: a machine reader arrives at llms.txt, an index listing every page with a one-line description, each pointing to its plain-text twin; beside it, llms-full.txt carries the whole site's text in one request](/diagrams/machine-ch3-map.svg) Alongside it sits the fuller version, `llms-full.txt`: the same map with the actual content inlined - the whole site's text, one request. For a reader that intends to genuinely digest the site rather than skim it, that single file replaces dozens of fetches. And the older machine surfaces still matter and still ship: a sitemap listing every indexable page for the search-engine crawlers, and a feed announcing new posts. Different readers arrive with different habits; the site keeps a door open for each of them. There's a subtlety in the summary line that's worth an owner's attention, because it's a writing decision, not a technical one. The paragraph at the top of the map is, for many machine readers, the first and sometimes only description of your business they'll ingest. It deserves the same care as your homepage headline: who you are, what you do, who it's for - in plain words, without slogans. A map that opens with marketing copy tells the machine nothing it can use; a map that opens with a clear description gets repeated back to buyers more or less verbatim. The takeaway travels to any site: you already tell human visitors where things are - that's your navigation. The machine equivalents are three small text files, cheap to generate and entirely invisible to your human audience. There is no trade-off to weigh. A site with the map serves both readers; a site without it merely hopes the crawler works everything out. > [!IMPORTANT] Don't make the reader guess > A machine reader gives every site a limited slice of attention. Spend that slice well: hand over an index that says what the site is and where the content lives, and the reader spends its attention on your words instead of on finding them. Next: beyond words and maps - pages that state, in a machine's own terms, what they are. ## Pages that say what they are Clean text tells a machine what a page says. It still leaves the reader inferring what the page *is*. Is this an article or a product listing? Is the business behind it a company or a person? Is the name in the byline the founder, an employee, a guest? People resolve this from a hundred visual cues. Machines resolve it from a purpose-built layer: structured data - a block of labels, in a standard vocabulary, in which the page states plainly what it is. Every page on this site carries that layer. This page declares itself an article, with its title, its description, its dates, and its publisher. The publisher is declared once as an organisation - name, address of record, logo, founder - and every other page points at that same declaration rather than repeating it. The founder is declared as a person, linked to an external profile so a machine can confirm that the Mark Brownlie here and the one on LinkedIn are the same Mark Brownlie. Each label references the others, so the pages assemble into one consistent story rather than a pile of disconnected claims. Here is what that actually looks like - an excerpt from this very page's declaration, trimmed for reading but otherwise as shipped: ```json { "@type": "TechArticle", "headline": "Your Next Visitor Is a Machine", "datePublished": "2026-07-19", "author": { "@type": "Person", "name": "Mark Brownlie" }, "publisher": { "@id": "https://backstage.playhouse.digital/#organization" } }, { "@type": "Organization", "@id": "https://backstage.playhouse.digital/#organization", "name": "Playhouse Digital", "founder": { "@type": "Person", "name": "Mark Brownlie", "sameAs": ["https://www.linkedin.com/in/markbrownlie/"] } } ``` Notice the `@id` doing the joining: the article doesn't repeat who the publisher is, it points at the one declaration of the organisation, and the organisation's founder carries the external link that pins down which Mark Brownlie. Four page types on this site - home, articles, listings, service pages - each declare their own shape, and all of them reference the same business entity. ![Diagram: a declared page carries linked labels - I am an article, published by this business, founded by this person - each referencing the others; an undeclared page sits surrounded by question marks: probably a company, maybe an article, author unknown](/diagrams/machine-ch4-declared.svg) Why go to the trouble? Because answer engines prefer material they can lift cleanly, and declared, structured content is the easiest material there is. The easier a claim is to extract and verify, the more likely it is to be used. Structured data is that idea applied to the page itself. The same logic reshapes the writing, which is where this stops being a developer topic and becomes an owner one. Content that machine readers quote well leads with the answer: a heading that asks the reader's actual question, followed by a self-contained couple of sentences that answer it before elaborating. Short paragraphs. Real tables for real comparisons. Claims with numbers, numbers with sources. None of this harms the human reading experience - it's roughly what good writing advice said anyway - but it now has a second audience with no patience for throat-clearing at all. For your own site, the question to ask is unglamorous: does any page on it say what it is, in the machine layer? Most owner-managed sites answer no - the page builder didn't add it and nobody asked. It's a modest piece of work with an outsized effect, because it upgrades every page from "words a machine must interpret" to "claims a machine can use". > [!IMPORTANT] What gets quoted gets chosen > Answer engines build from the most extractable material available. A page that declares what it is, leads with answers, and sources its numbers is doing the machine's work for it - and the sources a machine finds easy to use are the sources that end up in front of buyers. Next: the one thing underneath all of it that most sites get wrong - speed. ## Site speed Underneath the named surfaces are the ordinary things any decent site should do: one `h1` per page, headings in order, regions labelled for what they are, alt text on every image, a canonical URL so nothing's indexed twice. We do all of it, and it puts us ahead of most: these pages score 100 for accessibility on Google's Lighthouse tool, where the typical site sits around 85 ([HTTP Archive, 2025](https://almanac.httparchive.org/en/2025/accessibility)). The one thing here that actually matters is weight. Every page is built in advance as a finished file. Nothing's assembled when you arrive - no database call, no scripts stitching the content together after it loads. Measured live in July 2026, our pages come in at 12 to 23 kilobytes. The median web page is 2.6 megabytes ([HTTP Archive, 2025](https://almanac.httparchive.org/en/2025/page-weight)), and it climbs every year. A machine notices that. It reads to a budget - so much content per request, so much time per site. A page that arrives whole in 20 kilobytes gets read in full. A 2.6-megabyte page, half of it scripts the reader won't run, gets skimmed or skipped. The buyer from chapter one lands on the same light page and just feels it as fast. ![Diagram: two pages side by side. On the left, this site's page - gold-outlined, 12 to 23 kilobytes, arriving complete with no script to run, read in full. On the right, a typical page - 2.6 megabytes, most of its weight in scripts the reader won't run, content arriving late, sampled and then dropped by a reader working to a budget](/diagrams/machine-ch5-anatomy.svg) Want to check your own site? Run it through Google's PageSpeed Insights - the diagnostics show the total page size in kilobytes. Most page-builder sites come back heavy, weighed down with scripts. That's the page a machine gives up on. > [!IMPORTANT] Weight is the differentiator > Clean structure and alt text are basics - every decent build has them. Weight is what separates sites. A page that arrives whole in kilobytes gets read; a script-heavy, multi-megabyte one gets skimmed and dropped. Machines feel it first, buyers second. Next: all of this is easy to break later without anyone noticing - so nothing ships until the build has checked it. ## It checks itself All of this is easy to set up once and easy to break later. Add a page next month, forget its plain-text twin, and nothing errors - the map just goes on describing a site that's no longer there. Nobody notices, because the readers it fails don't complain. So none of it relies on anyone remembering. Every build runs the same check over every page: has it got its plain-text twin, is it on the site map, does it say what it is? Miss any one and the build stops and names the gap - nothing ships until it's fixed. The next page you add meets the same check, whether anyone's thinking about it or not. ![Diagram: pages flow toward publication through a gold gate asking three questions of every page - plain-text twin present, on the site map, says what it is; complete pages ship, and one bare page is bounced back with the missing pieces named](/diagrams/machine-ch6-guard.svg) We tested it the only way worth testing: added a broken page on purpose - no twin, no map entry, no declaration - and ran the build. It stopped and named the gaps. A check you've never seen fail isn't one you can trust; this one catches. For your own site, three quick checks, none of them technical. Put your address into an AI assistant and ask what the business does and who it's for - what comes back is what your site taught it, gaps and all. Add `/llms.txt` to your address and see if there's a map or a 404. And ask whoever builds your site: "what does a machine reader see when it visits us?" Text without scripts, a map, and pages that say what they are is a good answer. A blank look is also an answer. > [!IMPORTANT] No upkeep > You shouldn't have to remember any of this. Set it up once and let the build enforce it: every page checked automatically, nothing published until it passes. A site that was readable last year won't stay readable on its own. --- ## Posts # A self documenting wiki where the difference is the gold Source: https://backstage.playhouse.digital/blog/a-self-documenting-wiki-where-the-difference-is-the-gold Published: 2026-08-11 Andrej Karpathy published a write-up on what he calls the [LLM Wiki](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f): an LLM agent builds and maintains a persistent wiki over everything you collect, so knowledge is compiled once and kept current instead of re-derived every time you ask. If you run something like it, this takes the premise and moves it along. If you haven't read it yet, it is ten minutes well spent, but that one line is enough to follow along so don't go anywhere just yet. Follow the gist and you'll get a well put together summary of the articles you hold on any given subject. The destination is encyclopedic by nature (as intended): tidy topic pages, one calm voice, facts kept current. A disagreement between your sources has nowhere to live in a page like that, so it gets flattened into a neat paragraph and it reads beautifully. My version treats those disagreements as tensions: held and worked in a companion wiki page beside the main one. > [!IMPORTANT] The gold is the difference > Between sources that refuse to agree, between what a page claims and what its evidence can carry, and between the record of what everyone says and your own ruling on what to do. So take the LLM Wiki as written - the persistent wiki, the model doing the maintenance, you directing. Here are the four steps I use to extend it. ## Step 1: Give the tension its own page Karpathy's write-up spends two lines on contradiction. The agent should note "where new data contradicts old claims", and a periodic health check should look for "contradictions between pages". For facts, that is the right treatment: a price moved, a tool died, one value is wrong. Fix it and move on. Methods break that logic. When two respected practitioners disagree about how to do something, neither value is wrong. That contradiction is a tension: both positions can be valid in different contexts, and the deciding context is often the most valuable thing your collection holds. An agent running the fix-it instruction will repair the tension away, and with it the nuance you built the wiki to keep. So the first step is a routing rule, written into the agent's standing instructions - the schema file Karpathy himself credits with making the agent "a disciplined wiki maintainer". Mine says: a fact contradiction gets fixed; a method contradiction gets space of its own. That space is a second wiki page. Any subject that holds real disagreement gets a pair: the main write-up states what the sources collectively say, and a companion wiki page beside it holds every tension, cross-referenced, with the contradictions intact and the new evidence each one is forced to provide. The write-up stays readable, and the disagreement stays honest, in full, next door. ![One subject, two wiki pages, flowing top to bottom. Everything held on the subject is read in one pass and hits the routing rule in the agent's standing instructions: a fact contradiction goes to the main write-up as a fix, a method contradiction goes to the companion wiki page as a held tension with both positions verbatim. Each tension must land in one of three states - settled, reconciled by a named condition, or left open with the reason recorded - and where it stays open, your dated ruling sits on top with the record intact underneath. A held tension sooner or later yields a sentence none of your sources wrote: that sentence is the gold.](/diagrams/difference-is-the-gold-the-pair.svg) The companion has a discipline of its own. Every entry holds both positions verbatim, with citations, and then must land in exactly one of three states: - settled, - reconciled by a named condition, or - left open with the reason recorded "It depends" is banned as a resting place, and sooner or later that forcing produces a sentence none of your sources wrote: the condition under which both sides are right. My conversion companion holds the cleanest example. One respected school says a landing page keeps the full site navigation, because stripping it damages trust. Another treats removing navigation as a standard structural test and quotes lifts for it. Avinash Kaushik has argued the keep side; Sam Tomlinson works the strip side. Both quote their evidence correctly. Read together in one pass, the three-state rule forced the entry to land somewhere, and it landed on a sentence neither of them wrote: > The implied deciding variable is page type and traffic source ... written for homepages and full landing-pages that double as credibility scaffolding ... while the strip-it guidance is written for single-purpose paid-traffic landing-pages where every nav link is an exit. No source states this condition explicitly, so a reviewer scoring the same page against the two rules reaches opposite verdicts. Reading everything on a subject in one pass keeps finding smaller versions of this. Loss aversion turned up in my pages wearing four costumes: a guarantee rule, a risk-reversal move, urgency copy, and a friction principle and the companion now holds the principle named once, with all four vocabularies pointing at it. That is knowledge the collection contained but could not surface - it existed in no single source and answered to no single search term until a page wrote it down. ## Step 2: The page keeps listening None of this surfaces if pages are assembled from search results at question time. The pair comes from a full read: everything the collection holds on a subject, read in one pass, before a word is written. Buried positions and cross-source clashes only show up when the sources actually sit side by side. The companion has to keep listening. Every new source that lands gets checked against the held tensions: a new voice can settle an open one, or reopen one that looked settled, and the entry's state moves with it. Twenty-one subjects in my collection currently carry the pair. The navigation tension is why the listening matters. This year a third position arrived - Kaushik again - arguing that the disagreement is about the wrong object. Navigation and site search are both find-it furniture: they assume a visitor who already knows what they want. His question - "why still persist with the most prominent spot given to an inferior technology?" - is aimed at the search box holding a site's prime spot, when the visitor who arrives with a real question is better served by a surface that answers it. The companion now holds it as a third side, dated, with the first two intact. Karpathy's wiki would have caught the arrival too - filed as one more fact, folded into the calm paragraph. What a settled page cannot do is mark a new voice as a new side of an old disagreement, so the moment the tension moved would sit in plain text, waiting for a reader who already knew to look for it. ## Step 3: A page that knows where it is thin The second thing the pair buys is a page that grades the evidence it used to build itself. My conversion companion carries a section on where its own knowledge base is weak: "Of the fourteen issues now in scope, three carry any load" - the rest of that prolific source's output is acquisition content, PR, or pitch. Two outside voices carry nearly all the external weight; remove six documents and the outside layer adds little my own curated notes don't already say. That is thin coverage that knows it is thin, and says so on the page. Grading also means refusing. One real disagreement in that collection never made it into the companion as a named tension, because one side rests on a vendor's account inside a sponsored issue and the other on a single practitioner's assertion. Declining to promote weak evidence is grading too: a near-entry filed as nothing, with the reason recorded. And the grading runs before a page exists at all. > [!IMPORTANT] When a subject earns its pair > A subject earns its pair only when the collection disagrees with itself enough to hold a difference, at a scope one page can carry. When I graded my own candidates, I demoted one after a two-minute check showed about 80 percent of everything I held on it came from a single author - one voice is no difference at all, however much of it you have. Another is parked as too broad, because it spans several real subjects one page would flatten. The routing has four exits: build now, confirm density first, demote to a narrower page, or wait. "Wait" is a legitimate answer. ## Step 4: Your ruling sits on top A companion that leaves a tension open leaves you a problem: you still have to act. The last step is a ruling layer. Where the record stays open, a house ruling sits on top of the entry, dated and named as mine, with everything either source said left intact underneath. On the navigation question my ruling is that every page keeps its navigation, single-purpose paid landing-pages included. The two schools stay on the page in full, as the record. The page records what the sources say; the ruling decides what I do. Neither overwrites the other. When the evidence moves, I re-read the tension as it actually stood and re-rule, with nothing overwritten in the meantime. Karpathy's write-up has a line most readers skim: good answers "can be filed back into the wiki as new pages", so your explorations compound. A ruling is exactly that - an answer you file back. Starting costs three edits, all of them inside step one: - a routing rule in your schema file that sends method contradictions to a companion wiki page instead of a fix - one companion pair for the subject where your sources disagree most - the three-state rule, so every tension you hold must end settled, reconciled by a named condition, or open for a reason you wrote down > [!IMPORTANT] The gold > Run a full review on a subject and watch for the sentence that appears in nobody's source. That sentence is the gold. --- # The site is live, and it's AI-native by default Source: https://backstage.playhouse.digital/blog/ai-native-by-default Published: 2026-07-02 The site is live. And from the first commit, it was built to be read by machines as easily as by people. You don't turn that on. You don't think about it. It just does it. That's the part worth pointing at, because you'd never know it was there. Here's what's under the hood. ## Every page has a plain-text twin Add `.md` to the end of any page's URL and you get the clean markdown source: no layout, no clutter, just the words. It's the version an AI assistant can read without tripping over the design. You're reading the styled version now. Here's [this post as plain markdown](/blog/ai-native-by-default.md). Same words, machine-readable. ## A map for the machines There's a file at [/llms.txt](/llms.txt): a plain index of the whole site, written for AI agents, pointing at every page's markdown twin. It's the convention forming for "here is my site, in a form you can read." Most sites don't have one yet. This one did on day one. There's no sitemap.xml here yet: the same idea for search engines, and much older. I built this one first, which is probably the wrong way round. ## Structured data on every page Every page carries structured data that tells a search engine or an AI what it is looking at: that this is an article, who wrote it, when. Invisible to you, load-bearing for them. ## Measurement, wired in from the start Analytics is embedded across the site, with consent handled properly, from the beginning. Not because there's anything interesting to track yet, but because measurement is something you want in place before you need it, not bolted on after. ## Why bother People are starting to ask an AI assistant what they used to type into Google. When that happens, the site an AI can actually read is the one that gets represented properly. Building it in from the start costs almost nothing; retrofitting it later costs a lot. It's the same instinct as the work I do for clients: make the thing legible and measurable, then you can improve it. None of this is the point of the site. It's the plumbing. But it's the kind that's easy to skip and expensive to add back, so it went in first. More soon. --- # Building Playhouse in public Source: https://backstage.playhouse.digital/blog/building-in-public Published: 2026-06-30 This is the first post on a site I'm building out in the open. The plan is simple: show the work, not just the result. The work I do for clients turns evidence into decisions, and decisions into things that ship. The way to prove that is to do it in public, on my own site, first. ## What this site is It's a working example of the thing I sell. Underneath every page is a token engine: one source of brand identity that the whole site reads from, so colour, type and theme can change in one move rather than a redesign. The same engine that themes this site is the one I use to capture and rebuild any brand's design language. For a services business, the website is where the offer wins or loses. So reading a brand's design language, capturing it, and rebuilding a page around it - fast, and in the brand's own voice - is the delivery end of the work. The intelligence tells you what to say; this is how it gets said, live, without a three-month redesign. ## What comes next A short list of what's already in flight: - A live theme switcher, so the whole site can take on a different brand identity on demand. - Real pages for the work itself: what I do, who it's for, and the proof. - More posts like this one, written as each piece gets built. If you're reading this early, it will look rough in places. That's the point. The polish arrives in public, one commit at a time.