Resources

Capture, Distil, Decide

Every piece of marketing work rests on the same handful of facts about your business: who you are, who buys from you, what you’re up against. Those facts exist. They scatter the moment a project ends, so the next one goes looking for them all over again.

This article is about a way of working that stops that: capture what’s known about your business once, keep it in one place that stays current, and have every project read from it and add to it. The whole of it fits in a single sentence.

The whole method, in one sentence

Capture your site into the lake, distil the lake into the foundation, and turn the foundation into a plan.

Three layers, each made from the one before:

  • The lake - what we captured: every page, yours and your rivals’, dated.
  • The foundation - what we know: how your brand looks and sounds, your offer, your audience, your proof, your competitor set - every fact sourced.
  • The plan - what to do: the judgment calls and next steps, the only layer that ever reaches you.

The foundation isn’t filled in once and left. It grows out of the work itself: the brand reading, the competitor reports, the deep dives each leave it knowing more than they found. That’s what makes the second project cheaper than the first and the fifth sharper than the fourth.

The method in three actions, left to right: the websites, yours and your rivals', are captured into the lake, what we captured; the lake is distilled into the foundation, what we know - identity, offer, audience, proof, competitor set, every fact sourced and dated; the foundation is decided into the plan, what to do. Under each layer, how it fails: lose the lake and you keep what you know but can never prove it; lose the foundation and you rebuild it from the lake; lose the plan and the knowledge survives but nothing ships - the only layer that reaches you

The knowledge that goes missing

Whenever someone does marketing work for you, the deliverable is visible and the groundwork behind it isn’t. Every new piece of work - a report, a landing page, a campaign, a website - needs the same facts about your business to hand. Some are written down in an old proposal or a shared drive; some are in browser tabs and someone’s memory of a call. Rarely is any of it in one place, current, and complete. So each project spends its first hours gathering the pieces it needs again, rebuilding whatever the last one didn’t leave behind.

Two things make that expensive. It’s lossy - the piece that lived only in one person’s head is gone the moment they move on. And it’s siloed - the person writing your landing page and the person running your ads are each holding a different fragment, and no one is looking at all of it together. Nobody puts that cost on an invoice, but it sets the ceiling on how fast anything can ship, because the gathering always runs before the work.

Two panels: without a foundation, every project gathers the same scattered facts, delivers, and lets them disperse again; with a foundation, projects stack on top of a captured base, each starting where the last finished and writing back what it learned

What the fix has to do

The right-hand side of that diagram is the promise: projects stacking on a base that already knows your business, each one starting where the last finished. A shared drive full of documents doesn’t get you there, because nothing in it has to stay true, so it drifts. The fix has to do two jobs at once: hold what’s known in a form every project actually reads, and take back what every project learns. The arrows pointing down into the base matter as much as the base itself.

If you want to know where your own business stands, ask the last three people who did work for you what they worked from. Whatever comes back, and whatever doesn’t, is your foundation as it exists today.

Doing both jobs properly needs three distinct layers, one for each step in the opening sentence. The first holds the raw material.

A page is not a scrape

The first layer is the lake: every captured page, yours and your rivals’, held in one place with a date on each. The name comes with a test attached: after capture, nothing downstream ever needs the live site again, because everything after it reads the lake.

So what does “captured” have to mean for that to hold? When people say they’ve captured a website, they usually mean one of two things: a scrape, which rips the words out and throws away everything about how they were arranged, or screenshots, which keep the arrangement and throw away the machine-readable words. Each answers exactly one kind of question. The lake has to answer questions nobody has asked yet - so it needs both, plus one more thing neither provides.

Three files

In my system, capturing a page produces three files:

  • What it says - the words, as clean readable text
  • What it is - the structure: navigation, headings, links, metadata, the skeleton a rebuild needs
  • What it looks like - a full screenshot, because some truths only exist visually

Flow diagram: one page, visited once, becomes three files - text, structure, screenshot - landing in the lake alongside competitor pages, with a per-page ledger of when it was captured and when it actually changed; brand capture, intelligence reports and the site build all read from the lake and none revisit the website

Together, the three files are the page as it stood on the day: what it said, how it was built, how it looked. Enough for a later project to read its words, rebuild its structure or study its design without going back to the site. Not everything a website does lives in those three files - how fast it loads, how a form behaves, whatever sits behind a login - so anything in that class is captured on purpose or not at all.

Here are those three files for one real page, my own About page, as it was captured on 26 June 2026:

playhouse-digital__about-about-us.md       5.1 kilobytes   what it says
playhouse-digital__about-about-us.json     5.7 kilobytes   what it is
playhouse-digital__about-about-us.png    412.8 kilobytes   what it looks like

The first opens straight into readable text, no markup to strip:

# About “Us”

I bring value to supporting businesses of all sizes that market directly
to consumers where success is measured by generating a lead such as a
form fill or a phone call from their website...

The second holds the same page as structure and nothing else: its canonical address, its social metadata, one heading, three navigation blocks, twenty internal links, twenty-nine calls to action, and three structured-data declarations. No words, no design. The third is the screenshot, and at more than seventy times the size of either text file it carries what neither of them can.

The ledger

Each page in the lake also carries a small ledger: when it was first captured, when it was last captured, and when its content actually changed. Those are different facts. A page re-checked yesterday that hasn’t changed since March tells you something a scrape never could: the evidence is current and stable. It’s what lets the layers above stay trustworthy without endlessly re-reading everything - and it means every claim built on the lake can carry a date. “Their homepage said this, as of this day” is evidence. “I remember their homepage saying this” is not.

Any site. Yours - or your biggest rival’s

The capture doesn’t know whose site it’s reading. The same machinery that puts your website in the lake puts any website there. Point it at your own site and you get the base everything in this article builds on. Point it at your biggest rival and you get their words, their structure, their look - dated, complete, and sitting in the lake one folder away from yours.

So your pages and your competitors’ pages live side by side, in the same three-file format. It sounds like a filing detail, and it decides every comparison downstream: messaging gaps, design benchmarks, positioning all read like-for-like captured evidence rather than one fresh site against one remembered one.

That is what the test at the top of this chapter buys. Capture a site once and every later piece of work reads the capture instead of the site, which is why one pass over a site keeps paying out for weeks.

The lake holds everything and knows nothing. Turning what it holds into knowledge starts with the most visible thing a website carries: the brand.

Reading the brand: how it looks and how it sounds

Your design language is already on your website: every choice of typeface, colour, spacing and shape. What it has probably never been is collected in one place. It lives as an accumulation of decisions, held together by whoever built the site and your eye for “that looks right.”

The same is true of your voice. How your business talks - the words it reaches for, the promises it makes, how formal or plain it is with a buyer - is just as real, and wherever it has been written down, it is usually in the copy itself rather than in a guide anyone could hand to a new writer. Both halves of the identity are sitting in the lake, waiting to be read. One is read from the styles; the other from the words.

The look: read from the styles

The design half rests on one fact: to draw a page at all, a browser has to compute the final, resolved value of every style on every element. That resolved value is the result, whatever anyone intended. So the design language isn’t lost. It’s fully specified, machine-readable, and sitting in plain sight on the rendered page.

Reading those computed styles off a page yields thousands of raw values. Distilling them down produces a small set of role-assigned tokens: this typeface is the display face, this one the body, this colour is the accent, this is the surface everything sits on, these are the shapes. Alongside the tokens comes a design document - the brand’s rules, written down for the first time.

Flow diagram: a rendered page yields thousands of computed style values; distillation produces an identity file of role-assigned tokens, each graded exact, derived or defaulted, plus a readable design document; below, the same process pointed at a rival's site makes brand comparison a comparison of two files

And every token carries an honesty grade: exact (read directly off the page), derived (computed from what was read), or defaulted (nothing to read, filled from convention). The grade is what separates capturing a brand from guessing at one. Role by role, you can see how much of the file was read off your site and how much was filled in for you.

The voice: read from the words

The other half of the reading works on the text in those same files. Every page you’ve published is a sample of how your business speaks: what it leads with, what it calls your customers, whether it says “we help” or “solutions are delivered”. Distilled the same way, that becomes a written tone of voice - the words your brand uses, the ones it never would, and how it talks to the person you’re trying to win.

And because the rivals’ pages sit in the lake in the same format, the same reading runs across the whole market. How does everyone else in your space talk? Where do the voices blur into one interchangeable hum, and where would a different way of speaking stand out? Every line of that answer traces to something a competitor actually published, dated and on file.

The trust layer nobody could measure

In one study, people shown a page for 50 milliseconds rated its visual appeal much as they did when shown it for ten times as long (Lindgaard et al., 2006). That is too fast to be reading, and too fast to be reasoning. Whatever a visitor is judging that quickly, it isn’t your copy. The look of the page has already landed.

That has been difficult to act on, because nobody could say what was wrong. You could feel that a site was “off”; you couldn’t say by how much, or exactly where.

A look you can measure

Once a brand is a file, that changes. It can be measured, compared, versioned and handed over. The part of your business that was hardest to pin down becomes the part that travels best.

Point it at your rival

The lake chapter said the capture doesn’t know whose site it’s reading. Neither does this step - and that’s where it stops being an efficiency and becomes an advantage. Run the same reading over your biggest rival’s site and you get their identity as a file: their type scale, their palette, their level of craft, and their voice - graded with the same honesty.

“How does our brand hold up against theirs?” is usually answered with feelings. Now it’s a comparison of two files - role by role, measurably, with the evidence attached.

A brand nobody has written down is a brand nobody can hold to account. Yours is a file now, and so is theirs. Where both files go next is the second layer.

The foundation: what we know

Ask what any marketing project needs to know about your business and it comes down to the same five objects: the identity (how your brand looks and sounds), the offer, the audience, the proof, and the competitor set. In most businesses all five exist - scattered across formats and folders, each version written for one project and none maintained.

The second layer, the foundation, is those five objects held in one structured place: current, maintained, and with every fact carrying its source and its date. “Their guarantee changed in March” is a foundation fact, and the captured page that proves it is sitting in the lake underneath.

Built by the work, not before it

Nobody sits down and writes the foundation. It comes out of the work. The brand reading in the last chapter deposits the identity. A competitor report deposits the competitor set and the unclaimed positions. A deep dive deposits what one rival is really doing. The foundation is the accumulation of every reading and every report that’s ever run, which means it’s never a form someone filled in optimistically on day one and forgot. Every fact in it is there because a piece of real work needed it, found it, and left it behind.

That also means it starts faster than a documentation project would. The first capture opens it, the brand reading fills in the identity, and everything after that adds to it. The five objects are the same five in every business, whatever it sells, so the shape doesn’t change from one to the next.

Reading replaces re-learning

The first project pays to start the foundation. Every project after it begins by reading what’s already there. The second competitor report doesn’t begin at zero; it begins where the first one finished, because the identity, audience and competitor set are already sitting there, current and sourced.

What makes this practical now, rather than a nice theory, is that an AI can read the entire foundation before every task and write what it learned back afterwards. Reading everything known about a business before every single deliverable was never worth a person’s time, which is why nobody did it. Handing that reading to a machine is what changed.

A quick test for your own business: name the five objects, then see whether you could put your hands on a current version of each one in the next ten minutes. The ones you can’t find are the ones your next supplier will rebuild from scratch and bill you for.

That’s the foundation

One place where everything known about your business accumulates, sourced, dated and current. Nothing in it was written to impress anybody. Every line is there because a piece of real work needed it.

The three layers earn their separate names because they fail in three different ways. Lose the lake and you still know everything; you just can’t prove any of it, or date a claim. Lose the foundation and nothing is gone for good: it rebuilds from the lake, because it’s derived, never primary. Lose the plan and the knowledge survives, but nothing ships.

Holding the three apart by name is the working discipline behind everything else. Every piece of work can say which layer it reads and which it writes: the brand reading reads the lake and writes the foundation; a competitor report reads the foundation and feeds the plan. That one habit keeps evidence, knowledge and judgment from blurring into each other as the system grows.

It’s a question worth asking of anything you get sent: which of the three is this? Evidence, knowledge, or a decision. Work that can’t answer is usually none of them.

The branches that grow the foundation

With a foundation started, intelligence work stops being a series of one-off projects and becomes a set of branches. Each branch reads what’s already known, asks one question of the evidence, and writes its answer back. The competitor report asks where the messaging gaps and unclaimed positions are, across your whole rival set. The deep dive asks what one rival is doing, forensically. The design benchmark asks how your brand’s craft measures up, file against file. The performance analysis asks where your spend is leaking. Each asks a different question; all of them ask it of the same evidence.

Diagram: your foundation at the centre; four branches around it - competitor report, deep dive, design benchmark, performance analysis - each reading from the foundation in grey and writing findings back in gold; everything to the left culminates on the right in the plan: the assembled view of decisions, not data, and the reasoning behind them kept as long-term memory

Every deliverable is also a deposit

Reading is only half of what a branch does. It writes back what it learned. When a competitor report finishes, its findings about each rival file into the foundation as dated, sourced knowledge. Run the report again next quarter and it doesn’t start from a blank page; it starts from everything the last run established, and spends its effort on what changed.

Without the write-back, intelligence evaporates into PDFs. With it, every deliverable grows the foundation it was built from, and the fifth question you ask costs a fraction of the first, because everything before it has already paid in.

When the intelligence was wrong

The system has shipped wrong findings. One report read a routine lead time as a scarcity signal, and called a proof gap that turned out to be sitting one click from the homepage. Human review caught both, and the report was rebuilt from verified capture and re-issued. The fixed report was the small outcome. The lasting one was that the pipeline gained the checks that catch that class of error, and the corrections went back into the foundation with everything else.

So when you read a finding, the useful question is what would have happened if it were wrong. Somebody noticing is a weaker answer than something checking.

Every branch ends the same way: with answers about what’s wrong, what’s missing, what’s unclaimed. Answers on their own don’t run a business.

The plan: decisions, not data

The plan is what you actually receive: judgment calls, to-dos and next steps, benchmarked against your named rivals, assembled across every branch into one narrative with a single prioritised action list.

What you’re actually getting

Nobody needs a stack of findings. What you need is knowing what to do on Monday.

Getting there is assembly and judgment rather than more analysis: every branch already ends in answers, so the plan’s job is to weigh them against each other and put them in order.

Diagram: four intelligence branches - competitor report, deep dive, design benchmark, performance analysis - feed one gold document, the plan: judgment calls on what to claim and what to drop, to-dos in priority order, every line benchmarked against named rivals, the reasoning kept for the decision after; a set of decisions, not a stack of reports, ending in what to do on Monday

Of the three layers, the plan is the only one that ever reaches you. The lake is evidence and the foundation is memory; both exist so the plan is right, and so the next one is sharper. It’s a useful filter to apply to anything you’re sent: if a page of it doesn’t change a decision, it’s trivia, however interesting the data behind it.

The reasoning travels with the decision

What makes the plan compound is that it keeps the decision and the reasoning together: the evidence behind it, and what was ruled out. Come back two quarters later and the question “why did we drop that channel?” has an answer on file, so the next decision starts from the last one instead of from scratch. Decisions with their reasoning attached become long-term memory; decisions on their own are just this quarter’s opinions.

A plan still doesn’t move a business until something ships against it, and shipping has a discipline of its own.

From diagnosis to treatment

Every line in a plan splits the same way: a diagnosis - what the intelligence says is wrong, missing or unclaimed - and a treatment - the artifact that acts on it. The design read pairs with a rebuilt page. The competitor gap analysis pairs with positioning, which feeds generated landing pages. The search-term analysis pairs with restructured campaigns.

Diagnosis and treatment need each other

Diagnosis without treatment is trivia; treatment without diagnosis is guesswork with good production values.

Capture, distil and decide stop at the plan: they are how knowledge gets made. Treatment is how work gets made.

The pairing also means you can see why a piece of work is being proposed before you agree to it. The diagnosis names the problem and shows the evidence behind it; the treatment is the answer to that problem and nothing else.

Two-column diagram: diagnoses on the left - the design read, the competitor gap analysis, the search-term analysis - each with an arrow to its treatment on the right - the rebuilt page, positioning feeding generated pages, restructured campaigns; a gold gate on the middle arrow marks that vague positioning is refused, and a return arrow along the bottom notes that what ships writes back

Rebuilding the page properly

When the treatment is a website, new colours on the old structure won’t do it. The reading that captured your brand becomes the way into rebuilding the page itself: the structure, the order the argument unfolds in, the journey a visitor takes from arriving to enquiring - designed for the audience the foundation identified, saying what the voice reading says will land, looking the way the identity file says it should look.

All of that puts a condition on the work

A page built this way has to know what it’s arguing before anyone writes a line of it.

Variants, scored against each other

One more discipline when the treatment is a page: never generate one and hope. The system generates structurally different variants of the same page - different arguments, different structures, same locked positioning - and each one is scored independently, out of ten, on one dimension at a time.

Here is a real scoreboard from a run on my own landing page. The same proposition was built three ways: one written from the facts alone with no brand styling, one with the same facts inside the brand, and one with every structural module in place.

Dimension (out of 10)Facts onlyOn-brandFull buildWinner
Comprehension689Full build
Relevance559Full build
Differentiation348Full build
Trust994On-brand
Craft885On-brand
Design785On-brand

Read down the columns and the split is hard to miss. The version with the best message had the worst proof; the version with the best proof had the least to say. Neither of them shipped. The merged page took its hero and its differentiation from the full build, and its proof, its craft and its whole visual frame from the on-brand one. The facts-only version won nothing at all, which is still a result: it set the proof discipline the winner had to match.

Lifting the winning copy also meant losing some of it. A claim about how many reports had been delivered went, because at the time the honest count was one, and so did a competitor price comparison with no source behind it.

One identity, every surface

None of that consistency is maintained by hand. Because the identity is data, every surface the work produces for you wears it automatically: the rebuilt pages, the intelligence reports, the documents. Your brand is defined in one place and read from there by everything, so nothing drifts out of sync, and a brand change ripples everywhere at once.

You can check that claim against the page you’re reading. Here are five lines out of the identity file this site is built from, and where each one is on screen right now:

RoleValueWhere it is on this page
primary#16181Dthe chapter headings
body#565B62the text you’re reading
accentAccessible#8B6B1Athe gold labels above each highlighted box
binding#0B2545the deep panels heading tables in the reports
heading typefaceFrauncesevery heading you’ve read so far

Twenty values in total: twelve colours, two typefaces, one corner radius, five spacing rules. Swap that one file and the same page renders in someone else’s brand instead, with nothing else touched. Change a value inside it and every surface follows, which is the whole point of keeping a brand as data rather than as a PDF nobody opens.

Hub and spoke diagram: the identity file - about twenty decisions - at the centre, with one gold arrow in labelled edit here, once, and arrows out to your website, intelligence reports, landing pages, and documents, all sitting on a slab of house craft that never varies by brand

That invites a fair worry. If one small file is the only thing standing between your surfaces and anyone else’s, how much of what you get is actually yours?

The gates

What’s shared is the process: the capture, the readings, the reports, the checks. What comes out of it is built from your lake and your foundation, so the insight and the output are yours alone. A template starts from the same answers every time. This starts from the same questions, and your website, your rivals and your market make the answers different every time.

What makes that shared process safe to run at speed is the last piece of machinery: the gates. A system that can produce a report, a page or a whole site in minutes can produce a wrong one just as fast, so speed is only worth having if you can trust what comes out. That trust is built in as gates: automatic checks on every path to “shipped” that block rather than advise. Nothing reaches you that hasn’t passed them.

Flow diagram: generated work - pages, reports, sites - passes through a wall of three blocking gates: the readability floor, the completeness guard, and configuration traps; passing work ships, and any red gate stops the build entirely

Two of the three gates check ordinary hygiene, and they should. A readability floor: every text-and-background pairing has to clear the AA contrast level of the Web Content Accessibility Guidelines. A completeness guard: every page has to carry its machine-readable mirror, its structured data and its place in the sitemap. Any competent build clears both of those. That second gate runs on this site, and what a machine reader actually does with those pieces once they exist is the subject of Your Next Visitor Is a Machine. What differs here is what happens when one fails: the build stops and names the gap, instead of logging a warning into a report nobody opens.

The third is less ordinary. A site refuses to produce a production build while its enquiry form still points at an interim inbox, so the mistake that’s easy to make and expensive to discover live is simply not buildable.

Blocking is what makes them worth having. A gate that warns gets ignored on a busy day; these stop the line.

A gate before anything is built

The same discipline sits earlier than the build. On the diagnosis-to-treatment path, one arrow carries a gate of its own: the page generator refuses to run on vague positioning. If the positioning can’t say plainly who the page is for, what’s being claimed, and why it beats the alternatives, no page gets generated.

A beautiful page built on mush is still mush

It just fails more expensively, in front of paid traffic.

That refusal encodes something most marketing production gets backwards. The bottleneck in good pages sits upstream of the writing and the design, in the thinking. So the thinking becomes a hard dependency - you can’t skip the diagnosis to get to the treatment faster. The build gates stop bad work shipping; this one stops it starting.

Green ticks lie

Gates have a blind spot, and it’s worth knowing about even if you never go near one. A check that silently fails to run reads exactly like a check that passed. A build can be green because everything works, or green because the thing that should have failed never ran at all. Two habits guard against that.

Count, don’t just exit. A photo gallery built to skip any broken entry can pass its checks with no photos in it at all. Only a count of how many images actually rendered proves the green is real. Wherever absence can pass silently, the gates count.

Measure the artifact. One build here looked right on every page and passed every audit while quietly shipping every photo at several times the resolution it needed. Nothing caught it until someone listed the built files by size and asked why they were so big. Looking at what was actually produced, rather than at the ticks, is the habit that finds that class of problem.

Gates that know their limits

When a rival’s captured brand fails the readability floor, the system warns instead of failing - their identity is what it is, and the capture’s job is fidelity. The blocking floor applies to what ships to you. A verification system that can’t tell “this is our standard” from “this is their reality” would fail at both jobs.

For your own site, the question to ask whoever builds it is what happens when a check fails. If the answer is a warning in a report, the check is advisory, and advisory checks get skipped on the day everyone is busy.

That’s what the gates buy: a process that repeats for any business, at any speed, without the output ever being less than checked, and without it ever being anyone else’s.

The loops

Run end to end, on your business, it goes like this.

The capture comes first: every page of your existing site into the lake as three files, your rivals’ alongside them. The brand is read out of the capture - your identity as a file, graded for honesty, down to the difference between the shade you display and the shade you can set text in. The foundation grows from there - identity, offer, audience, proof, competitor set, every fact carrying its source. The branches run, each asking its one question of the same evidence and writing the answer back, so the second run starts where the first finished. Their answers assemble into the plan: what to claim, what to fix, what to build, in order. And the treatment ships against it, through the gates, on the identity the capture read - with the files yours at the end of it.

Flow diagram: your website is captured once into a lake of pages as three files each; the lake is distilled into the foundation - identity, offer, audience, proof, competitor set; intelligence branches read the foundation and write back what they learn; their answers assemble into the plan - decisions, not data; everything renders from the same foundation and ships as public proof, and the loop begins again

Each stage’s output is the next stage’s input. Nothing gets re-learned at any point along the way.

The loops inside

Inside that end-to-end run, each branch is a loop of its own. It reads the foundation, asks its question, and writes the answer back, so the next run of that branch starts where the last finished rather than from a blank page.

The same write-back applies after something ships. Performance data, responses, your own reactions all return to the foundation, so the next diagnosis starts sharper than the last. Every treatment that goes out makes the one after it better informed. What leaves writes back.

The system grew while it ran

Some of the gates in the last chapter exist because a real build needed them. The trap that refuses a production build with the wrong form inbox was written the day that mistake became possible. Every engagement leaves the system with better defences than it found.

Why the checks keep changing

Processes that never change are processes nobody’s checking.

Your next step

If you’re wondering what any of this means for you on Monday, it starts with one small step: the capture. An afternoon, no meetings, nothing needed from you but the website you already have. Everything else compounds from there.