Friday morning, 31 July. Vera's daily briefing is on my phone with the others, and one line in it asks for a decision. Twelve days earlier she had raised five bids on our competitor ads by a quarter, with a prediction written down before she touched anything: how many clicks should follow, what a customer may cost, and the number at which she'd give up. The clicks came. The customers didn't, not at a price that made sense. Now the cost had crossed the line she'd written down herself, and she wanted the bids back where they were.
She could have done that without me. Rolling a bid back is inside what she's allowed to touch. She asked anyway, as the switch that lets her change the real account had only been live for a couple of weeks and she'd been told to be careful while it was new. I typed yes. At 09:22 the live Google Ads account changed, and the log line underneath said no budget had moved.
Vera isn't a person. She's a folder on a Mac mini.
Two agents, both of them folders
The mini is the same machine that runs the nightly tests from part two. Next to that work live two agents. Vera is the CMO. Sisi runs ops. Franz is named after an emperor, so the ops agent got his wife's name, and her identity file says she's "quietly running the empire behind the curtain". History has it that the real Sisi wanted nothing to do with running the empire. Ours does.
Each of them is a folder with a handful of markdown files at the top. One says who she is. One says what she may and may not do. One lists her tools and where the data lives. The longest one is her operating doctrine, the loop she runs every morning, in numbered steps. Below those sits a knowledge base, split into facts, playbooks, hypotheses and decisions, and next to it a ledger of experiments. Both agents run on a GPT model through an agent runtime called OpenClaw, and both wake up on a schedule. Vera wakes at seven every morning for her ops turn, on Mondays for content, on Wednesdays for research and on Sundays to grade herself. Each wake is one turn: do the work, leave the folder tidy, stop.
They read revenue, not vanity metrics
Around six every morning, scripts on the mini pull fresh snapshots into Vera's folder: Search Console, the ads account, web traffic, and the full product funnel from signups to paying customers, the same numbers I look at on my own dashboard. Since July there's also a join between ad clicks and the customers those clicks actually became. Most marketing I've seen works from clicks and conversions the way the ad platform counts them, which is quite a different thing. Vera has to cite the revenue file before she's allowed to raise a bid.
A manifest says which collectors ran. If one failed, her doctrine tells her to say so in the briefing and not pretend the data is fresh. A ton of dashboards could use that sentence.
None of this needs a model. Fetching data on a schedule is a shell script. The model is what reads the snapshots at seven, notices that a campaign's cost per customer has drifted, and decides whether to hold, propose or act.
Hands, inside a fence
Here's the part people ask about first. Vera can change the real ads account. Bid changes and negative keywords apply live, and since 13 July they can apply without me.
What makes that survivable is a fence, and the fence is code on the host, not a rule in her prompt. She can't raise a budget, ever. A single move can change a bid by at most a quarter. There's a ceiling per click. She can only touch entities on an allowlist, a batch either applies completely or not at all, and every call is logged. Budgets, new campaigns and new keywords stay proposals, which means a card on the board and a decision from me. The whole thing ran in shadow mode for the first weeks, where she proposed and nothing applied, before I flipped it live.
The other half of the fence is in her doctrine, and it's one sentence: "An action without a falsifiable prediction is not allowed." Before she touches a bid she writes an experiment file. What she believes, what should change, by how much, measured how, and the date she'll check. The July step I opened with had all of that, including the stop-loss. When it triggered she rolled back, then scored the rollback as a win for cost control and explicitly refused to count it as growth. Those are two different things, and her playbook has a rule that says so.
If you take one thing from this post for your own product, take the fence. Find the actions where the worst case is bounded and boring, and let the agent do those for real. Keep the rest as proposals.
The board is where I come in
The folder is the truth, and it's mirrored into Notion as one board for all agents. Four roles can own a card: me as CEO, Vera, Sisi, and an external maintainer. A card moves from assigned to in progress to done, or gets handed to the next role, or parks as "blocked, needs consent". A small poller on the mini checks the board every few minutes and wakes an agent only when a card is assigned to it.
Anything that leaves the machine parks. Publishing a post, opening a pull request, sending a mail, spending money. The agent writes exactly what it needs approved into the card, sets the decision field to awaiting, and stops. My view of the board is a filter called "Needs my decision". I set the field to approve, reject or needs changes, sometimes with a comment, and the poller routes the card back to the agent. Quite often this happens on my phone in the coffee shop, in the same few minutes where I merge the dependency updates from part three.
Two lines from the board protocol tell you more about the design than the rest. The role list includes a CTO and a CDO, marked as roles that don't exist yet and must never be assigned to, because "it would dead-end". And: a missing tool is not a handoff. Every agent runs in the same sandbox and hits the same wall, so an agent that can't do something parks the card for me and names the capability it needs. Fixing it adds a script on the host, never a role on the org chart.
Vera opens pull requests, and writes an RFC every Wednesday
Vera has her own copy of the Franz repo. The main mirror is read-only; beside it she has a worktree she can write to. In June she noticed the website had no llms.txt, the file that's supposed to tell AI assistants what a site is about, and predicted that adding one would lift referrals from ChatGPT and Gemini by a fifth. She wrote the route, wrote the test and opened the pull request, and it was merged on 17 June. Thirty days later the referrals hadn't gone up by a fifth. They'd fallen by more than half. She scored it refuted and wrote the lesson into the knowledge base: cheap hygiene, not a traffic lever. If you're adding an llms.txt to your site this month, you now know more than most people doing it.
Wednesdays are for research. Since June she has written 18 RFCs on channels I never had time to look at properly: the Microsoft Store, Setapp, Flathub, IndexNow, a Portuguese localisation pilot, AlternativeTo reviews, GitHub as a release surface, the Mac App Store. Two arrived this morning. Each one is a card on the board with the evidence attached, and each one is work that used to happen once in an eternity, the week before something forced it.
Sisi's side
Sisi's mornings start with the nightly QA report. She reads one file and relays it to my phone: what was tested, what passed, which pull requests the night opened. Her instructions forbid her from checking anything herself. If the file doesn't say it, she doesn't say it. That boundary is the one part two was about, seen from her end.
The rest of her day is the board. When I assign her a bug, she hands it to the dev station, the runner that also repairs tests at night, and the job comes back as a pull request with proof attached: a screenshot of the fixed screen for anything visual, the verifying command output for anything else. The same proof lands on the board card before I'm asked to review. She can boot a real Linux VM on the mini, install the latest release into it, check that the window paints and four other things, and tear it down. In July she restored a messaging recipe that a fix had dropped by accident, prepared a preview release, and merged two green dependency updates after I said yes in the morning digest.
None of that is glamorous, which is the point. It's the ops work that belonged to nobody when Franz had five people, and it now belongs to a fox emoji with a folder.
What Vera can't see
The first question any company asks about a setup like this is what the agent gets to see. So here's what Vera sees of my customers: nothing.
The snapshots in her folder are aggregates. No names, no mail addresses, no per-user rows. I actually checked while writing this, by searching every snapshot file for an @ sign. The only one in there is the address the collector logs in with, which is mine.
In July she asked a question I'd never asked: are people who run five or more WhatsApp accounts in Franz a real segment, or noise? She couldn't query the database. What she could do was ask for the product to compute one more aggregate, the same kind it already computes for my dashboard: how many accounts sit in that bucket, how many of them pay, how many were active in the last month. The answer came back the same day. The segment is real, big enough to matter, and it converts at several times the rate of everyone else.
Then she decided not to build a landing page for it, because search demand for that phrasing is tiny, and wrote that decision down too.
You can give a marketing agent customer insight without giving it customers. Hand over the aggregates you already compute for yourself, and nothing else.
Sunday she rewrites her own rulebook
Every Sunday at nine, Vera scores every experiment that reached its date, then grades herself. The calibration playbook in her folder has 20 rules now, and every rule ends with the date and the miss that produced it. The first one came out of her first week: she proposed fixing wasted mobile ad spend that I'd already fixed three days earlier, because she'd read a seven-day average that straddled the fix. Rule one: check the days after a change before diagnosing from a window that contains it.
The score she publishes is quite unflattering. On metric predictions she's wrong seven times out of ten, and she's the one who writes that down, every Sunday, next to the sentence that confidence is low. Twice this summer the number got worse because she audited her own audit and found experiments she'd forgotten to count. Marketing usually doesn't get a number like that, because nobody wants to be the one writing it down. She does not mind.
Both agents also keep a dream diary. At three in the morning something writes a short entry about the day, with a haiku in it. Sisi's entry after the two dependency merges describes a drawing of a robot watering two sprouts. I'm not turning that off.
If your team wants to try this
The part that would transfer to your product isn't the model. It's the folder: a doctrine, a fence written in code, data that includes revenue, one board with a needs-my-decision column, and a Sunday where the agent grades itself. Every company already has most of that. It's spread across five people's heads instead of one folder.
Vera and Sisi are the half of my day I hand over. The other half I still do myself on this laptop, with agents next to me, and it's changed what a working day feels like. That's part six.
If you're a founder working out how much of this you can hand over, and where the line should sit, that's the work I do with teams. Let's talk.