← All writing
Daily Driver 5 / 7

Vera runs marketing, Sisi runs ops, and I say yes or no

One of them changes live ad bids inside a fence made of code. The other merges pull requests after I say yes in the morning. Both are folders on a Mac mini, and the folder is the part any company could copy.

Sep 2, 2026 11 min read Original

Friday morning, 31 July. Vera's daily briefing is on my phone with the others, and one line in it asks for a decision. Twelve days earlier she had raised five bids on our competitor ads by a quarter, with a prediction written down before she touched anything: how many clicks should follow, what a customer may cost, and the number at which she'd give up. The clicks came. The customers didn't, not at a price that made sense. Now the cost had crossed the line she'd written down herself, and she wanted the bids back where they were.

She could have done that without me. Rolling a bid back is inside what she's allowed to touch. She asked anyway, as the switch that lets her change the real account had only been live for a couple of weeks and she'd been told to be careful while it was new. I typed yes. At 09:22 the live Google Ads account changed, and the log line underneath said no budget had moved.

Vera isn't a person. She's a folder on a Mac mini.

Two agents, both of them folders

The mini is the same machine that runs the nightly tests from part two. Next to that work live two agents. Vera is the CMO. Sisi runs ops. Franz is named after an emperor, so the ops agent got his wife's name, and her identity file says she's "quietly running the empire behind the curtain". History has it that the real Sisi wanted nothing to do with running the empire. Ours does.

Each of them is a folder with a handful of markdown files at the top. One says who she is. One says what she may and may not do. One lists her tools and where the data lives. The longest one is her operating doctrine, the loop she runs every morning, in numbered steps. Below those sits a knowledge base, split into facts, playbooks, hypotheses and decisions, and next to it a ledger of experiments. Both agents run on a GPT model through an agent runtime called OpenClaw, and both wake up on a schedule. Vera wakes at seven every morning for her ops turn, on Mondays for content, on Wednesdays for research and on Sundays to grade herself. Each wake is one turn: do the work, leave the folder tidy, stop.

When they wake, and for what
Every day
06:00 scripts fresh snapshots land in the folder07:00 Vera ops turn: score due experiments, read, monitor, actmorning Sisi relays the nightly QA report to my phone03:00 both dream diary, with a haiku
Monday
08:00 Vera content: one or two drafts to the preview
Wednesday
08:00 Vera research: one RFC on a channel
Sunday
09:00 Vera self-review: grade every prediction, rewrite the rules
Each wake is one turn. The 06:00 row is shell scripts; everything else is a model reading what the scripts left behind.

They read revenue, not vanity metrics

Around six every morning, scripts on the mini pull fresh snapshots into Vera's folder: Search Console, the ads account, web traffic, and the full product funnel from signups to paying customers, the same numbers I look at on my own dashboard. Since July there's also a join between ad clicks and the customers those clicks actually became. Most marketing I've seen works from clicks and conversions the way the ad platform counts them, which is quite a different thing. Vera has to cite the revenue file before she's allowed to raise a bid.

A manifest says which collectors ran. If one failed, her doctrine tells her to say so in the briefing and not pretend the data is fresh. A ton of dashboards could use that sentence.

None of this needs a model. Fetching data on a schedule is a shell script. The model is what reads the snapshots at seven, notices that a campaign's cost per customer has drifted, and decides whether to hold, propose or act.

Hands, inside a fence

Here's the part people ask about first. Vera can change the real ads account. Bid changes and negative keywords apply live, and since 13 July they can apply without me.

What makes that survivable is a fence, and the fence is code on the host, not a rule in her prompt. She can't raise a budget, ever. A single move can change a bid by at most a quarter. There's a ceiling per click. She can only touch entities on an allowlist, a batch either applies completely or not at all, and every call is logged. Budgets, new campaigns and new keywords stay proposals, which means a card on the board and a decision from me. The whole thing ran in shadow mode for the first weeks, where she proposed and nothing applied, before I flipped it live.

Hands, inside a fence
never
¼ downthe bid¼ upceiling
one move out of reach for one move out of reach, full stop
19 July, one step up 31 July, back where it was
Checked on every call never a budget increaseonly entities on an allowlistall of a batch or noneevery call logged
Outside the fence a bigger budget, a new campaign, new keywords or ads, publishing a post, opening a pull request, sending anything. Each one is a card on the board and a decision from me.
The fence is code on the host. Shadow mode first, live since 13 July. The model never gets to negotiate with it. Not to scale.

The other half of the fence is in her doctrine, and it's one sentence: "An action without a falsifiable prediction is not allowed." Before she touches a bid she writes an experiment file. What she believes, what should change, by how much, measured how, and the date she'll check. The July step I opened with had all of that, including the stop-loss. When it triggered she rolled back, then scored the rollback as a win for cost control and explicitly refused to count it as growth. Those are two different things, and her playbook has a rule that says so.

If you take one thing from this post for your own product, take the fence. Find the actions where the worst case is bounded and boring, and let the agent do those for real. Keep the rest as proposals.

An action without a falsifiable prediction is not allowed.

The board is where I come in

The folder is the truth, and it's mirrored into Notion as one board for all agents. Four roles can own a card: me as CEO, Vera, Sisi, and an external maintainer. A card moves from assigned to in progress to done, or gets handed to the next role, or parks as "blocked, needs consent". A small poller on the mini checks the board every few minutes and wakes an agent only when a card is assigned to it.

Anything that leaves the machine parks. Publishing a post, opening a pull request, sending a mail, spending money. The agent writes exactly what it needs approved into the card, sets the decision field to awaiting, and stops. My view of the board is a filter called "Needs my decision". I set the field to approve, reject or needs changes, sometimes with a comment, and the poller routes the card back to the agent. Quite often this happens on my phone in the coffee shop, in the same few minutes where I merge the dependency updates from part three.

One board, four roles, real cards from this summer
Assigned2 Sisi Prepare the preview releaseVera Sunday self-review: score everything that reached its date
In progress2 Vera Wednesday RFC: Setapp as a Mac marketplaceSisi Restore the messaging recipe a fix dropped by accident
Needs my decision3 Vera Roll five competitor bids back to where they wereVera Publish the Mac post that has been in preview since JulySisi Merge two green dependency updates
Done2 Vera llms.txt route and test, merged 17 JuneVera Windows content experiment scored: confirmed, but weak
Anything that leaves the machine parks in the third column. Approve, reject or needs changes, and the poller wakes the agent again.

Two lines from the board protocol tell you more about the design than the rest. The role list includes a CTO and a CDO, marked as roles that don't exist yet and must never be assigned to, because "it would dead-end". And: a missing tool is not a handoff. Every agent runs in the same sandbox and hits the same wall, so an agent that can't do something parks the card for me and names the capability it needs. Fixing it adds a script on the host, never a role on the org chart.

Vera opens pull requests, and writes an RFC every Wednesday

Vera has her own copy of the Franz repo. The main mirror is read-only; beside it she has a worktree she can write to. In June she noticed the website had no llms.txt, the file that's supposed to tell AI assistants what a site is about, and predicted that adding one would lift referrals from ChatGPT and Gemini by a fifth. She wrote the route, wrote the test and opened the pull request, and it was merged on 17 June. Thirty days later the referrals hadn't gone up by a fifth. They'd fallen by more than half. She scored it refuted and wrote the lesson into the knowledge base: cheap hygiene, not a traffic lever. If you're adding an llms.txt to your site this month, you now know more than most people doing it.

Wednesdays are for research. Since June she has written 18 RFCs on channels I never had time to look at properly: the Microsoft Store, Setapp, Flathub, IndexNow, a Portuguese localisation pilot, AlternativeTo reviews, GitHub as a release surface, the Mac App Store. Two arrived this morning. Each one is a card on the board with the evidence attached, and each one is work that used to happen once in an eternity, the week before something forced it.

Sisi's side

Sisi's mornings start with the nightly QA report. She reads one file and relays it to my phone: what was tested, what passed, which pull requests the night opened. Her instructions forbid her from checking anything herself. If the file doesn't say it, she doesn't say it. That boundary is the one part two was about, seen from her end.

The rest of her day is the board. When I assign her a bug, she hands it to the dev station, the runner that also repairs tests at night, and the job comes back as a pull request with proof attached: a screenshot of the fixed screen for anything visual, the verifying command output for anything else. The same proof lands on the board card before I'm asked to review. She can boot a real Linux VM on the mini, install the latest release into it, check that the window paints and four other things, and tear it down. In July she restored a messaging recipe that a fix had dropped by accident, prepared a preview release, and merged two green dependency updates after I said yes in the morning digest.

None of that is glamorous, which is the point. It's the ops work that belonged to nobody when Franz had five people, and it now belongs to a fox emoji with a folder.

What Vera can't see

The first question any company asks about a setup like this is what the agent gets to see. So here's what Vera sees of my customers: nothing.

The snapshots in her folder are aggregates. No names, no mail addresses, no per-user rows. I actually checked while writing this, by searching every snapshot file for an @ sign. The only one in there is the address the collector logs in with, which is mine.

In July she asked a question I'd never asked: are people who run five or more WhatsApp accounts in Franz a real segment, or noise? She couldn't query the database. What she could do was ask for the product to compute one more aggregate, the same kind it already computes for my dashboard: how many accounts sit in that bucket, how many of them pay, how many were active in the last month. The answer came back the same day. The segment is real, big enough to matter, and it converts at several times the rate of everyone else.

Then she decided not to build a landing page for it, because search demand for that phrasing is tiny, and wrote that decision down too.

What reaches Vera when she asks about customers
Never leaves the product NamesMail addressesPer-user rowsA database connection
The aggregate, and only that how many accounts sit in the buckethow many of them payhow many were active in the last monthhow that compares with everyone else
Same numbers my own dashboard shows. Enough to find a segment worth caring about, and nothing that could identify a person in it.

You can give a marketing agent customer insight without giving it customers. Hand over the aggregates you already compute for yourself, and nothing else.

Sunday she rewrites her own rulebook

Every Sunday at nine, Vera scores every experiment that reached its date, then grades herself. The calibration playbook in her folder has 20 rules now, and every rule ends with the date and the miss that produced it. The first one came out of her first week: she proposed fixing wasted mobile ad spend that I'd already fixed three days earlier, because she'd read a seven-day average that straddled the fix. Rule one: check the days after a change before diagnosing from a window that contains it.

The score she publishes is quite unflattering. On metric predictions she's wrong seven times out of ten, and she's the one who writes that down, every Sunday, next to the sentence that confidence is low. Twice this summer the number got worse because she audited her own audit and found experiments she'd forgotten to count. Marketing usually doesn't get a number like that, because nobody wants to be the one writing it down. She does not mind.

Both agents also keep a dream diary. At three in the morning something writes a short entry about the day, with a haiku in it. Sisi's entry after the two dependency merges describes a drawing of a robot watering two sprouts. I'm not turning that off.

If your team wants to try this

The part that would transfer to your product isn't the model. It's the folder: a doctrine, a fence written in code, data that includes revenue, one board with a needs-my-decision column, and a Sunday where the agent grades itself. Every company already has most of that. It's spread across five people's heads instead of one folder.

Vera and Sisi are the half of my day I hand over. The other half I still do myself on this laptop, with agents next to me, and it's changed what a working day feels like. That's part six.

The series Daily Driver Seven parts on running a one-person software company where agents do most of the execution, and where the line still sits.

If you're a founder working out how much of this you can hand over, and where the line should sit, that's the work I do with teams. Let's talk.

Stefan Malzner
Stefan Malzner

Product designer in Vienna · founder of Franz. I write about product, design, and building software that lasts.

Get in touch

Let's talk.

Whether you're shipping a product or planning an event, the fastest way to reach me is email. I read everything that lands.