← All writing
Daily Driver 7 / 7

My AI team costs €400 a month. One file added €500.

The whole setup runs on two subscriptions from two companies, one building and one checking. In September a single planning file burnt through a week of tokens in a few hours, and I paid €500 extra to finish the release on time.

Sep 28, 2026 7 min read Original

September was a statistical outlier, and I'm making sure it stays one. In the middle of a release I ran out of my weekly Codex allowance, on the Pro plan, a few hours into the week.

The cause was simple once I looked at it. One of the planning files for the release had grown to 1.5 MB of markdown, and the agents kept pulling it into their context. Imagine a team where everybody has to read the complete company wiki out loud every morning before they're allowed to touch a ticket. That was my release for a few hours, and it burnt through a whole week of tokens.

I could have waited for the allowance to reset, but the release was going well and I wanted to ship it on time with the same setup. So I did what you do at a slot machine and inserted some coins. €500 in total over September, and about €75 of it is still in the machine as I write this.

The bill

In part six I promised the bill for all of this, line by line. It's a lot shorter than six posts about agents would suggest.

The bill
Franz · AI team · monthly Claude€180Codex€220 Baseline€400 OpenRouter, experiments€5-10Mistral, support inbox€1-2 September, coins inserted€500 Still in the machine~€75
A normal month is a bit over €400. September was the outlier.

The baseline is two subscriptions: Claude for €180 a month and Codex for €220. On top of that I spend €5 to €10 a month on OpenRouter to try out new models, prompts and workflows before any of them get near Franz, and the support inbox runs on Mistral for €1 to €2.

Until the summer the bill looked different. I had two Claude Max plans, which I maxed out regularly, and a smaller Codex plan for OpenClaw. In September I switched to one large plan from each company. The baseline got a bit cheaper, but the bigger change is in how the work gets checked.

A release manager from one company, engineers from another

Codex is my release manager now, and Claude Code does the engineering. The prompt I start a release with is a lot longer in reality, but stripped down it's pretty much this:

You're the release manager for Franz 6.x. These features, fixes and improvements go in: [the planning docs]. Split them into meaningful work packages and hand each one to Claude Code to build. Your job is to make sure all of it gets verified and tested afterwards. Close the release with a retro, so the next release manager can pick up where you stopped. The release is scheduled in two weeks. Go for gold!

Each package goes to its own Claude Code session. When one comes back, Codex runs the tests and reads what changed, and anything that doesn't hold up goes back to Claude with a note on what's wrong. Only then does the package go into the release.

One release
Codex splits the release
Claude Code builds a package
Claude Code builds a package
Claude Code builds a package
Codex checks every package Doesn't hold up? Back to Claude Code.
Release
Retro, with my input, read first by the next release manager
One release manager, several engineers, one check everything has to pass. How many Claude Code sessions run depends on the release.

Cutting a release into packages that don't step on each other's toes, and deciding if one is really finished, are judgment calls. That's why this needs a model and not a script.

What surprised me is how much the quality went up once the checking happened across companies. In part six I mentioned that the model repairing a failing test at night comes from a different family than the one that built the code. Now the whole release works like that. Two models from different companies look at the same code in quite different ways, and Codex finds a ton of things before they ever reach a customer: bugs, edge cases nobody thought of, things that pass every test and still break once you click through the app.

At the critical checkpoints Codex also tests like a user, with end-to-end tests and computer use, where the agent sits in front of the app and clicks through it. That's the expensive part token wise, and it has found so many issues, in new code and in parts of the app that had been out there for a while, that it's worth every coin.

The retro

The last line of the prompt is the one I'd copy first. Every release ends with a retro for the next release manager. The release manager drafts it, and I add my side, as a lot of what went wrong is only visible from where I sit. In the retro for Franz 6.9, the release manager's own assessment started like this:

The engineering and final artifact verification were stronger than my orchestration.

So an agent wrote its own performance review and graded itself below the code it was managing. Fair enough. It also listed what it had got wrong, including the moment it told me an app had been closed while I was still looking at it, which is the kind of thing you'd hear from a colleague exactly once.

The most useful part was about tokens: where they went, what was wasted and what to do differently. That turned into a short guide every new release manager reads before it starts. One of its rules says the file that holds the current state of a release should stay around 6 KB.

The file that sent me to the coin slot was 1.5 MB.

Planning file, by size
6 KBWhat the guide says now
1.5 MBThe file in September
About 250 times the target. Areas are to scale.

This is the self-improving loop from part one, one level higher. Every release manager starts with the mistakes of the previous one written down in front of it, and that's also how I'm getting rid of the extra coins.

What's still on my desk

None of this means I've handed over the technical side of Franz. I care about the architecture, about how things get built and what they're built on, as much as I did before. What changed is my process. I used to do most of it myself, step by step. Now I spend that time planning, reviewing and pushing back on what comes back.

A few things are also mine by rule, written into the release manager's instructions: every change a customer will see, anything that touches privacy or security, anything that costs money, and publishing. The release manager has to ask before it crosses any of those.

The other direction needed a rule too. At some point I had to write down that a finished work package isn't a reason to stop. It had done a good job, reported back very nicely and then waited for me, while the next package could have started long ago. With a colleague I'd have been happy about the update. It's in the rules now.

If your team wants to try this

Every team I've worked with has a backlog that's full to bursting, and there's nothing bad about that. The backlog is where features go to die, and most of the ideas in there aren't bad. They end up there for two reasons. Some simply never get the time next to the work that's already planned. Others are a product call: somebody decided, often for good reasons, that an idea doesn't fit the strategy right now, or isn't worth the risk.

Now picture a setup like this with the team you already have, same people. Every developer gets a release manager and a few engineers from a different company, and nothing reaches a human before a model that didn't write it has checked it. Your people spend their days on the decisions, the architecture and the reviews. Every release ends with a retro, drafted by the agent and completed by the people who were there.

That helps with both reasons. The team has time for the meaningful work again. And the product calls get easier, as trying an idea costs a lot less than it used to. A "not now" that was mostly about the risk can turn into a small experiment, and you decide on what actually happened instead of on a guess.

At around €400 a month per developer, that's a cheap way to find out. Just put a size limit on your planning files first. I learned that one for €500.

The series Daily Driver Seven parts on running a one-person software company where agents do most of the execution, and where the line still sits.

If you're a founder working out how much of this you can hand over, and where the line should sit, that's the work I do with teams. Let's talk.

Stefan Malzner
Stefan Malzner

Product designer in Vienna · founder of Franz. I write about product, design, and building software that lasts.

Get in touch

Let's talk.

Whether you're shipping a product or planning an event, the fastest way to reach me is email. I read everything that lands.