AI isn’t magic. It’s orchestration.
I run a multi-model AI team the way a bank runs risk. One person signs. Nobody grades their own work.
I’ve used AI 20+ hours a week for four years. I run coding agents every day, three or four sessions in focus and more in the background.
Here’s what I know. The magic isn’t the model. Everyone can buy the same models. Somebody still has to sign their name to the work. That somebody ain’t the model.
So I run AI the way a bank runs risk. I spent eight years in banking. The business owns the risk. A second line challenges it. Internal audit tests both. Nobody marks their own homework.
Who does what
I’m the orchestrator. I set the scope, write the brief and make the final call. Once I approve a brief, the agents work inside it. They come back when an answer would change the product or cross a commercial line. Not before.
Strategy is my challenge partner. One session holds my rulings and pulls the record. I jam with it, and it pushes back before I commit to a brief. It never writes the brief and never touches code. A persuasive recommendation isn’t a decision just because an agent wrote it down.
The implementer builds. First line. It can’t certify its own work.
Control checks. Second line. Every claim gets verified before it reaches me. “Tested and passing” is a claim. Control goes and looks: runs the checks, reads the change, uses the build the way a customer would. It’s read-only on the product, so it can’t quietly fix what it judges.
A fresh auditor stands at the gate. Third line. A new session from outside the working conversation, called at a merge or a release. It checks what tests can’t: does this do what the brief said, on the exact version, with the evidence to show it. A pass isn’t permission. The release call stays mine. The verdict goes on the record. If I challenge a no, the disagreement and the resolution go on the record.
It says no. Five of the first nine audits on my current build came back no-go. One piece of work took three audits to clear.
Some seats run on models from different AI companies. That helps, because they miss different things. But a different logo isn’t an independent review. A session that didn’t do the work is.
How it runs
It isn’t a pipeline. It’s loops, and I’m in them. I review, I redirect, I decide. What I don’t do is carry messages, or check a claim an agent can check.
Build and check go back and forth until the evidence holds. A no-go goes back to the builder with the evidence.
The sessions hand off to each other. I’m not the clipboard. No “got it”. Silence is never approval.
Every request cites the decision that authorized it. No citation, and the agent may read and report. Nothing else. There are 200+ written rulings on my current build. An agent quoting one from memory is wrong until it cites it.
Everything reads from one record. Rulings, status, evidence, one home each. That’s the company brain. Agents trust the record. Not their memory. Not mine.
Why I’m strict about it
AI reports “done, tested, correct” in the same tone whether it is or not.
I’ve watched one session spawn 74 agents, lose its rules when its context compacted, and burn a week of credits in a few hours. Nobody was there to sign their name. Me included.
The rules don’t live in a conversation. They live in the record, and every session reads them again. When a session starts repeating a claim I corrected, or reopening a question I settled, I don’t argue. I replace it. The new one starts from the record.
What it gets me
Three of us run CoBuy. No engineering department.
- 350%+ faster shipping in six months. We’re faster than that now.
- 90%+ off customer response times.
- One platform being built to replace every product we have, behind 2,393 tests and five release gates. A release can’t close without a go from a session that didn’t build it.
- 19 of 190 research records, one in ten, caught after the first filter passed them as clean.
- A working product for garage-door companies, a trade I’d never worked in.
- Paid client work for other businesses, run the same way.
The people still matter
None of this is autonomous. AI has no judgment and no taste. The lazy path is to let it loop, or to accept output that looks legit. Orchestration, review and governance stay with humans. So does managing the resource: cost, efficiency, quality.
I’ve led 15+ engineers over ten years, on five continents, and ran a 180+ applicant search to find my technical co-founder. The people I hire now have to understand AI and its limits. That isn’t optional. One engineer who reported to me called it “a culture of growth and collaboration”. Their words, not mine.
The seats, the gates and the record aren’t specific to my company. I run them on my own products and on client work.
A bank putting AI to work, or a fund doing it across a portfolio, needs the same three things: a named owner, an independent check and a written record. That’s what I build.
People call it magic when they can’t see the controls.
I built the controls.