I have a lot of side projects that I vibe code with Claude and Codex. When I want to change something, I describe it, the agent builds it, I look at it, and it ships. It’s one of the most fun parts of my week.
I’d love to work on Obot the same way. I have plenty of ideas for small changes, like a tweak to a screen or a cleaner flow. But if I started pushing that kind of change into the Obot codebase, Darren, my co-founder and our Chief Architect, would flip his lid. And he’d have a point. My side projects don’t have customers running them in production. Obot does. A PR that “works” on my laptop can still take him longer to review and fix than it took me to make.
Try Obot today — the open-source MCP gateway. Self-hostable, MIT licensed, full audit trail at every tool call.
So when Tammuz Dubnov, founder and CTO of AutonomyAI, joined us on The Context this week to talk about what happens when non-developers start shipping code into real products, I was paying close attention. It basically describes the gap between how I build on my own and how I’m allowed to build at work.
Three ways to make things worse
Tammuz has watched this play out at a lot of companies. AutonomyAI works with more than 180 organizations, mostly 200 people and up, on big existing codebases. Some of those monorepos are twenty years old. He described three stages companies go through when they hand AI to their PMs and designers.
Stage one: better tickets.
PMs use ChatGPT to write PRDs. The tickets get longer, but they aren’t tied to the code. They’re full of hallucinations, and engineering pushes back. You spent all that effort making engineering faster, and the handoff to engineering got slower.
Stage two: prototypes.
PMs and designers start building HTML mockups so they can show what they want, not just describe it. Better, but the mockups don’t match the real product. Workflows get buried in them. Then in product review someone asks, “Why didn’t you build this? It’s in the mockup.”
Stage three: the pull request.
This is the stage I’d jump to if Darren let me. The builder points a coding agent at the repo and lets it run. It doesn’t know what should be reused or where things belong. The PR is huge. An engineer looks at it and says it could be 70% smaller. The builder says, “Who cares? It works.” We all know why it matters. Maintainability. Fewer bugs. Code a human can live with next year. But the person who opened the PR doesn’t know that, and the agent didn’t tell them. His summary was that all this AI makes the org louder. More noise, more friction. The goal is a good one, but it falls apart if you don’t set it up right.
The slide that stuck with me
Tammuz shared one slide that I keep thinking about. It was titled Why “just use Claude Code.” That’s how you get slide 2. (Slide 2 was the mess I just described.)
On one side, a generic coding agent in a PM’s hands:
Full repo and shell access on a laptop. No guardrails for a non-engineer. Secrets and env files in reach.
Starts cold every session. Doesn’t know your components, so it invents new ones.
No visual check. It compiles, but nobody knows if it looks right.
Quality depends on the prompt. Every PM gets a different standard, and your reviewers catch all of it.
On the other side, what AutonomyAI built:
An isolated workspace. No terminal. Engineering sets the guardrails.
A shared model of your codebase. It reuses your components first, and every PM is held to the same standard.
It renders and checks the UI before every PR.
Scoped, auditable PRs in your git. Nothing merges without approval.
Look at that list again, because the point isn’t really about PMs. Read it as a security person would. Shell access on a laptop, secrets in reach, no guardrails, results that change with every prompt, no consistent standard. That’s not only why you get sloppy PRs. It’s why agents are giving CISOs a headache right now.
Context is the product
The biggest idea in the episode is that the model isn’t the product. Tammuz said it plainly: they use the same models everyone else does. “If the models did a good job out of the box, we wouldn’t have a business.”
What they’ve built is what he calls a product harness. R&D teams have spent the last year building their own harness for coding agents: the rules, the skills, the MCP servers, all the context that makes Claude Code useful on a real codebase. Product teams never got one. So AutonomyAI builds it for them. It models your codebase with static analysis and dependency graphs, pulls out your design system and component library, and writes its own design and product docs. It connects to more than a thousand data sources, like Mixpanel, Linear, Jira, Figma and your call recordings, so the agent is working from real customer signal instead of a vague prompt.
One detail I loved: they tried fine-tuning on customers’ coding standards, and the engineering teams begged them to stop. “Our coding standards aren’t good. Please do better.” Codebases keep changing, so the context has to keep changing too. It’s a living layer. It learns from reviewer feedback and even writes its own linters so it doesn’t make the same mistake twice. The result: 86% of PRs opened through their platform get merged.
The 14% that don’t are interesting too. Tammuz said they usually come from overconfident PMs, often ones with a technical background, who override the safeguards and tell the agent to go change the backend as well. The agent assumes they know what they’re doing.
Darren’s pushback, and why we’re both right
This is where the episode got fun. Darren pushed back hard, in the way only someone who has maintained real systems for twenty years can.
His argument: roles exist for a reason. A PM doesn’t care about every little detail of how an app gets built, and that’s why they hired engineers. Even a “simple” UI change is probably tied to a design system, and that scaffolding has to be in place first. Somebody has to own the code after it merges. If engineering can move faster, maybe engineering should just move faster.
Tammuz’s reply was the one he uses in every sales call. Great, you made engineering faster. Do you still have a backlog? Are you responding to the market quickly? Engineers work in code, which is the source of truth. PMs work in tickets and designers work in Figma, and both drift away from the code. All the review loops we built exist to fix that drift, and they’re the slow part.
My take: I think the PM wins this one in the long run. Business requirements always win. But Darren’s point about scaffolding is the part that decides whether it goes well or goes badly. Tammuz’s pitch works because engineering sets the guardrails and everyone else works inside them. Take away the guardrails and you’re back to “just use Claude Code.”
Where Obot comes in
This is the same argument we make at Obot, just one layer down.
Early in the talk Tammuz said something that could have come from one of our customer calls. MCP is great because it brings the right context into the agent’s work. But not everyone is technical enough to know which MCP servers to set up, and a lot of setups don’t have the guardrails an enterprise needs to feel safe.
That’s the gap we’re working on. AutonomyAI governs what a PM’s agent does to the codebase. Obot governs what every agent in the company can reach: which MCP servers, skills and tools are approved, who gets access to which systems, which work runs in an isolated sandbox instead of on someone’s laptop, and a full audit trail of what each agent actually did. Engineering and IT set the guardrails once. Everyone else gets the right context without having to set it up themselves.
And it matters more with what AutonomyAI launched the day we recorded. They call it the autonomous product delivery layer. It’s a loop: agents take in signals from customers, the market and your execs, prioritize against the codebase, prototype options, open a mergeable PR, look at how the change performed after deploy, and score it against the original goal. It runs on a schedule. Tammuz described the human as “not in the loop, but on the loop.”
I think that’s where a lot of work is heading. And when agents own whole areas of your product and run on their own, the question stops being “is this PR good?” It becomes: what can these agents reach, who approved it, and can I prove what they did? Their slide says “engineering sets the guardrails.” To be fair, AutonomyAI also offers bring-your-own-model and on-prem for customers with stricter security needs. But the more agents run on a loop without a person starting each task, the more those guardrails need to be enforced by the platform rather than written into a prompt.
That’s the layer we’re building at Obot. Whether you think the PMs or the engineers should win, both of them need it.
Watch the episode
Thanks again to Tammuz for joining us, and congrats to the AutonomyAI team on the launch. We went almost an hour on what’s normally a 25-minute show, which tells you how much there was to argue about. The full episode is on the Agentic AI Foundation YouTube channel. Watch it for the demo alone: seven versions of one component, rendered live in a sandbox running their whole infrastructure, before a small PR was ever opened.