burgndy.ai
← All articles

September 19, 2026 · 6 min read

What Building a Real AI Coding Agent Actually Taught Us

AI coding agentAI app builderBurgundy AIEngineering

Most "AI builds a real app" claims are marketing copy with no numbers behind them. Before Studio (Burgndy.ai's freeform AI coding agent) shipped, we ran it against real prompts and measured what actually happened — cost, success rate, failure modes — independently verified rather than trusting the agent's own "I'm done" claim. Here's the real, sometimes unflattering, data.

The first real test: 8 simple apps, 8 successes

Round one was deliberately simple: 8 varied prompts (a waitlist form, a small CRUD note-taking app, that kind of thing) through a real agentic loop with real tool access — writing files, running shell commands, reading its own output. A separate harness independently curled the running server afterward rather than trusting the agent's self-report. Result: 8 for 8, genuinely running servers responding with real HTTP 200s, averaging 4-5 turns, about 66.5 seconds, and roughly $0.13 in real API cost per app.

Then we made it harder, and the cost curve got real

Round two moved to harder prompts: real file-backed persistence, multi-page routing, a REST API driving a frontend, file uploads, and a login/session/password-hashing flow. 5 of 6 passed with substantive, multi-assertion self-written verification (not rubber-stamp checks). The 6th — the login/auth prompt — hit the turn budget and never finished, after burning $1.38, the most expensive attempt of the six. Authentication turned out to be a meaningfully harder category than everything else we'd tested, and we said so plainly rather than quietly dropping that data point.

Why we built a safety net for exactly one category first

Rather than just giving the agent a bigger turn budget and hoping, we gave it something more useful for the one category that had actually failed: a real, pre-tested reference implementation of email/password auth it's instructed to adapt rather than generate from scratch — the same underlying idea as a senior engineer handing a junior one a known-good pattern instead of leaving them to reinvent it. Built and wired in for both the mobile (Flutter) and web (Node) generation paths.

The honest result of testing that fix: genuinely mixed

We re-ran the exact failed scenario with and without the reference pattern available, on both platforms. On mobile, the agent called the reference tool, then built a custom auth system anyway — same result, just more expensive. On web, both versions succeeded this time (likely because other prompt improvements had already fixed the root issue) — the reference version was cheaper and faster, but the baseline wrote a more thorough self-check and finished cleanly on its own. We're not spinning this into a win because the plan predicted one; it's a real, mixed result, and it's still true today.

What actually moved the needle

Root-causing a real failure did more than the reference-pattern fix. A later hard test — combined auth, payments, reviews, and a role-aware dashboard in one prompt — failed once at the very edge of its turn budget. Digging into the actual event log (not guessing) showed the coding step had deferred writing any tests until it was almost out of turns, then got stuck on a known Flutter-testing pitfall. Two small, targeted prompt fixes later — verify early, and a specific warning about that exact pitfall — the identical prompt succeeded cleanly, with real tests written incrementally throughout the run instead of crammed in at the end.

Why we're publishing the messy version

It would be easy to only publish the 8-for-8 number. We're publishing the $1.38 failure, the mixed A/B result, and the actual root-cause investigation too, because that's what real engineering on a system like this actually looks like — and because a newly launched product asking for your trust should show its work, not just its highlight reel.