burgndy.ai
← All articles

September 23, 2026 · 6 min read

Vibe Coding's Security Problem: What the Data Actually Shows

Vibe codingAI coding agentsSoftware security

Andrej Karpathy coined "vibe coding" in a February 2025 tweet, describing a specific and fairly narrow thing: fully giving in to an AI coding tool for a throwaway project, to the point of not reading the code it produces. "I just see stuff, say stuff, run stuff, and copy-paste stuff, and it mostly works," as he put it. He was explicit that this was for weekend hacks and low-stakes experiments, not production software. The term has since been applied far more broadly — including to real, funded startups shipping to real users — and that broadening is where a real, measurable problem shows up.

The trust gap

Recent industry surveys put a number on something most developers would probably admit if asked directly: 96% say they don't fully trust that AI-generated code is even functionally correct. Only 48% actually review it before committing. That's fewer than half. The majority of developers are, by their own account, shipping code they don't trust without checking it first.

What that gap looks like in practice

That's not abstract distrust — independent security testing backs it up with real failure rates. In one study, 45% of AI-generated code samples failed basic OWASP Top 10 checks, the industry-standard list of the most common and most dangerous web vulnerabilities. Thirty-eight percent contained at least one real security flaw: hardcoded secrets, broken authentication, straightforward injection vulnerabilities. Separately, AI-co-authored code was measured at nearly 2.74 times the security vulnerability rate of code written by a human alone.

It's accelerating, not leveling off

Georgia Tech's Vibe Security Radar, a project that tracks confirmed CVEs tied to vibe-coded software, logged 35 in March 2026 alone — up from 6 in January of the same year. That's close to a six-fold jump in two months. Adoption of AI coding tools is growing just as fast on the other side of the ledger: 90% of developers now regularly use at least one AI coding tool at work, and over half use one daily. Adoption is close to vertical. Review discipline is flat. That gap between the two curves, not the tools themselves, is the actual source of the risk.

"Runs" is not the same test as "safe"

A lot of the confusion in this conversation comes down to conflating two different bars. AI coding tools are often genuinely good at producing code that runs — the button works, the form submits, the page renders. Whether that same code safely handles a malformed input, an unauthenticated request, or a database query built from raw user text is a completely separate question, and it's the one the OWASP and CVE numbers above are actually measuring. "It worked when I clicked it" and "this won't leak a user's data in six months" are not the same test, and treating them as interchangeable is where most of these incidents start.

What actually helps

  • Treat AI-written code like a pull request from a developer you've never worked with before — read it before it merges, every time, especially anything touching authentication, payments, or user data.
  • Run an actual security check against the OWASP Top 10 before shipping anything that's reachable from the internet, rather than assuming a working demo means it's safe.
  • Keep a real mental (or written) line between "I'm testing whether this idea works" and "real users or real money are now involved" — the review bar should change the moment you cross it.

None of this is an argument against using AI to write code — the speed gains are real and not going away. It's an argument for closing the specific, measured gap between how much these tools are trusted and how often they're actually checked, which right now is the real source of the breaches making headlines, not the underlying models themselves.