burgndy.ai
← All articles

September 28, 2026 · 8 min read

Comparing AI Coding Models in 2026: What the Benchmarks Don't Tell You

AI coding modelsLLM benchmarksStartup building

Every few weeks there's a new "best AI for coding" post, usually built around one leaderboard screenshot and a headline number. If you're actually building something on top of these models, rather than just reading about them, that number tells you less than it looks like it does. What actually matters splits into three separate questions — how a model performs on real engineering work, what it costs once you account for how these models are actually used, and whether the platform hosting it will even still offer it next month. They move independently, and most comparisons only ever look at the first one.

The benchmark that reflects real engineering work

Most public leaderboards still lean on SWE-Bench Verified, which by now is comparatively easy and fairly saturated. SWE-Bench Pro is the harder, newer version — real GitHub issues in production-grade codebases, not curated toy fixes. On it, Kimi K2.6 currently leads frontier models at 58.6%, just ahead of GPT-5.4 in extended-thinking mode at 57.7%, with Gemini 3.1 Pro at 54.2% and Claude Opus 4.6 at 53.4%.

Worth sitting with, though: on the Artificial Analysis Intelligence Index, a broader measure of general capability, GPT-5.5 leads at 60 against K2.6's 54. Kimi's K2 family is a genuine coding specialist, not an all-round frontier model — leading a hard, coding-specific benchmark and leading general intelligence are two different claims, and it's worth checking which one a ranking is actually making before treating "best at coding" as a settled question.

The pricing gap most comparisons skip

Headline per-token prices vary a lot, and on their own they say very little about what a real build costs:

  • Kimi K2.6 / K2.7 Code: $0.95 input / $4.00 output per million tokens, with cache-hit input priced down at $0.16.
  • Qwen3-Coder-Next (Alibaba, reachable via OpenRouter): $0.12 input / $0.80 output per million — one of the cheapest models actually purpose-built for agentic coding rather than general chat.
  • GLM-4.6: $0.60 input / $2.20 output per million, 200K context, with a real reputation for holding up inside agentic tools like Claude Code and Cline, not just scoring well in isolation.
  • DeepSeek V4-Flash: $0.14 input / $0.28 output per million — currently the cost floor among genuinely capable coding models.
  • Gemini 3.x Flash: $0.75 input / $3.75 output per million at current promotional pricing, with cached input as low as $0.075.
  • Claude Sonnet 5: $3 input / $15 output per million — priced as a frontier agentic coder, not a budget pick.

None of those numbers, by themselves, tell you what a real build will actually cost. That depends on how the model gets used, turn after turn.

Why the cheaper model can end up costing more

A coding agent doesn't send one prompt and stop — it resends the whole growing conversation on every turn, file contents and all, and most providers give a discount on whichever part of that history a previous turn already sent (a "cache hit"). How deep that discount goes varies far more between providers than the sticker price does: Gemini discounts cached tokens down to roughly 10% of its base rate, while Groq's own published discount on cached tokens is a flat 50% off. On a long coding session where most of what's sent on any given turn is repeated context rather than new content, that gap in caching discount can matter more than which model's price looked lower on the pricing page. It's an easy thing to miss if the only thing being compared is the two numbers at the top of each vendor's site.

A good reminder that this information goes stale fast

Kimi K2 really was available on GroqCloud's self-serve platform in late 2025 — real launch post, real pricing, real availability — and it's still cited as current by plenty of articles that still rank well today. Checking Groq's own live model documentation directly, as of this writing, shows it's gone: Groq's self-serve catalog is down to two models, GPT-OSS 120B and GPT-OSS 20B, with Llama 3.1 and 3.3 moved behind an enterprise-only conversation. That's not a knock on Groq — providers add and drop models on their own schedule — but it's a real reminder that "where a model is actually available" is one of the fastest-moving facts in this whole space, and it's worth checking a provider's current docs directly rather than trusting a search result, however recent it looks.

What actually matters when you're picking one

A handful of honest questions get you further than any single leaderboard number: Is the benchmark you're looking at testing realistic engineering work, or an easier one that's mostly saturated by now? Is the ranking really about coding, or about general intelligence — because those two don't always agree? What's the real cache-hit ratio your own workload will actually see, and what does each provider discount that down to, not just its advertised base rate? And is the model still on the platform you're planning to build against, confirmed against that platform's own current documentation rather than a blog post from a few months ago?

There's no single right answer to all four at once — which is exactly why a coding pipeline built to route between models, instead of betting everything on whichever one happens to be winning this month's leaderboard, tends to hold up better as this market keeps moving as fast as it currently is.