Blog

AI Coding for Java Teams: Get the AI Speed Without Losing Control

By  
Roman Kałkowski
Roman Kałkowski
·
On Oct 6, 2026, 2:50:14 PM
·

AI assistants now write code faster than any of us can read it. Whether we can trust that code is a different question, and for a lot of Java teams the answer is still "not really".

If you build software with an AI assistant, you control two things: the prompt you write and the code you accept. What happens in between (which libraries get pulled in, how the code is structured, what gets tested) is mostly out of view. We call that space the AI control gap, and it's where a lot of enterprise risk is quietly piling up.

What is the AI control gap?

Put simply, it's the difference between what you asked the assistant to build and what it actually built. Architecture, dependencies, security decisions, tests: plenty of things end up in the codebase that nobody consciously chose. And the gap grows with speed. The faster the assistant writes, the less of its output a team can realistically review.

Why don't developers trust AI-generated code?

Almost everyone uses these tools now. Far fewer trust them. Google's 2025 DORA report found that 90% of technology professionals use AI at work, but also that AI adoption still goes hand in hand with less stable software delivery. In the 2025 Stack Overflow Developer Survey, 46% of developers said they distrust the accuracy of AI tools, and only 3% said they highly trust it.

The research gives them good reason:

  • Veracode's 2026 GenAI Code Security Report tested more than 100 models. In 44% of security-relevant coding tasks, the generated code still introduced a vulnerability, about the same as a year earlier.
  • CodeRabbit found that pull requests co-written with AI had roughly 1.7 times as many issues as human-only ones, including 75% more logic and correctness problems.
  • GitClear looked at 623 million code changes and found duplicated code up 81% since 2023, while refactoring dropped to just 3.8% of changes.
  • And 66% of developers say their biggest frustration is AI output that's "almost right, but not quite" (Stack Overflow). If you've reviewed an AI-generated pull request, you probably know the feeling.

Where does AI speed cost you control?

It tends to slip away in five places:

  1. Architecture. Assistants are good at adding code and bad at reshaping it, so the structure drifts a little with every prompt.
  2. Technology choices. In a split frontend/backend stack, the assistant is picking libraries in two languages and guessing at the API contract between them.
  3. Correctness. Code that compiles isn't necessarily code that works.
  4. Testing. When tests are slow or missing, "almost right" code sails through until a user finds the bug in production.
  5. Visibility. You see what went in and what came out, but not the reasoning in between.

Is AI-generated Java code less secure?

Right now, somewhat. In Veracode's 2026 tests, Java had the lowest security pass rate of any language, at 30%. Python managed 63%. The encouraging part is that Java was also the only language clearly getting better. Cross-site scripting was a weak spot everywhere: across all languages, the defenses held up just 15% of the time.

So should Java teams steer clear of AI? We don't think so. What these numbers show is that the setup around the assistant matters more than which assistant you pick.

What is control by construction?

Most advice on AI-generated code is about scanning it after it's written. Scanning matters, but by the time it flags something, the code already exists. Control by construction comes at the problem from the other side: you shape the development environment so the assistant has fewer ways to get things wrong in the first place.

As Jurka Rahikkala, CEO of Vaadin, puts it: "AI has made writing code cheap. It hasn't made trusting code cheap. The answer is to give the assistant less to guess: one language, one codebase, one compiler that objects before anything ships."

How do you keep control of AI-generated code in a Java web app?

Whatever stack you're on, six habits close most of the gap:

  1. Give the assistant fewer layers to juggle. Every extra language, repository or hand-written API contract is one more thing the model has to guess.
  2. Let the compiler catch mistakes first. When code is typed end to end, from the UI down to data access, whole categories of AI errors fail at build time instead of at runtime.
  3. Feed it current documentation. Models learn from older versions of APIs. An MCP server or agent skills that serve your framework's current version cut down on outdated or made-up method calls.
  4. Test inside the loop. Spec, generate, verify, refine, with tests fast enough to run on every change.
  5. Keep changes reviewable. A full-stack change that arrives as a single pull request is one a reviewer can actually read.
  6. Think in decades. Business apps usually outlive the model that helped write them, so build on foundations that will still be maintained.

How Vaadin fits in

This is the thinking behind Vaadin, our full-stack Java web framework:

  • One typed Java codebase. UI, business logic and data access live in one language, checked by one compiler and covered by one security model.
  • No API layer to hand-build. With Vaadin's server-side Java UI, the framework handles communication between browser and server, and your business logic stays on the server. [Scope check: Flow]
  • Up-to-date context for your assistant. The Vaadin MCP server and agent skills work with Claude, GitHub Copilot, Cursor, Windsurf and Codex.
  • Verification built in. Browserless UI tests run as plain JUnit tests, and Vaadin Copilot turns visual edits you make in the browser back into Java code.
  • Enterprise governance. Vaadin Enterprise Edition comes with 15 years of maintenance per major release, security and accessibility reviews, ISO certificates, and an air-gapped MCP server for restricted environments.

Does better context actually help?

We wanted to know too, so we measured it. VaadinBench is our open benchmark of real Vaadin tasks. On its own, Claude Opus 5 passed 50–67% of them. With Vaadin's agent skills or MCP server, it reached up to 100%. Claude Sonnet 5 went from 44% to 92% with agent skills alone, so the models that needed help most gained the most.

To be fair, the benchmark is small: five tasks, with 8 to 12 attempts per setup. That's why the code and results are public, so you can rerun it yourself.

Wrapping up

Speed from AI only pays off if your team can still explain, test and stand behind what gets built. The code gets written faster, but it's still yours to own. If you'd like to try this approach, start building with Vaadin and AI, or come see it live at Vaadin Create/26 in Barcelona on 27–28 October.

Frequently asked questions

What is the AI control gap?

It's the difference between what a developer asks an AI coding assistant to build and what the assistant actually builds: the architecture, libraries, security decisions and tests that nobody explicitly chose. The gap gets wider as assistants get faster, because code arrives faster than teams can review it.

Is AI-generated Java code secure?

Not by default. Veracode's 2026 GenAI Code Security Report found Java had the lowest security pass rate of any language tested, at 30%, and cross-site scripting defenses held up only 15% of the time. Java teams get safer results when they put structure around the assistant: typed code, current framework context, fast tests and reviewable changes.

How do enterprise teams keep control of AI-generated code?

They cut down the number of layers the assistant has to coordinate, rely on a compiler that checks the code end to end, give the assistant version-specific framework context through an MCP server or agent skills, test inside the loop, and keep each change small enough to review in one pull request. At Vaadin, we call this control by construction.

What is the best Java framework for AI-assisted web development?

Look for a framework that keeps the UI and backend in one typed language, gives AI assistants current API context, and makes automated UI tests fast. Vaadin is a full-stack Java web framework built this way, with an official MCP server and agent skills for Claude, GitHub Copilot, Cursor, Windsurf and Codex.

Is vibe coding safe for enterprise business apps?

Prompting an app into existence without reading the code is fine for prototypes. For business-critical systems that have to be secure and maintained for years, it's risky. Enterprise teams can keep most of the speed safely by pairing AI assistants with a typed stack, version-accurate context, automated tests and a human review of every change.

Does an MCP server make AI coding assistants more accurate?

It can. In VaadinBench, Vaadin's open benchmark, Claude Opus 5 passed 50–67% of real Vaadin tasks on its own and up to 100% with Vaadin's agent skills or MCP server. Weaker models gained the most. The benchmark is small and public, so teams can rerun it.

Roman Kałkowski
Roman Kałkowski
Roman spends his days at the intersection of product strategy and developer advocacy. With a background rooted in product management and marketing, he specializes in translating complex technical capabilities into clear human value. When he’s not helping developers build better web apps, he’s likely obsessing over user experience or finding new ways to bridge the gap between "how it works" and "why it matters.
Other posts by Roman Kałkowski