How to Verify AI-Generated Code Before Merging It to Production

··12 min read
How to Verify AI-Generated Code Before Merging It to Production

Last quarter, a developer on my team merged a slick-looking pull request generated almost entirely by an AI assistant. It passed CI. It looked idiomatic. It even had comments. Two weeks later we traced a nasty data leak back to a single line where the AI had silently swapped a parameterized query for string concatenation. The tests never caught it because the AI also wrote the tests, and those tests asserted the wrong behavior with total confidence.

That is the trap. AI-generated code is not usually wrong in the obvious way a junior developer is wrong. It is wrong in the plausible way, the confident way, the way that survives a quick skim. A 2024 study from Stanford found that developers using AI assistants wrote code with more security vulnerabilities than those who did not, yet felt more confident their code was secure. That confidence gap is exactly where production incidents live.

This guide is about closing that gap. I'll walk you through a practical, repeatable process to verify AI-generated code before it hits your main branch, including a worked review example, a comparison of verification approaches, and the specific tooling I keep in my own pipeline. None of it requires distrust of AI. It just requires the same discipline you'd apply to any code from a fast, tireless contributor who has never once run the thing they wrote.

Key Takeaways
  • Never trust AI-authored tests to validate AI-authored code. Separate the two, or write the assertions yourself.
  • Run static analysis and dependency scans on every AI diff, treating the AI like an untrusted external contributor.
  • Read the diff line by line for the "plausible wrong" patterns: silent security downgrades, hallucinated APIs, and off-by-one logic.
  • Reproduce the behavior manually at least once before merging. AI cannot tell you what it never executed.
  • Pin and audit any dependencies the AI introduces, since suggested packages are a growing malware vector.
  • Codify the process into a checklist so verification does not depend on who happens to be reviewing that day.

Why AI-Generated Code Needs a Different Review Standard

Human-written code fails in patterns you learn to anticipate. A tired engineer forgets an edge case. A rushed one skips error handling. You develop instincts for where people cut corners.

AI-generated code fails differently. Large language models produce statistically likely code, not correct code. When the most common pattern in the training data happens to be insecure, the model reproduces it fluently. Three failure modes show up over and over:

  • Hallucinated APIs. The model invents a method that sounds right (user.hasPermission()) but does not exist, or calls a real method with imaginary parameters.
  • Silent security downgrades. It quietly replaces a safe construct with an unsafe but simpler one, like string-interpolated SQL or disabled TLS verification.
  • Confident fabrication. It writes documentation and tests that describe correct behavior while the actual code does something else.

The extra danger is speed. AI lets a single developer generate hundreds of lines in minutes, which overwhelms the human review capacity that would normally catch these issues. You cannot fix a volume problem with the same manual habits you used for hand-written code. You need a layered process, similar to how you'd vet open-source software before adding it to your stack.

Step 1: Read the Diff Like You Don't Trust It

Before any tool runs, read the code. Not the AI's explanation of the code, the code. Explanations are generated text too, and they will happily rationalize a bug.

Here is my line-by-line pass, in order:

  1. Verify every API and method actually exists. If you don't recognize a call, look it up in the real documentation, not the AI's memory. Hallucinated methods are the fastest thing to catch and the easiest to miss on a skim.
  2. Trace the data flow for user input. Follow anything that comes from outside the system. Does it get sanitized? Parameterized? Escaped at the right boundary?
  3. Check every branch and boundary. Off-by-one errors, empty-array handling, null checks, and the always-false condition are classic AI slips.
  4. Question anything that got "simpler". If the AI refactored existing safe code into something cleaner, ask what it removed. Simplicity often means a check disappeared.
  5. Look for secrets and hardcoded values. Models frequently inline API keys, default passwords, or debug flags they saw in training data.

A Worked Example

Say you ask an assistant to write a login endpoint. It returns 40 lines that look great. On the line-by-line pass you find three issues that CI would never flag:

  • Line 12: SELECT * FROM users WHERE email = '" + email + "'" — string concatenation instead of a parameterized query. A SQL injection, introduced silently.
  • Line 23: password comparison uses == on plaintext, because the AI "forgot" the hashing step it described in its own comment on line 20.
  • Line 31: on failure it returns the message "No user with that email", leaking which accounts exist to an attacker enumerating the endpoint.

Three vulnerabilities in 40 lines, each individually plausible, none flagged by the tests the AI wrote alongside it. This is not an unusual result. It is the median result when you look carefully.

Step 2: Run Automated Analysis Before You Approve

Human eyes catch logic and context. Machines catch scale. Treat every AI diff as if it came from an anonymous external contributor and route it through the same gates.

Your minimum automated stack:

  • Static application security testing (SAST). Tools like Semgrep, CodeQL, or Bandit flag injection patterns, unsafe deserialization, and weak crypto. Run these on the diff, not just the whole repo, so noise stays low.
  • Dependency and supply-chain scanning. npm audit, pip-audit, or Snyk catch known-vulnerable packages the AI may have pulled in.
  • Linters and type checkers. TypeScript's compiler, mypy, or a strict ESLint config catch hallucinated signatures and type mismatches fast.
  • Secret scanning. gitleaks or trufflehog catch the hardcoded keys from Step 1 that slipped past you.

Wire these into your pre-merge checks so no diff, human or AI, reaches production unscanned. If you run web properties, the same defensive posture belongs at the perimeter too, which is why teams often pair code-level scanning with runtime protection like SiteGuard Pro or a full WordPress security plugin stack.

Step 3: Never Let AI Grade Its Own Homework

This is the single highest-leverage rule in this article. The incident I opened with happened because the same model wrote both the code and the tests, and the tests asserted the buggy behavior as correct.

If a test suite was generated by AI, it validates that the code does what the AI thought it should do, not what your system actually requires. That is circular. Break the loop with one of these approaches:

  • Write the test assertions yourself, then let the AI fill in scaffolding. You own the definition of correct.
  • Use a different model or session to generate tests than the one that wrote the code, so they don't share the same blind spots.
  • Add adversarial tests manually: the empty input, the giant input, the malicious input, the concurrent request. AI rarely writes these unprompted.

Then reproduce the behavior by hand at least once. Run the endpoint. Trigger the error path. Watch the actual output. An LLM cannot tell you what happens at runtime because it never ran anything. You are the first entity in the chain that actually executes the code.

Verification Approaches Compared

Not every change deserves the same depth of review. A one-line copy tweak is not a new auth flow. Here's how the common approaches stack up so you can match effort to risk.

Approach Catches Hallucinations Catches Security Flaws Catches Logic Bugs Time Cost Best For
Skim + CI pass Rarely No Rarely Minutes Trivial, low-risk edits only
Line-by-line human read Yes Often Yes Moderate Business logic, refactors
SAST + dependency scan Partial Yes Partial Automated Every diff, always on
Independent test suite Partial Partial Yes High Critical paths, auth, payments
Manual reproduction Yes Yes Yes Moderate Anything touching production data

The right answer is almost never one row. For anything touching authentication, payments, or user data, I run all five. For a CSS tweak, a scan and a skim are fine. The failure is applying "skim + CI pass" to code that deserved the full stack.

Step 4: Audit Every Dependency the AI Suggests

AI assistants love to reach for packages. Sometimes those packages don't exist, which sounds harmless until you learn about slopsquatting: attackers register the exact package names LLMs commonly hallucinate, then wait for developers to install them.

Research in 2024 found that a significant fraction of AI-suggested package names were nonexistent, and that models hallucinate the same fake names repeatedly, which makes them predictable targets. Before you install anything an AI recommends:

  1. Confirm the package is real and canonical on the official registry, not a lookalike with a swapped character.
  2. Check download counts, last publish date, and maintainer history. A "popular" package published three days ago with 200 downloads is a red flag.
  3. Cover image: The Torch Graduate circuit board (bottom) by Chris Whytehead, licensed under BY-SA 3.0 via Openverse.

Recent Posts

View all →

Most Popular Software

View all →

Browse by Platform

View all →