Written by Gowtham Raj, Director at TartLabs, who leads custom software development engagements where AI coding tools are now part of everyday delivery.
The Short Answer
Not by default. In Veracode's July 2026 report, AI models produced secure code in only 56% of test cases, while their code compiled and passed syntax checks almost every time. Code that runs is not code that is safe. The data does not say AI code is worse than human code, because none of these studies ran a matched human baseline. It says the model will not catch security problems for you, so the review process has to.
Key Takeaways
- Security pass rate is 56%, and syntax is near 100% (Veracode GenAI Code Security Report, 28 July 2026, tested on raw models with no guardrails).
- The 2025 edition found 45% of test cases introduced a vulnerability, and that larger models did not perform significantly better than smaller ones.
- IOActive's April 2026 whitepaper found 31.6% of AI-generated samples fully exploitable across 27 tools and 730 prompts. Security-aware wrappers and system prompts improved results by up to 25 percentage points.
- Georgia Tech's Vibe Security Radar had linked 74 CVEs to AI coding tools by March 2026, up from 6 in January. The researcher estimates the true count is 5 to 10 times higher.
- What works: treat AI output as untrusted input, scan every change automatically, and keep a human reviewer on authentication, authorization and data handling. The controls are our view, not a standard.
What the Data Says
Four sources, four methods. They agree on direction, so read them side by side rather than adding them up.
| Source | Finding | What it measures |
|---|---|---|
| Veracode, July 2026 | 56% average security pass rate; GPT-5.5 highest at 68%, Qwen3.7-max lowest at 50% | Raw models on standardized security tasks, no agents or guardrails |
| Veracode, July 2025 | 45% of cases introduced a vulnerability; Java failed over 70% of the time | 80 curated tasks across 100+ LLMs |
| IOActive, April 2026 | 59% average security performance; 31.6% of samples fully exploitable | 27 tools, 730 prompts, nearly 20,000 samples |
| Georgia Tech SSLab via Infosecurity Magazine, 26 March 2026 | 74 confirmed CVEs tied to AI coding tools; 6 in January, 15 in February, 35 in March | Real public advisories traced back through Git history |
Two cautions. The first three rows are lab tests on tasks chosen because they can go wrong, so they show what a model does when nobody is steering it, not what ships to production. The fourth row counts real vulnerabilities but only those it can attribute, and the researcher estimates the true number is five to ten times higher. Veracode and IOActive both sell security products, which is worth knowing when you read their conclusions.
Three Patterns in the Numbers
1. Models learned syntax, not security
Veracode reports syntax correctness at roughly 100% in 2026, while the security pass rate sits at 56%. In 2025 it already noted that security was not improving as syntax did. Newer models are not fixing this on their own. The 2026 edition found reasoning models averaged 56% against 51% for non-reasoning models, a gain, but nowhere near a safe baseline.
2. Size and price do not buy safety
Veracode's 2025 report states that larger models do not perform significantly better than smaller ones. In 2026 the gap between the top model (68%) and the bottom (50%) was 18 points, and even the best still failed almost one test in three. Choosing a premium model is not a security control.
3. Language and code type matter
Veracode's 2026 edition shows Python passing 63% of the time and Java 30%. IOActive found infrastructure and DevOps code had vulnerability rates between 70% and 97%. If your team generates Terraform, Kubernetes manifests or Java services with AI, those are the places to look first.
Where AI Code Goes Wrong Most
From the 2025 Veracode data, models failed to defend against cross-site scripting in 86% of relevant cases and log injection in 88%. These are the flaws that need context a prompt rarely contains: where data comes from, who may see it, what gets logged. The model writes the happy path and leaves out the defence.
A second risk sits outside the code itself. The Cloud Security Alliance's April 2026 note, citing a March 2025 study of 576,000 code samples from 16 models, reports that about 19.7% of AI-suggested packages in Python and JavaScript did not exist. Open-source models hallucinated about 22% of the time, commercial models about 5%. We have not read the original study, so treat the figure as secondhand. The danger is that an attacker can register a name a model keeps inventing.
Controls That Close the Gap
This is our view of a practical minimum. It is not a certification checklist.
- Treat AI output as untrusted input. It gets the same scrutiny as a pull request from a new contractor, whoever pressed the button.
- Scan every change automatically. Run static analysis, dependency checks and secret scanning in CI so no AI-authored change reaches main unscanned.
- Pin and verify dependencies. Check that any package the model suggests exists, is maintained and is the one you meant, before it lands in a lockfile.
- Put humans on the risky paths. Authentication, authorization, payments, personal data and anything that parses user input should have a named reviewer who understands the threat, not just the diff.
- Add security to the prompt and the wrapper. IOActive found security-aware wrappers and system prompts lifted results by up to 25 percentage points. That helps, but a prompt is not a guarantee, so it supplements the controls above.
- Test for abuse, not just function. Add negative tests: malformed input, missing permissions, oversized payloads.
- Keep an AI-code log. Note which changes were AI-assisted so incidents and audits can trace them. Georgia Tech can only count what leaves a signature.
Team size and review capacity change how much of this you can do. Our piece on how many engineers you need now that AI writes the code covers that side of the question, and the broader planning view is in our AI strategy for software companies guide.
A Self-Check You Can Run This Week
- Do we know which repositories receive AI-generated changes?
- Does every pull request run static analysis and secret scanning before merge?
- Can a reviewer see which lines were AI-assisted?
- Do we verify every new dependency a model proposes?
- Is there a named owner for authentication and authorization code?
- Have we tested abuse cases on the last AI-built feature?
If three or more answers are "no", the risk is in your process, not in a particular model.
The Bottom Line
AI coding tools have solved syntax and left security largely where it was. Veracode's 2026 pass rate of 56%, IOActive's 31.6% exploitable samples and Georgia Tech's rising CVE count all point the same way: generated code is fast, and it is not self-securing.
Build the checks around the model, not inside it. Scan every change, verify every dependency, and keep a person on the code that guards your data. If you want a second opinion on your pipeline, get in touch, read our AI strategy for software companies guide, or see how a dedicated team can review AI-built code alongside your engineers.




