- Researchers tested AI-generated patches on six CVEs with poor success rates
- Many fixes failed, altered behavior, or introduced new vulnerabilities
- Guidance improved outcomes, leading to FLAWED evaluation harness release
When using Generative Artificial Intelligence (GenAI) to fix vulnerabilities, security professionals are most of the time just robbing Peter to pay Paul, experts have warned.
Researchers from 1Passwords Off-by-1 Labs analyzed fixes proposed by two frontier models – ChatGPT 5.5 at “medium” effort, and Claude Opus 4.8 at “high” effort.
As an experiment, the researchers took six recently disclosed CVEs and produced 6,080 patches using two frontier, cyber-capable reasoning models. The results were underwhelming to say the least – of all the proposed patches, just a quarter (26%) fully resolved the issue.
Latest Videos FromTechRadar
FLAWED work?
This obviously leaves plenty to be desired, as half (49.3%) of the patches failed to fix at least one existing exploit path. A fifth (20.1%) fixed the original issue but changed application behavior, while 2.3% introduced new security issues. Funny enough, 2.2% failed to fix the vulnerability while also introducing additional exploit paths, as well.
Even among the patches that might be considered (26% of clean ones and 20.1% of those that changed app behavior), more than a third were fragile and not entirely addressing the underlying problem.
The researchers created an acronym for automated LLM patches: FLAWED (Fix-Like Artifacts With Embedded Defects), and warned against letting AI work without human oversight: “The expected value of a fully LLM-generated, non-human-reviewed patch is a net-negative by a considerable margin.”
Results drastically improved when the AI was given better context, the researchers further explained. Before working on any patch, human developers are usually given initial guidance. When AI is given proper guidance, its success rate rises to 65%. Incorrect guidance, on the other hand, drops the success rate down to 15.2%. The difference between humans and AI is that humans are better at catching misleading information and poor guidance.
This doesn’t mean developers will, or should, abandon AI. Worst case scenario is that developers will spend more time reviewing AI-generated fixes which could increase cognitive load and still end up being net negative. Therefore, the researchers released a patch evaluation harness called FLAWED, which organizations can now use to determine the effectiveness of their AI-generated fixes.
Via The Register
The best antivirus for all budgets
Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.
https://cdn.mos.cms.futurecdn.net/TaxPLZc75WiicpmgZNzWzL-2560-80.jpg
Source link




