onlyTrustedInfo.comonlyTrustedInfo.comonlyTrustedInfo.com
Font ResizerAa
  • News
  • Finance
  • Sports
  • Life
  • Entertainment
  • Tech
Reading: AI models still struggle to debug software, Microsoft study shows
Share
onlyTrustedInfo.comonlyTrustedInfo.com
Font ResizerAa
  • News
  • Finance
  • Sports
  • Life
  • Entertainment
  • Tech
Search
  • News
  • Finance
  • Sports
  • Life
  • Entertainment
  • Tech
  • Advertise
  • Advertise
© 2025 OnlyTrustedInfo.com . All Rights Reserved.
Tech

AI models still struggle to debug software, Microsoft study shows

Last updated: April 10, 2025 3:10 pm
OnlyTrustedInfo.com
Share
4 Min Read
AI models still struggle to debug software, Microsoft study shows
SHARE

AI models from OpenAI, Anthropic, and other top AI labs are increasingly being used to assist with programming tasks. Google CEO Sundar Pichai said in October that 25% of new code at the company is generated by AI, and Meta CEO Mark Zuckerberg has expressed ambitions to widely deploy AI coding models within the social media giant.

Yet even some of the best models today struggle to resolve software bugs that wouldn’t trip up experienced devs.

A new study from Microsoft Research, Microsoft’s R&D division, reveals that models, including Anthropic’s Claude 3.7 Sonnet and OpenAI’s o3-mini, fail to debug many issues in a software development benchmark called SWE-bench Lite. The results are a sobering reminder that, despite bold pronouncements from companies like OpenAI, AI is still no match for human experts in domains such as coding.

The study’s co-authors tested nine different models as the backbone for a “single prompt-based agent” that had access to a number of debugging tools, including a Python debugger. They tasked this agent with solving a curated set of 300 software debugging tasks from SWE-bench Lite.

According to the co-authors, even when equipped with stronger and more recent models, their agent rarely completed more than half of the debugging tasks successfully. Claude 3.7 Sonnet had the highest average success rate (48.4%), followed by OpenAI’s o1 (30.2%), and o3-mini (22.1%).

Microsoft AI debugging benchmark
A chart from the study. The “relative increase” refers to the boost models got from being equipped with debugging tooling.Image Credits:Microsoft

Why the underwhelming performance? Some models struggled to use the debugging tools available to them and understand how different tools might help with different issues. The bigger problem, though, was data scarcity, according to the co-authors. They speculate that there’s not enough data representing “sequential decision-making processes” — that is, human debugging traces — in current models’ training data.

“We strongly believe that training or fine-tuning [models] can make them better interactive debuggers,” wrote the co-authors in their study. “However, this will require specialized data to fulfill such model training, for example, trajectory data that records agents interacting with a debugger to collect necessary information before suggesting a bug fix.”

The findings aren’t exactly shocking. Many studies have shown that code-generating AI tends to introduce security vulnerabilities and errors, owing to weaknesses in areas like the ability to understand programming logic. One recent evaluation of Devin, a popular AI coding tool, found that it could only complete three out of 20 programming tests.

But the Microsoft work is one of the more detailed looks yet at a persistent problem area for models. It likely won’t dampen investor enthusiasm for AI-powered assistive coding tools, but with any luck, it’ll make developers — and their higher-ups — think twice about letting AI run the coding show.

For what it’s worth, a growing number of tech leaders have disputed the notion that AI will automate away coding jobs. Microsoft co-founder Bill Gates has said he thinks programming as a profession is here to stay. So has Replit CEO Amjad Masad, Okta CEO Todd McKinnon, and IBM CEO Arvind Krishna.

You Might Also Like

Ever Seen a Bug Use Its Nose Like a Power Tool? Meet the Weevil

Amazon unveils a new AI voice model, Nova Sonic

Global Tiger Day Reminds Us We Need to Protect These Cats

Scientists Are Pretty Close to Replicating the First Thing That Ever Lived

Hurricane Erin stirs up strong winds, floods part of main highway as it creeps along the East Coast

Share This Article
Facebook X Copy Link Print
Share
Previous Article 2025 NBA Mock Draft: Where will Cooper Flagg, Ace Bailey land? 2025 NBA Mock Draft: Where will Cooper Flagg, Ace Bailey land?
Next Article My favorite iOS 18 upgrade for Apple’s Reminders app is tiny My favorite iOS 18 upgrade for Apple’s Reminders app is tiny

Latest News

PFL Brussels 2026: Why the Odds Are Stacked Against the Underdogs in a Night of Dominant Favorites
PFL Brussels 2026: Why the Odds Are Stacked Against the Underdogs in a Night of Dominant Favorites
Sports May 23, 2026
Ja Morant Spotted at WNBA’s Dream vs. Wings: What His Presence Means for the NBA Star and Women’s Basketball
Ja Morant Spotted at WNBA’s Dream vs. Wings: What His Presence Means for the NBA Star and Women’s Basketball
Sports May 23, 2026
WWE Clash in Italy: Rhea Ripley vs. Jade Cargill Rematch Confirmed—Why This Title Showdown Matters
WWE Clash in Italy: Rhea Ripley vs. Jade Cargill Rematch Confirmed—Why This Title Showdown Matters
Sports May 23, 2026
Gerrit Cole’s Triumphant Return: 6 Shutout Innings After 569-Day Absence, But Yankees Fall to Rays
Gerrit Cole’s Triumphant Return: 6 Shutout Innings After 569-Day Absence, But Yankees Fall to Rays
Sports May 23, 2026
//
  • About Us
  • Contact US
  • Privacy Policy
onlyTrustedInfo.comonlyTrustedInfo.com
© 2026 OnlyTrustedInfo.com . All Rights Reserved.