一种基于代码语言模型的程序缺陷自动修复方法及系统

By constructing an automatic program defect repair method based on a code language model, and using open-source project corpora and defect-repair code datasets for multi-round iterative fine-tuning training, the method solves the problems of long repair time, high cost and poor applicability in existing technologies, and achieves efficient and universally applicable program defect repair.

CN116755753BActive Publication Date: 2026-07-17HARBIN INST OF TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HARBIN INST OF TECH
Filing Date
2023-06-19
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing methods for fixing program defects are time-consuming, costly, and lack universality. Traditional manual repair relies on the professional skills of developers, heuristic rule-based methods have limited applicability, and patches generated by neural networks are of poor quality and have a low success rate.

Method used

We construct an automatic program defect repair method based on a code language model. We build the model using an open-source project corpus, construct a dataset using program defects and correct code samples, establish a defect repair difficulty discrimination strategy and a sample scheduling function, conduct multiple rounds of iterative fine-tuning training, and generate and verify candidate patches.

Benefits of technology

It improves the success rate of program defect repair and the quality of generated patches, enables unified repair of multiple types of defects, reduces the time cost of manual repair, and provides auxiliary tools for developers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116755753B_ABST
    Figure CN116755753B_ABST
Patent Text Reader

Abstract

本发明涉及一种基于代码语言模型的程序缺陷自动修复方法,包括如下步骤:步骤一、利用开源程序项目源代码语料库,构建代码语言模型;步骤二、利用程序缺陷代码样本和正确代码样本,构建缺陷‑修复代码数据集;步骤三、构建基于缺陷修复难度判别策略的缺陷‑修复代码样本分类器,对数据集样本进行评估并划分缺陷修复难度等级。本发明基于课程学习机制的微调训练框架对数据集进行缺陷修复难度划分,使用处理后的数据集对代码语言模型进行多轮迭代微调训练,能够有效提升模型生成补丁的质量,提高程序缺陷修复成功率。
Need to check novelty before this filing date? Find Prior Art