基于代码变更预训练模型的即时缺陷预测方法及存储介质

By using a pre-trained model based on code changes, and employing five pre-training tasks to establish semantic connections between code changes and submission comments, this approach solves the problem of inconsistency between pre-training tasks and fine-tuning targets in existing technologies, and achieves efficient real-time defect prediction and code change understanding.

CN117785207BActive Publication Date: 2026-07-17NAT UNIV OF DEFENSE TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NAT UNIV OF DEFENSE TECH
Filing Date
2023-11-14
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing pre-trained models cannot effectively utilize domain knowledge of code changes in real-time defect prediction, resulting in an inconsistency between the pre-training task and the fine-tuning objective, making it difficult to achieve efficient real-time defect prediction.

Method used

We employ a code change-based pre-trained model. Through the steps of constructing training data, preprocessing, model pre-training, and real-time defect prediction, we utilize five pre-training tasks to establish semantic connections between code changes and submission comments. These tasks include mask language modeling for code changes, mask language modeling for submission comments, NL→PL generation, PL→NL generation, and structure-aware pre-training tasks, thereby enhancing the model's ability to understand code changes.

Benefits of technology

It achieves real-time defect detection with simple principles, easy implementation, and wide applicability. It can effectively utilize domain knowledge of code changes in a large amount of unlabeled data, thereby improving the efficiency and accuracy of code change-related tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117785207B_ABST
    Figure CN117785207B_ABST
Patent Text Reader

Abstract

本发明公开了一种基于代码变更预训练模型的实时缺陷预测方法及存储介质,该方法包括:步骤S1:构建训练数据并且预处理;从网络公开资源中提取模型训练所需数据,并对数据进一步预处理;步骤S2:模型预训练;将步骤S1获得的训练数据加以调整后,对模型进行预训练,获取神经网络模型;步骤S3:实时缺陷预测;对于一个提交变更代码的项目或程序,使用步骤S2中训练好的神经网络模型,来检查这些变更,识别是否有缺陷的代码变更;所述代码变更相关的理解任务是一种二元分类:预测给定的代码变更是否存在缺陷。该存储介质存储了用来执行上述方法的计算机程序。本发明具有原理简单、易实现、适用性广、能够保证实时性缺陷检测等优点。
Need to check novelty before this filing date? Find Prior Art