Software defect detection method, software defect repair method and software defect repair system
The software is detected and repaired through a large language model, which solves the problems of inaccurate software defect detection and unblocked repair links in the existing technology, and achieves efficient and accurate defect detection and repair, improving software quality and security.
Patent Information
- Application Number
- CN202510176627.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to detect software defects efficiently and accurately, and the link for defect repair after defect detection is not smooth.
The software is detected and repaired by a large language model, and suspected defect fragments are located through static analysis, and the large language model is used for secondary verification and repair code generation.
It reduces the false alarm rate of code defect detection, enhances the process integration of defect detection and repair, and improves software quality, reliability and security.
Smart Images

Figure CN120104464A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software testing, and in particular to a software defect detection method, a repair method and a repair system. Background Art
[0002] Software defects are usually introduced by R&D personnel during the software design and writing stages. These defects can be exploited by hackers to bypass the system's access control, illegally obtain access rights, leak user privacy data, steal digital assets, and cause huge economic losses to companies and individuals. Before the software is officially launched, quickly and efficiently locating potential defects and repairing them is an effective way to improve software quality, reliability and security and avoid software system risks.
[0003] Mainstream software defect detection methods include static analysis and dynamic analysis. Static analysis refers to scanning the software through lexical analysis, syntax analysis, rule matching and other methods without running the code, so as to detect potential defects. Dynamic testing refers to a software defect detection method that provides unexpected input to the running program under test and monitors the abnormal result output. At present, there are some limitations of static and dynamic analysis tools, such as high false positives in static analysis, few defect coverage types, low code coverage rate in dynamic analysis, and high resource overhead. In the past few years, machine learning-based methods have been widely explored in vulnerability detection. Due to the need for a large amount of training data and the need for clear annotations in feature engineering to approximate the complexity of the analysis system, its feasibility in practical applications is limited. In addition, after the software defect is detected, the tester cannot directly provide a repair patch. The defect still needs to be submitted to the relevant developer for patch correction, which leads to a blocked link between defect detection and defect repair. Summary of the invention
[0004] In order to overcome the deficiencies of the prior art, the present invention provides a software defect detection method, a repair method and a repair system to solve the problems existing in the prior art such as difficulty in efficiently and accurately detecting software defects.
[0005] The technical solution adopted by the present invention to solve the above problems is:
[0006] A software defect detection method performs defect detection on software based on a large language model.
[0007] As a preferred technical solution, the following steps are included:
[0008] S1, defect detection: static analysis of the software to be tested is performed to locate suspected defect fragments;
[0009] S2, defect verification: Use a large language model to perform secondary verification on suspected defect fragments to eliminate false positive code defects.
[0010] As a preferred technical solution, in step S2, the large language model retrieves defect samples from the knowledge base, inputs the suspected defect fragments and the corresponding defect samples into the large language model, and allows the large language model to make a judgment based on the prompt words. If it is judged as a false alarm by the large language model, it jumps to judge the next defect; if it is verified as a real code defect by the large language model, the defect is repaired.
[0011] As a preferred technical solution, the data of the knowledge base includes: defect information, defect code snippets, and code repair snippets.
[0012] As a preferred technical solution, in step S2, combined with the static analysis results, the large language model first searches the knowledge base for defect information and code snippets corresponding to the defect information, and retrieves multiple 缺陷信息 ,C 缺陷代码 >Yes, input the prompt samples, static analysis results and instructions into the large language model to obtain the verification result of the defect.
[0013] A software defect repair method, comprising the software defect detection method, further comprising the following steps:
[0014] S3, Defect Fix: Use a large language model to generate defect fix code that is verified as defect code.
[0015] As a preferred technical solution, in step S3, defect samples and defect repair samples corresponding to the defect samples are retrieved from the knowledge base, and the large language model is used to generate defect repair codes corresponding to the defect codes in a few-shot prompt word manner.
[0016] As a preferred technical solution, the large language model retrieves the defect information, the code snippet corresponding to the defect information, and the code snippet after the defect is fixed in the knowledge base, and retrieves multiple 缺陷信息 ,C 缺陷代码 ,P 缺陷补丁 >Triple, input the prompt sample, static analysis result and instruction into the large language model to obtain the repair result of the defect.
[0017] As a preferred technical solution, step S3 includes the following steps:
[0018] S31, extract a portion of sample data from the project X in the knowledge base to construct the training sample set χ of the large language model 0 T , the vulnerability exploit samples collected directly from open source projects 0 T Mark it as label Y, and take the sample set in X except for training as χ 0 E , train to minimize the loss function L, and find χ 0 E The candidate result of the label closest to the sample pair in each input information;
[0019] S32, using the trained large language model to evaluate the candidate results to obtain the confidence probability of the candidate results, and then generating deep learning context learning prompt words based on the output candidate results and their confidence probabilities.
[0020] A software defect repair system is used to implement the software defect repair method, including a large language model and a knowledge base that communicate with each other.
[0021] Compared with the prior art, the present invention has the following beneficial effects:
[0022] (1) To address the problem of high false positives in static analysis, the present invention introduces a large language model defect verification mechanism, performs secondary judgment on the detection results through intelligent means, and reduces false positives in code defect detection;
[0023] (2) The present invention proposes a code defect detection and repair method based on a large language model, which enhances the process of code defect detection and repair. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 for Figure 1 Flowchart of the software defect detection and repair method driven by a large language model;
[0025] Figure 2 Generate a schematic diagram for the prompt word. DETAILED DESCRIPTION
[0026] The present invention will be further described in detail below in conjunction with embodiments and drawings, but the embodiments of the present invention are not limited thereto.
[0027] Example 1
[0028] like Figure 1 to Figure 2 As shown, the present invention focuses on building a workflow driven by a large language model. By exploring the possibility of integrating large language models and traditional code security detection technologies in code security detection task scenarios, a large language model-driven software defect detection and repair framework is formed to accelerate the process of code defect detection and repair, and improve the efficiency of code defect detection to help developers identify and repair code defects.
[0029] include:
[0030] 1. Study the application methods of large language models in the field of software defect detection and repair;
[0031] 2. Reduce the high false alarm rate of traditional static defect detection;
[0032] 3. Through the introduction of a large language model, the link between defect detection and defect repair is connected.
[0033] The present invention provides a code defect detection and repair system based on a large language model. The system is composed of an intelligent core module and a knowledge base module. The system takes an open source large language model as the intelligent core, and combines the prompt word project to enhance its ability in the field of code defect detection, and adapts to the intelligent detection and repair needs in the system. The knowledge base is mainly composed of code defect fragments and code defect repair fragment data, which are used to supplement the context information of the prompt words and support the accuracy of reasoning and generation of the large language model. The system divides the software defect detection and repair process into three stages. First, a static analysis is performed on the software to be tested to locate suspected defect fragments. Secondly, the suspected defect fragments are secondary verified in combination with the large language model to eliminate false positive code defects. Finally, the large language model is used to generate code fragment patches verified as defects.
[0034] 1. Defect detection and repair framework based on large language model:
[0035] The typical features of the defect detection and repair framework based on the large language model include: automation and intelligent integration. The large language model is used as a core intelligent tool, which is deeply integrated with the software automated security detection process to ensure the efficiency of defect detection and repair. The model continuously optimizes its performance in large-scale data analysis and defect identification through adaptive learning. Secondly, the framework design is oriented to be available to everyone, and it realizes real-time recommendation and iteration, collaboration and verification, and simplification of complex tasks through intuitive human-computer interaction mechanisms. Finally, intelligent gain is introduced by combining the knowledge base with the large language model. Through the two-way transfer of knowledge, the framework is designed to address problems such as high false alarm rate of software defect detection, difficult results to verify, and independent detection and repair.
[0036] The main process of the system is as follows Figure 1As shown in the figure, it is divided into three stages: defect detection stage, defect verification stage and defect repair stage. In the defect detection stage, due to the problems of word input limitation, high detection time consumption and commercial model charging in the large language model, the traditional static analysis tool is still used to analyze the software under test to obtain the defect detection result of this stage. This result has the inherent defects of traditional static analysis technology, that is, too high false positives and too low precision. Then enter the defect verification stage. In the traditional static analysis process, verification mainly relies on manual verification by security testers, which has certain requirements on the professional level and resource investment of testers. In the defect detection and repair system based on the large language model, defect samples are retrieved from the knowledge base, and all suspected defect fragments are polled in a few-shot prompt word manner, and are input into the large language model in turn, so that the large language model can make a judgment based on the prompt word. If it is judged as a false positive by the large language model, it will jump to judge the next defect. If it is verified as a real code defect by the large language model, it will enter the third stage, that is, the defect repair stage. In the traditional code defect repair process, security personnel mainly confirm code defects and submit them to developers, who then repair them manually. The entire detection and repair chain is disconnected. In the defect detection and repair system based on the large language model, defect samples and corresponding defect repair samples are retrieved from the knowledge base, and the large language model generates the corresponding defect repair code in the form of few-shot prompt words.
[0037] 2. Defect detection and repair knowledge base:
[0038] The defect and detection and repair knowledge base plays a vital role in promoting contextual understanding and long-term information retention. It provides background knowledge for specific code defects, mainly including certain types of code defect information, defect code snippets, and defect repair methods. Through the contextual learning prompt words, the large language model can obtain short-term knowledge, enhance the large language model's understanding of interactive contexts, and achieve coherent communication.
[0039] The data structure of the knowledge base comes from defect information, defect code snippets, and code repair snippets. Among them, defect information includes: Common Weakness Enumeration (CWE), Common Vulnerability Exposure (CVE) and other data related to code vulnerability information; defect code snippets come from two aspects, namely artificially constructed defect code snippets and real defect code snippets. Artificially constructed defect code snippets mainly come from SARD. The code snippets cover 118 types of code defects, with a total of 64,099 code defect cases for C / C++, a total of 28,881 test cases for Java, and a total of 28,942 test cases for C#. The artificially constructed defect code snippets are not as logically complex and abstract as real codes, but they cover a wide range of defect types, have complete annotations, and have corresponding repair methods for the defects, which can provide rich information for the prediction of large language models. The real code snippets come from the code hosting platform. First, parse the reference hyperlinks in the defect information CVE. Some hyperlink URLs will be marked with the patch tag, indicating that the link is related to the code patch. Then find the hyperlinks from github (a code hosting platform) and other code hosting platforms that can obtain source code in the hyperlinks with the patch tag. Then, you can download the code of the version corresponding to the patch and the version before the patch. These two versions of the code correspond to the code snippets after and before the software defect is fixed. After determining the source of knowledge, carry out knowledge collection and data format unification. Through data cleaning and standardization, process the data into knowledge and store it in the vector database. Design a task-sensitive and scenario-related knowledge retrieval mechanism to query the knowledge required by the prompt word through keyword matching or similarity retrieval.
[0040] 3. Large language model defect detection and repair prompt words:
[0041] Prompt word engineering is crucial in the field of large language models. The mainstream prompt word methods include thought chain prompt words and context learning prompt words. This patent uses context learning to design vulnerability detection and repair prompt words. By splicing and defect detection and defect repair prompt words, the large language model can learn online to understand the relationship between defect types and specific code snippets, and then generate results based on this. In the defect verification stage, combined with the static analysis results, the system first retrieves the defect information and corresponding code snippets in the knowledge base, and retrieves multiple 缺陷信息 ,C 缺陷代码 >Yes, the prompt sample, static analysis results and instructions are input into the large language model to obtain the verification result of the defect. In the defect repair stage, the system retrieves the defect information, the corresponding code snippet and the code snippet after the defect is repaired in the knowledge base, and retrieves multiple <I缺陷信息 ,C 缺陷代码 ,P 缺陷补丁 >Triple, input the prompt sample, static analysis result and instruction into the large language model to obtain the repair result of the defect. And the preset template S as input. First, extract a part of sample data from the project X in the knowledge base to construct the training sample set Vulnerability exploit samples collected directly from open source projects Labeled as label Y. The sample set used for training is removed from X as represents the training set, Y represents the label information corresponding to the training set, and the loss function L is defined in the training to measure the gap between the model's predicted output and the true label Y. The training goal is to minimize the loss function L by adjusting the parameters of the model M, that is, to make the model's prediction as close to the true label as possible. For each input information in the set, the locality-sensitive hashing algorithm (LSH) is used to find the candidate result that is closest to the truth. Then the trained large language model is used to evaluate these candidate results to obtain the confidence probability of the results. The prompt word automatic generation model uses the output candidate results and probabilities of the large language model to enhance and generate deep learning context learning prompt words.
[0042] Example 2
[0043] like Figure 1 to Figure 2 As shown, based on Example 1, this example provides a more detailed implementation method.
[0044] The present invention provides a large language model driven software defect detection and repair method and technology, comprising the following steps:
[0045] 1. Build a defect knowledge base. In this embodiment, the defect knowledge base contains three types of knowledge, namely code defect knowledge, code defect snippets, and code defect repair snippets. After determining the source of knowledge, knowledge collection and data format unification are carried out. Through data cleaning and standardization, data is processed into knowledge and stored in a vector database. In addition, a task-sensitive and scenario-related knowledge retrieval mechanism is designed;
[0046] 2. Select a large language model, weigh the options of open source and closed source large language models, and make the best choice in terms of model performance and model resource consumption;
[0047] 3. Combined with the system framework, build a software defect detection and repair system driven by a large language model;
[0048] 4. Upload the software under test to the system and conduct the first phase. Use static analysis tools to conduct preliminary analysis on the software under test to obtain defect detection results with a high false alarm rate. Then enter the defect verification phase. Use a large language model to replace the traditional static analysis process that relies on manual verification by security testers. Retrieve defect samples from the knowledge base, and use a few-shot prompt word method to poll all suspected defect fragments and input them into the large language model in sequence. Let the large language model make a judgment based on the prompt word. If it is judged as a false alarm by the large language model, jump to judge the next defect. If it is verified as a real code defect by the large language model, enter the third phase, retrieve defect samples and corresponding defect repair samples from the knowledge base, and use a few-shot prompt word method to let the large language model generate the corresponding defect repair code. Finally, complete the code defect detection and repair tasks of the software under test.
[0049] As described above, the present invention can be preferably implemented.
[0050] All features disclosed in all embodiments in this specification, or steps in all methods or processes implicitly disclosed, except for mutually exclusive features and / or steps, can be combined and / or expanded or replaced in any manner.
[0051] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. According to the technical essence of the present invention, within the spirit and principles of the present invention, any simple modification, equivalent replacement and improvement made to the above embodiment still falls within the protection scope of the technical solution of the present invention.
Claims
1. A software defect detection method, characterized in that: Software defect detection based on large language models.
2. A software defect detection method according to claim 1, characterized in that: The following steps are involved: S1, defect detection: static analysis of the software to be tested is performed to locate suspected defect fragments; S2, defect verification: Use a large language model to perform secondary verification on suspected defect fragments to eliminate false positive code defects.
3. A software defect detection method according to claim 2, characterized in that: In step S2, the large language model retrieves defect samples from the knowledge base, inputs the suspected defect fragments and the corresponding defect samples into the large language model, and allows the large language model to make a judgment based on the prompt words. If it is judged as a false alarm by the large language model, it jumps to judge the next defect; if it is verified as a real code defect by the large language model, the defect is repaired.
4. A software defect detection method according to claim 3, characterized in that: In step S2, the data of the knowledge base includes: defect information, defect code snippets, and code repair snippets.
5. A software defect detection method according to any one of claims 2 to 4, characterized in that: In step S2, combined with the static analysis results, the large language model first searches the knowledge base for defect information and code snippets corresponding to the defect information, and retrieves multiple 缺陷信息 ,C 缺陷代码 >Yes, input the prompt samples, static analysis results and instructions into the large language model to obtain the verification result of the defect. 6. A software defect repair method, characterized in that: A software defect detection method according to any one of claims 2 to 5, further comprising the following steps: S3, Defect Fix: Use a large language model to generate defect fix code that is verified as defect code.
7. A software defect repair method according to claim 6, characterized in that: In step S3, defect samples and defect repair samples corresponding to the defect samples are retrieved from the knowledge base, and the large language model is used to generate defect repair codes corresponding to the defect codes in a few-shot prompt word manner.
8. A software defect repair method according to claim 7, characterized in that: In step S3, the large language model retrieves the defect information, the code snippet corresponding to the defect information, and the code snippet after the defect is fixed in the knowledge base, and retrieves multiple 缺陷信息 ,C 缺陷代码 ,P 缺陷补丁 >Triple, input the prompt sample, static analysis result and instruction into the large language model to obtain the repair result of the defect. 9. A software defect repair method according to any one of claims 6 to 8, characterized in that: Step S3 includes the following steps: S31, extract a part of sample data from project X in the knowledge base to build a training sample set for the large language model Vulnerability exploit samples collected directly from open source projects Marked as label Y, remove the sample set used for training from X as The training is carried out with the goal of minimizing the loss function L. The candidate result that is closest to the real one for each input information; S32, using the trained large language model to evaluate the candidate results to obtain the confidence probability of the candidate results, and then generating deep learning context learning prompt words based on the output candidate results and their confidence probabilities.
10. A software defect repair system, characterized in that: A software defect repair method for implementing any one of claims 6 to 9, comprising a large language model and a knowledge base that communicate with each other.
Citation Information
Cited By
Code defect reasoning instruction template generation method and system for large language model
CN120631737A
Software defect positioning method and system driven by thinking chain iteration verification
CN121412101A
A thought chain iteration verification driven software defect positioning method and system
CN121412101B