Program defect automatic repairing method combining source code semantics and exception feedback

By combining program source code semantics and exception feedback to build proprietary training data, and designing composite loss function to guide the model to learn program semantics and context information, the problem of insufficient automation of existing program defect automatic repair methods and large differences in training data semantics is solved, and more efficient program defect repair is achieved.

CN120123128APending Publication Date: 2025-06-10BEIJING INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510170194.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing automatic program defect repair methods are insufficiently automated, and the method based on repair mode is high in labor costs and has limited application scope. However, the method based on neural machine translation is difficult to effectively learn repair strategies due to the large difference between the training data and the program to be repaired, and lacks unique identifiers and semantic relationships when generating repair codes.

Method used

Combining the semantic features and exception feedback of the source code of the program to be repaired, proprietary training data is constructed, and a composite loss function is designed to guide the model to learn the semantic and context information of the program to generate correct repair code. The specific steps include generating defect variant programs with different semantics, extracting defect context and exception feedback, vectorizing and splicing as training data, and optimizing the composite loss function to train the defect repair model.

Benefits of technology

The number of defects that are correctly repaired is improved, and the problems of high labor costs and insufficient automation based on the repair mode method are alleviated. At the same time, the problem of difficult model learning program characteristics caused by large semantic differences in training data based on neural machine translation methods is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123128A_ABST
    Figure CN120123128A_ABST
Patent Text Reader

Abstract

The invention relates to an automatic program defect repairing method combining source code semantics and abnormal feedback, and belongs to the technical field of computer programs. The method comprises the following steps: firstly, selecting a latest released version of a to-be-repaired program, and randomly adding disturbance to the version program to generate a plurality of defect variant programs; secondly, extracting elements such as a defect row and a defect neighbor row of the defect variant program to construct a defect context; compiling errors or function errors after execution of the defect variant program are obtained to serve as abnormal feedback; finally, the defect context and the abnormal feedback are vectorized and spliced into training data, a loss function is designed to minimize the semantic difference between the repair code and the latest issued program, and a defect repair model is trained and generated. Aiming at the problem that a model is difficult to learn normal program characteristics due to the fact that open source training data and to-be-repaired program semantic differences are large in an existing method, the number of correctly repaired defects is increased by constructing proprietary training data and designing a target function to guide the model to learn semantic and contextual information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an automatic program defect repair method combining source code semantics and exception feedback, belonging to the technical field of computer programs. Background Art

[0002] Program defects refer to problems such as syntax errors, logical errors, and imperfect designs existing in program source code, which cause exceptions during project operation. With the wide application of software systems in fields such as industrial control, traffic management, medical assistance, and financial transactions, potential program defects in the system may cause serious damage to the economy and society. However, manual repair of program defects is both time-consuming and prone to introducing new defects, making subsequent repairs more difficult. Therefore, researching efficient and reliable automatic program defect repair methods can improve software system stability, reduce operation and maintenance costs, and is of great significance for ensuring the security of key fields and enhancing society's trust in software technology.

[0003] In the research of automatic program defect repair methods, the Generate-and-Validate method does not depend on a specific program structure or programming language, is applicable to a variety of application scenarios and defect types, and thus has received extensive attention. The Generate-and-Validate method generates a large number of candidate patches for program defects and uses a test suite to screen out the candidate patches that correctly repair the defects. According to the generation method of candidate patches, the Generate-and-Validate method is classified into a repair pattern-based method and a neural machine translation-based method.

[0004] 1. Repair Pattern-Based Method

[0005] The repair pattern-based method analyzes the repair methods of known defects, extracts reusable code repair patterns, and applies them to similar new defects. This type of method uses manually predefined repair templates or patterns automatically extracted from historical repair records to repair specific types of defects. Due to the inclusion of a large amount of domain knowledge, this type of method can generate high-quality patches. However, the automation degree of this type of method is insufficient. The method of predefined templates has a high labor cost and a limited scope of application, and the templates generated by the method of automatically extracting repair patterns also have limitations such as insufficient generalization.

[0006] 2. Neural Machine Translation-Based Method

[0007] The method based on neural machine translation defines defect repair as the process of translating the program to be repaired into the repaired program, and applies natural language processing technology to capture the complex mapping relationship between the program to be repaired and the repaired program to achieve automatic repair, which is applicable to various programming languages and different defect types. However, the training data of existing methods usually comes from historical repairs in open-source projects, with a large semantic difference from the program to be repaired, resulting in the generated repair code lacking the semantic relationship between the specific identifiers, idiomatic expressions, and domain identifiers of the program to be repaired, making it difficult to correctly repair defects; moreover, the training data is only a single defect-repair code pair and cannot guide the model to learn defect types.

[0008] In summary, the existing methods for automatic program defect repair mainly have the following problems: (1) The method based on repair patterns has insufficient automation: The manual cost of defining templates is relatively high and the applicable scope is limited, and developers still need to spend a lot of time and energy; (2) The open-source training data used by the method based on neural machine translation has a large semantic difference from the program to be repaired: Only considering the static source code as the input for model training, lacking the utilization of the semantic and exception feedback of the defective program source code, making it difficult for the model to effectively learn repair strategies and generate repair code. Therefore, the present invention proposes an automatic program defect repair method that combines source code semantics and exception feedback. Summary of the Invention

[0009] The purpose of the present invention is to construct proprietary training data by combining the semantic features of the source code of the program to be repaired and the defect information of the exception feedback, and design an objective function to guide the model to learn the semantics and context information of the program to be repaired, so as to generate correct repair code, in order to address the problem that the existing methods use open-source training data with a large semantic difference from the program to be repaired, resulting in the model being difficult to learn the characteristics of normal programs, and improve the number of defects repaired correctly.

[0010] The design principle of the present invention is as follows: First, select the latest release version of the program to be repaired, randomly add character, word, statement, and code block-level perturbations to the program of this version to generate defective variant programs with different semantics; secondly, extract the defective lines, defective adjacent lines, variables in the scope of the defective function, defective function names, and class names of the defective variant programs to construct the defect context; then compile the defective variant programs, and use the test suite to execute the programs that pass the compilation to obtain exception feedback including compilation error information and functional error information; vectorize and splice the defect context and the exception feedback as training data, design a contrastive loss in combination with the difference between the latest release program and the defective variant program, and weight the contrastive loss and the cross-entropy loss to obtain a composite loss function, and train the defect repair model under the optimization of the composite loss.

[0011] The technical solution of the present invention is realized through the following steps:

[0012] Step 1, select the latest release version of the program to be repaired. The program corresponding to this version is P, and perturb P to generate a set of defective variant programs P'.

[0013] Step 1.1, select the latest release version of the program to be repaired. This version should include a test suite and the program P corresponding to this version can correctly execute the test suite.

[0014] Step 1.2, perform character replacement and word replacement on the identifiers of P, and add the perturbed program to the set P'.

[0015] Step 1.3, add or delete words from the statements of P or replace statements, and add the perturbed program to the set P'.

[0016] Step 1.4, change the order of statements, delete statements or code blocks in P, and add the perturbed program to the set P'.

[0017] Step 1.5, add existing statements or code blocks to P, and add the perturbed program to the set P'.

[0018] Step 1.6, filter the set P' so that the programs in the set P' are not repeated and are different from P.

[0019] Step 2, for each program in the set of defective variant programs P', extract the defective line, defective neighboring lines, defective function scope variables, defective function names, and class names to construct the defective context, and convert the defective context into a token sequence. Add each defective context token sequence to the set C.

[0020] Step 3, compile all the programs in the set of defective variant programs P', use the test suite of the latest release program P to execute the programs that pass the compilation, obtain the compilation error information and functional error information of the programs with compilation errors and test suite execution errors, convert them into token sequences, and add them to the exception feedback set D. The items in the defective context set C and the exception feedback set D correspond one by one.

[0021] Step 4, vectorize and splice the defective context and the exception feedback as training data, design a composite loss function that combines the contrastive loss and the cross-entropy loss, and train to generate a defective repair model.

[0022] Step 4.1, splice the token sequences corresponding one by one in the defective context set C and the exception feedback set D as training data.

[0023] Step 4.2, regard the latest release program as the positive sample and the defective variant program as the negative sample, design the contrastive loss, weight the contrastive loss and the cross-entropy loss to obtain the composite loss function, and train to generate a defective repair model with the training data under the optimization of the composite loss function.

[0024] Beneficial effects

[0025] Compared with the method based on the repair mode, the guiding model of the present invention learns to use valid identifiers, reuse the existing code in the same program, generate new code lines according to the context and delete code lines, alleviating the problems of high labor cost for defining repair templates and limited scope of application.

[0026] Compared with other methods based on neural machine translation, the present invention generates defective variant programs for the latest release versions of defective programs, constructs proprietary training data and designs a target loss function to guide the model to learn the semantic features and abnormal feedback of the source code of the program to be repaired, alleviating the problem that the open-source training data used in the existing methods has a large semantic difference from the program to be repaired, resulting in the difficulty for the model to learn the features of the program to be repaired, and improving the number of defects correctly repaired. Description of the drawings

[0027] Figure 1 It is a schematic diagram of the automatic program defect repair method combining source code semantics and abnormal feedback of the present invention.

[0028] Figure 2 It is a diagram of the definition and example of the perturbation rule of the present invention. Detailed implementation manners

[0029] In order to better illustrate the purpose and advantages of the present invention, the implementation manners of the method of the present invention will be further described in detail below with reference to examples.

[0030] The experimental data comes from 2 publicly available datasets mainly used in the field: Defects4J v1.2 dataset (2014), Defects4J v2.0 dataset (2020). The Defects4J dataset is a collection of reproducible defects and their supporting architectures, in which real defects in multiple open-source projects are collected, and each project contains a corresponding test suite. The Defects4J v1.2 dataset contains 391 available defects from 6 open-source projects. The Defects4J v2.0 dataset adds 444 defects from 12 open-source projects on the basis of the v1.2 version, for a total of 835 available defects in 17 open-source projects. For each defect, the Defects4J dataset collects the code of both the defective version and the repaired version. The defective version is the code version found and marked as defective by the project developer, and the repaired version is the code version submitted by the developer after repairing the defect. The detailed data of the Defects4J v1.2 dataset and the Defects4J v2.0 dataset are shown in Table 1 and Table 2, where the data volume of the Defects4J v2.0 dataset is the newly added data volume relative to the v1.2 version.

[0031] Table 1. Defects4J v1.2 Dataset

[0032]

[0033] Table 2. Defects4J v2.0 Dataset

[0034]

[0035] Among them, for each project, the fixed version of its first defect is selected to generate the perturbed training data. That is, a total of 17 fixed-version programs from 17 projects are selected to generate the training data, and the remaining 818 defects are used for model prediction. (The Defects4J v2.0 dataset provides the fixed version for each defect. To avoid selecting a version that has fixed a large number of defects in the dataset, the fixed version of the first defect of each project is selected. During the experiment, the repair effect of the defect repair model on the first defect of each project will not be examined, but the repair effect on the remaining 818 defects will be examined.) The experiment is set as perfect fault localization, that is, the program file and line number of code where each defect belongs are known.

[0036] The experiment uses the number of correctly fixed defects, Correct, as the evaluation metric. A correct fix means that the repaired code can successfully pass compilation, successfully execute the test suite, and is semantically equivalent to the fix code provided by the developer. The more defects are correctly fixed, the stronger the repair performance of the method.

[0037] The hardware and software environments used in the experiment are shown in Table 3 and Table 4 respectively.

[0038] Table 3. Experiment Hardware Environment

[0039]

[0040] Table 4. Experiment Software Environment

[0041]

[0042] The specific process of this experiment is as follows:

[0043] Step 1, select the fixed version after the first defect occurs in each project. The program corresponding to this version is P, and perturb P to generate a set of defect variant programs P'.

[0044] Step 1.1, select the fixed version after the first defect occurs in each project. This version contains a test suite and the program P corresponding to this version can correctly execute the test suite.

[0045] Step 1.2, define the perturbation rules as Figure 2As shown (the codes given in the figure are only examples). Rule 1 modifies the correct declaration type with an incorrect declaration type; Rule 2 modifies the correct operator with an incorrect operator; Rule 3 modifies the correct literal with an incorrect literal; Rule 4 modifies the correct constructor with an incorrect or overloaded constructor; Rule 5 modifies the correct parameters of a function to incorrect parameters; Rule 6 swaps the positions of two parameters in the same function; Rule 7 modifies the correct function call with an overloaded function call; Rule 8 replaces the correct call with another call that appears in the program file; Rule 9 superimposes the perturbation methods of Rules 1 to 8 to generate compound statements; Rule 10 deletes a clause in a boolean expression; Rule 11 copies an existing binary expression in the function scope to expand the correct boolean expression; Rule 12 copies a similar statement from the function scope to replace the correct statement; Rule 13 modifies the order of the correct statements; Rule 14 deletes the target statement; Rule 15 deletes the if statement and only keeps the then statement; Rule 16 deletes the entire code block; Rule 17 randomly inserts lines of code from the surrounding area before or after the target statement; Rule 18 includes the target statement within a conditional expression; Rule 19 copies the entire code block before and after the target statement.

[0046] Step 1.3, add perturbations as shown in Figure 2 Rules 1 to 9 to P, guiding the model to learn to use valid identifiers during training in Step 4. Add the perturbed program to the set P'.

[0047] Step 1.4, add perturbations as shown in Figure 2 Rules 10 to 12 to P, guiding the model to learn to reuse the existing code in the same program during training in Step 4. Add the perturbed program to the set P'.

[0048] Step 1.5, add perturbations as shown in Figure 2 Rules 13 to 16 to P, guiding the model to learn to generate new lines of code according to the context during training in Step 4. Add the perturbed program to the set P'.

[0049] Step 1.6, add perturbations as shown in Figure 2 Rules 17 to 19 to P, guiding the model to learn to delete lines of code during training in Step 4. Add the perturbed program to the set P'.

[0050] Step 1.7, filter the set P' so that the programs in the set P' are non-repetitive and different from P.

[0051] Step 2, for each program in the defective variant program set P′, extract the defective line, defective neighboring lines, defective function scope variables, defective function name, and class name as the defective context. Use the Byte Pair Encoder (BPE) method to convert the defective context into a token sequence and generate a vocabulary. Add each defective context token sequence to the set C.

[0052] Step 3, compile all programs in the defective variant program set P′, execute the compiled programs using the test suite of the latest release program P, obtain the compilation error information and functional error information of the programs with compilation errors and error executions in the test suite. Also use the BPE method to convert the error information into a token sequence and expand the vocabulary. Add the token sequence of the error information to the exception feedback set D. Each item in the defective context set C corresponds one-to-one with the items in the exception feedback set D.

[0053] Step 4, concatenate the defective context and the exception feedback as training data, design a composite loss function that combines the contrastive loss and the cross-entropy loss, and train to generate a defective repair model.

[0054] Step 4.1, concatenate the token sequences that correspond one-to-one in the defective context set C and the execution feedback set D to obtain the training dataset where k is the number of data in the dataset, and x i is a piece of training data.

[0055] Step 4.2, use the training dataset X to train the Transformer model. For each training data x i , the output of the model at each training is y gen = f(x i , θ), where θ is the current model parameter, and f(x i , θ) is the output of the model when the parameter is θ and the training data is x i , that is, the generated repair code. When optimizing the model, regard the repaired version program P as the positive sample y pos , and the defective variant program as the negative sample y neg , and design the contrastive loss as shown in Equation (1):

[0056]

[0057] where y pos is the repaired version program P of the positive sample, y neg is the defective variant program of the negative sample, y gen is the repair code generated during model training, τ is the temperature coefficient used to adjust the smoothness of the distribution, and sim(a, b) represents the cosine similarity between vectors a and b, as shown in Equation (2):

[0058]

[0059] where ‖a‖ 2 and ‖b‖ 2 are the L2 norms of vectors a and b respectively. The cross-entropy loss of the model is shown in Equation (3):

[0060]

[0061] where T is the length of the generated sequence, i.e., the total length of the generated repair code, y t is the true token expected to be generated at the t-th time step, is the predicted probability distribution of the model at the t-th time step, representing the probability of generating the y i token given the input x 1 , …, y t-1 . The contrastive loss and the cross-entropy loss are weighted to obtain a composite loss function, and the formula of the composite loss function is shown in Equation (4): t

[0062] L = λL ce + (1 - λ)L con Equation (4)

[0063] where λ is the weight coefficient used to control the contribution degrees of the two losses. The Transformer model is trained with the training data under the optimization of the composite loss function to obtain the defect repair model.

[0064] Given the defect location, the trained defect repair model is used to predict the repair code of the program to be repaired. During the experiment, the defect context of the program to be repaired is also extracted and the abnormal feedback is obtained. The two are converted into token sequences by the BPE method and concatenated. The sequence is input into the trained model, and beam search is used to output 50 best repair code lines. If the repaired code can successfully pass compilation, successfully execute the test suite, and is semantically equivalent to the repair code provided by the developer, it is considered that a correct repair has been completed.

[0065] Table 5 shows the comparison between the present invention and the existing automatic program defect repair methods, where TBar is the method based on the repair pattern, and SequenceR, CoCoNuT, CURE, and RewardRepair are the methods based on neural machine translation.

[0066] Table 5 Comparison experimental results between the present invention and the existing methods

[0067]

[0068] ​Analysis of test results: The experiment was based on an automatic program defect repair method that combines source code semantics and exception feedback. The repair predictions for 818 defects in the Defect4J v1.2 dataset and the Defect4J v2.0 dataset were carried out. The results show that the present invention can construct proprietary training data by combining the semantic features of program source code and exception feedback, and design an objective function to train a defect repair model to repair more defects.

[0069] The above specific description further details the purpose, technical solution and beneficial effects of the invention. It should be understood that the above is only a specific embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for automatically repairing program defects by combining source code semantics and exception feedback, characterized in that The method comprises the following steps: Step 1: Select the latest released version of the program to be repaired and its test suite. The corresponding program of this version is P. Define the perturbation rules, add perturbations to P to generate a defect variant program set P′; Step 2: For each program in the defect variant program set P′, extract the defect line, defect neighbor line, defect function scope variable, defect function name and class name to construct the defect context, and convert the defect context into a token sequence, and add each defect context token sequence to the set C; Step 3: compile all programs in the defect variant program set P′, use the test suite of the latest released program P to execute the compiled programs, obtain the compilation error information and functional error information of the programs with compilation errors and execution test suite errors, convert them into token sequences, add them to the exception feedback set D, and the defect context set C corresponds to each item in the exception feedback set D one by one; In step 4, the defect context and abnormal feedback are vectorized and concatenated as training data. The latest released program is regarded as a positive sample and the defect variant program is regarded as a negative sample. The contrast loss is designed, and the contrast loss and the cross entropy loss are weighted to obtain a composite loss function. The defect repair model is trained and generated under the optimization of the composite loss.

2. The method for automatically repairing program defects by combining source code semantics and exception feedback according to claim 1, characterized in that: In step 3, the exception feedback is obtained as part of the training data to guide the model to learn defect type information during training; in step 4, the defect context and the exception feedback are spliced ​​to construct proprietary training data for the program to be repaired.

3. The method for automatically repairing program defects by combining source code semantics and exception feedback according to claim 1, characterized in that: Step 4: In the model training phase, the latest released program is regarded as a positive sample, and the defective variant program is regarded as a negative sample. The contrast loss is designed as shown in formula (1): where y pos is the latest release program of the positive sample, y neg is the negative sample defect variant program, y gen is the repair code generated during model training, τ is the temperature coefficient, which is used to adjust the smoothness of the distribution, and sim(a,b) represents the cosine similarity between vectors a and b, as shown in formula (2): Where ‖a‖2 and ‖b‖2 are the L2 norms of vectors a and b respectively.

Citation Information

Cited By

  • Software defect positioning and repairing recommendation method based on code semantic vector

    CN122364101A