Code repairing system and method based on preference learning

By utilizing a preference-based code repair system, which leverages efficient fine-tuning of large language models and preference-based training, the system addresses the shortcomings in code understanding and generation capabilities in automated code repair technologies. This enables efficient and accurate code repair that meets developers' needs, thereby improving code consistency and development efficiency.

CN120872791APending Publication Date: 2025-10-31COMPUTER INNOVATION TECH RES INST OF ZHEJIANG UNIV +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511015455.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-05-15
Filing Date
2025-07-23
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing automated code repair technologies lack good code semantic understanding and code generation capabilities, making it difficult to effectively repair complex and diverse code errors. Furthermore, they do not consider the consistency of code before and after repair, often resulting in large-scale rewriting of the original code, which does not meet the needs and coding style of developers.

Method used

A code repair system based on preference learning is constructed. By constructing a code dataset, a large language model is efficiently fine-tuned and trained using preference learning to generate repair code that meets the needs of developers. This includes efficient fine-tuning of the code dataset D1, construction of the preference learning dataset D2, and training of the final code repairer M2. The Lora and DPO-Positive algorithms are used to optimize the model parameters.

Benefits of technology

Generates fix code that matches the developer's style and involves minimal changes, improving the accuracy and consistency of code fixes, reducing unnecessary modifications, and enhancing development efficiency and code quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872791A_ABST
    Figure CN120872791A_ABST
Patent Text Reader

Abstract

The invention discloses a preference learning-based code repair system and method. A code repair enhancement module of the system trains the large language model based on the code data set so as to construct an initial code repairer; a repair preference data generation module generates candidate repair codes from the programming tasks and the error codes through an initial code repairer, and constructs a preference learning data set after preference data selection; the code preference learning module is used for training the initial code restorer through the preference learning data set so as to obtain a final code restorer; and the code repair generation module generates a repair code with a small modification range for the to-be-repaired code data through the final code repairer. According to the method and the device, the repair code which is accurately repaired and has a smaller modification range can be intelligently generated according to actual requirements in a development scene, so that code problems are quickly positioned and solved, unnecessary modification of the code is effectively reduced, the consistency and readability of the code are kept, the time cost of manual debugging is reduced, and the development efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a code repair system, which falls under the field of natural language processing technology, and specifically to a code repair system and method based on preference learning. Background Technology

[0002] With the rapid iteration and increasing complexity of software development, code repair has become a crucial task in the software development process. Traditional code repair methods rely on manual intervention, increasing developers' time costs. To address these issues, automated code repair technologies have gained increasing attention and importance. However, due to a lack of robust code semantic understanding and code generation capabilities, traditional automated code repair technologies still face many limitations when dealing with complex and diverse code errors, making them ineffective in repairing code.

[0003] In recent years, the development of deep learning technology, especially large-scale pre-trained language models (such as GPT-4 and Codex), has demonstrated enormous potential in the field of code understanding and generation. These models, through learning from massive amounts of code data, can accurately understand the syntax and structure of code, thereby enabling code repair. Despite these significant advancements in code repair, they still face numerous challenges in meeting the practical needs of code repair. Current code repair methods often fail to consider the consistency of code before and after the repair, frequently involving large-scale rewriting of the original code, which does not align with developers' actual needs and coding styles. Summary of the Invention

[0004] To address the problems existing in the background art, this invention provides a code repair system and method based on preference learning. The purpose of this invention is to design a code repair preference learning framework to train a large language model, enabling it to output repair code that meets the developer's needs, i.e., repair code with minimal changes before and after the repair.

[0005] The technical solution adopted in this invention is:

[0006] I. A code repair method based on preference learning:

[0007] 1) Construct a code dataset D1 for code repair learning. Each piece of code data in the code dataset D1 contains a programming task Q and its error code C, a unit test example T, and a correct code S.

[0008] 2) The Large Language Model (LLM) is efficiently fine-tuned and trained using the code dataset D1 to obtain the initial code repairer M1 for code repair.

[0009] 3) The programming task Q and error code C of each code data in the code dataset D1 are processed by the initial code repairer M1 to obtain several candidate repair codes for error code C. After selecting each candidate repair code through the unit test sample T of error code C, the preference learning dataset D2 is constructed.

[0010] 4) Train the initial code repairer M1 using the preference learning dataset D2 to obtain the final code repairer M2 for code optimization and repair. The final code repairer M2 has the ability to modify a small range of code.

[0011] 5) The code data to be repaired is processed by the final code repairer M2. After processing, the final code repairer M2 generates repair code with a small scope of modification and displays it on the screen to achieve code repair.

[0012] In step 1), for each piece of code data in the code dataset D1, there is a programming task Q, which corresponds to several different error codes C. Each error code C corresponds to several unit test cases T and a correct code S. Each unit test case T of each error code C is run, and one of the unit test cases T that fails to run is added to the code dataset.

[0013] In step 2), firstly, several first training texts of code dataset D1 are constructed. For each piece of code data in code dataset D1, the first training text of the code data includes explanatory text and textual representations of programming task Q, error code C, and correct code S. Then, code dataset D1 and its various first training texts are input into the Large Language Model (LLM), and the LoRa efficient fine-tuning method is used to train the weights only on the additional parameter matrix of the Large Language Model (LLM). At the same time, the log-likelihood function is used as the final loss function, i.e., the negative log-likelihood function. Then, error backpropagation is used to adjust the weights and biases of the Large Language Model (LLM). During the training process, the set of model parameters corresponding to the minimum final loss function is used as the final model parameters of the Large Language Model (LLM), and finally, the initial code repairer M1 for code repair is obtained.

[0014] In step 3), firstly, several second training texts are constructed for each programming task Q and error code C in the code dataset D1. The second training texts include explanatory text and textual representations of programming task Q and error code C. Then, each programming task Q and error code C in the code dataset D1 and its second training texts are input into the initial code repairer M1 for processing. After processing, several candidate repair codes for error code C are output.

[0015] In step 3), for each unit test case T of error code C, the unit test case T is input into each candidate repair code of error code C and run to obtain several candidate repair codes R that pass and fail. Then, the Git tool is used to perform a "git diff" operation on each candidate repair code R and its corresponding error code C to obtain the total number of unmodified and undeleted lines of code L in candidate repair code R and the total number of lines of candidate repair code R. And the total number of lines L of error codes in the candidate repair code R. c Thus, the retention rate after fixing the i-th error code C is obtained. Select a set of preference pairs from each candidate repair code R. <R w ,R l >, R w R is the candidate repair code R with the highest retention rate among all candidate repair codes R that have passed the test. l To select the candidate repair code R with the lowest retention rate among all failed candidate repair codes R, and to determine it as the best and worst repair code respectively, the programming task Q, error code C, and preference pairs are constructed as preference data {Q,C,R}. w ,R l Finally, the preference data for each error code C was selected and constructed into a preference learning dataset D2.

[0016] In step 4), several third training texts of the preference learning dataset D2 are first constructed, and the optimal repair code R is then developed for each preference data in the preference learning dataset D2. w And worst-case fix code R l Optimal fix code R w The third training text includes explanatory text, programming task Q, error code C, and optimal fix code R. w The text indicates that the worst-case fix code is R. l The third training text includes explanatory text, programming task Q, error code C, and worst-case fix code R. l The text represents this.

[0017] The preference learning dataset D2 and its various third training texts are input into the initial code repairer M1, and trained using the Direct Preference Optimization with Positive (DPO-Positive) algorithm. The loss function of the DPO-Positive algorithm is used as the final loss function, with the hyperparameter set to 0.5 in the specific implementation. Then, backpropagation is used to adjust the weights and biases of the additional matrix of the initial code repairer M1. During the training process, the set of model parameters corresponding to the minimum final loss function is used as the final model parameters of the initial code repairer M1, ultimately obtaining the final code repairer M2 for code optimization and repair.

[0018] In step 5), the code data to be repaired includes the programming task Q′ to be repaired and its error code C′. A repair text is constructed for the code data to be repaired, which includes explanatory text and textual representations of the programming task Q and the error code C. The code data to be repaired and its repair text are input into the final code repairer M2 for processing to generate a repair code with a small modification range, which is then displayed on the monitor.

[0019] II. A code repair system:

[0020] The code repair enhancement module trains the large language model LLM on the code dataset D1 used for code repair learning, and constructs the initial code repairer M1 for code repair.

[0021] The preference data generation module is used to generate several candidate repair codes from the programming task Q and error code C of each code data in the code dataset D1 through the initial code repairer M1, and then construct the preference learning dataset D2 after preference data selection.

[0022] The code preference learning module trains the initial code fixer M1 using the preference learning dataset D2 to obtain the final code fixer M2 for code optimization and repair.

[0023] The code repair generation module, through the final code repairer M2, generates minor repair code for the code data to be repaired and displays it on the monitor.

[0024] 3. An electronic device comprising: a memory and a processor coupled to each other, wherein the memory stores program data, and the processor invokes the program data to execute the method described above.

[0025] IV. A computer-readable storage medium storing program data thereon, characterized in that the program data, when executed by a processor, implements the method described above.

[0026] The code repair system and method based on preference learning of the present invention trains a large language model with data on different repair preferences, which can understand the logic of repair code and the preferences of repair needs, thereby generating repair code that conforms to the developer's style and requires less modification, helping developers to locate and fix problems in the code more quickly, and improving the overall quality and efficiency of software development.

[0027] The beneficial effects of this invention are:

[0028] The code repair system and method based on preference learning of the present invention can generate precise code repair solutions with a smaller scope of modification according to the actual needs of the development scenario. It learns the needs and preferences of the code repair scenario through preference learning, and trains a large language model with code data to enable it to have code repair capabilities. It can intelligently generate repair code that meets the needs of developers, help developers quickly locate and solve code problems, effectively reduce unnecessary code changes, maintain code consistency and readability, reduce the time cost of manual debugging, thereby improving development efficiency and reducing code maintenance costs. Attached Figure Description

[0029] Figure 1 This is a flowchart of the present invention;

[0030] Figure 2 This is a schematic diagram illustrating the data construction process for preference learning in this invention;

[0031] Figure 3 This is a flowchart of the training process of the present invention. Detailed Implementation

[0032] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0033] like Figure 1 As shown, the code repair method based on preference learning of the present invention specifically includes the following steps:

[0034] 1) Construct a code dataset D1 for code repair learning. Each piece of code in D1 contains a programming task Q, its error code C, unit test cases T, and correct code S. For each piece of code in D1, the programming task Q corresponds to several different error codes C, and each error code C corresponds to several unit test cases T and one correct code S. Run each unit test case T for each error code C, and select one unit test case T that fails to run and add it to the code dataset. In specific implementation, the programming task Q is: "Given a character s containing only lowercase letters, find the length of the longest consecutive substring of the same character and return the length. The length of the string does not exceed 10."5 The optimal repair code can be obtained by inputting the programming task Q and error code C into the code repairer M2.

[0035] 2) The Large Language Model (LLM) is efficiently fine-tuned using the code dataset D1 to obtain the initial code repairer M1 for code fixing. First, several first training texts are constructed from the code dataset D1. For each piece of code data in D1, the first training text includes explanatory text and textual representations of the programming task Q, error code C, and correct code S. Then, the code dataset D1 and its first training texts are input into the Large Language Model (LLM), and the LoRa efficient fine-tuning method is used to train the weights only on the additional parameter matrix of the LLM. The log-likelihood function is used as the final loss function, i.e., the negative log-likelihood function. Backpropagation is then used to adjust the weights and biases of the LLM. During training, the set of model parameters corresponding to the minimum final loss function is used as the final model parameters of the LLM, ultimately obtaining the initial code repairer M1 for code fixing. In practice, for each piece of code data in the code dataset D1, the first training text is constructed in the following format: "Instructions: Given a programming task and an error code, please provide the code to fix it. Programming task: Q, Error code: C, Fix code: S." CodeLlama is used as the pre-trained LLM model in this implementation.

[0036] 3) The programming task Q and error code C of each code data in the code dataset D1 are processed by the initial code repairer M1 to obtain several candidate repair codes for error code C. After selecting each candidate repair code using the unit test example T of error code C, a preference learning dataset D2 is constructed. First, several second training texts are constructed for each programming task Q and error code C in the code dataset D1. The second training texts include explanatory text and textual representations of programming task Q and error code C. Then, each programming task Q and error code C in the code dataset D1, along with their second training texts, are input into the initial code repairer M1 for processing. After processing, several candidate repair codes for error code C are output. For each unit test example T of error code C, the unit test example T is input into each candidate repair code of error code C and run. Several candidate repair codes R that pass and fail are obtained. Then, a "git diff" operation is performed on each candidate repair code R and its corresponding error code C using Git tools to obtain the total number L of unmodified and undeleted lines of code in candidate repair code R and the total number of lines in candidate repair code R. And the total number of lines L of error codes in the candidate repair code R. cThus, the retention rate (Retention) after fixing the i-th error code C is obtained. i , Select a set of preference pairs from each candidate repair code R. <R w ,R l >, R w R is the candidate repair code R with the highest retention rate among all candidate repair codes R that have passed the test. l To select the candidate repair code R with the lowest retention rate among all failed candidate repair codes R, and to determine it as the best and worst repair code respectively, the programming task Q, error code C, and preference pairs are constructed as preference data {Q,C,R}. w ,R l Finally, the preference data for each error code C was selected and constructed into a preference learning dataset D2.

[0037] 4) Train the initial code repairer M1 using the preference learning dataset D2 to obtain the final code repairer M2 for code optimization and repair. The final code repairer M2 has the ability to modify a small range of code. First, construct several third training texts from the preference learning dataset D2, and then select the optimal repair code R for each preference data in the preference learning dataset D2. w And worst-case fix code R l Optimal fix code R w The third training text includes explanatory text, programming task Q, error code C, and optimal fix code R. w The text indicates that the worst-case fix code is R. l The third training text includes explanatory text, programming task Q, error code C, and worst-case fix code R. l The text represents this.

[0038] The preference learning dataset D2 and its various third training texts are input into the initial code repairer M1, and trained using the Direct Preference Optimization (DPO-Positive) algorithm. The loss function of the DPO-Positive algorithm is used as the final loss function, with the hyperparameter set to 0.5 in this implementation. Then, backpropagation is used to adjust the weights and biases of the additional matrix of the initial code repairer M1. During training, the set of model parameters corresponding to the minimum final loss function is used as the final model parameters of the initial code repairer M1, ultimately obtaining the final code repairer M2 for code optimization and repair. Specifically, the optimal repair code R... w The third training text is: "Instructions: Given a programming task and an error code, please provide the fix code. Programming task: Q, Error code: C, Optimal fix code: R" w "; Worst-case fix code R lThe third training text is: "Given a programming task and an error code, please provide the fix code. Programming task: Q, Error code: C, Worst-case fix code: R" l .

[0039] 5) The code data to be repaired is processed by the final code repairer M2. After processing, the final code repairer M2 generates repair code with a small modification scope and displays it on the screen, thus achieving code repair. The code data to be repaired includes the programming task Q′ to be repaired and its error code C′. A repair text is constructed for the code data to be repaired, which includes explanatory text and textual representations of the programming task Q and error code C. The code data to be repaired and its repair text are input into the final code repairer M2 for processing to generate repair code with a small modification scope, which is then displayed on the screen. In specific implementation, the repair text is: "Explanation: Given a programming task and an error code, please provide the repair code. Programming task: Q, Error code: C".

[0040] like Figure 2 As shown, this invention uses a code repairer M1 to repair the erroneous code C, generating five candidate codes, and then selects the preferred pair from them. <R w ,R l >, where R w The candidate fix code that passed correctly and had the highest retention rate (91.67%) was selected by R. l The candidate repair code with the lowest retention rate (20%) is the one with the most errors.

[0041] like Figure 3 Given a programming task Q, an error code C, and a fix code S, the model is fine-tuned using pre-trained weights W via the LoRa tuning method. The model weights W are adjusted by adding random noise A = N(0, σ). 2 The model is optimized by adjusting the offset B=0 and updating the offset B=0. N() represents the normal distribution function, σ represents the standard deviation of the noise, and the weight increment ΔW is generated to obtain the code fixer M1.

[0042] The performance of the method of the present invention on the adjusted preference learning dataset is shown in Table 1 below:

[0043] Table 1

[0044]

[0045] Among them, Fine-Tuning, Few-shot Learning, and CoT (Chain of Thought) are commonly used optimization methods for large models; Accuracy (ACC) ∈ [0,1], the larger the value of Accuracy, the higher the proportion of completely correct error localization; Improvement metric ∈ [0,1], the larger the value, the higher the proportion of code passing test points; Consistency metric ∈ [0,1], the larger the value, the better the consistency before and after code fixes;

[0046] As can be seen, compared with the Fine-Tuning model, the metrics ACC, Improvement, and Consistency of this invention on the same dataset are improved by 5.69%, 24.46%, and 23.38%, respectively; compared with Few-shot Learning, the metrics ACC, Improvement, and Consistency of this invention on the same dataset are improved by 5.46%, 25.30%, and 24.02%, respectively; and compared with the Few-shot Learning model, the metrics ACC, Improvement, and Consistency of this invention on the same dataset are improved by 11.06%, 22.02%, and 5.16%, respectively.

[0047] This invention also designs a code repair system, including a code repair enhancement module, a repair preference data generation module, a code preference learning module, and a code repair generation module. The code repair enhancement module trains a large language model (LLM) based on a code dataset D1 used for code repair learning to construct an initial code repairer M1 for code repair. The repair preference data generation module generates several candidate repair codes from the programming task Q and error code C of each code data in the code dataset D1 through the initial code repairer M1, and then constructs a preference learning dataset D2 after preference data selection. The code preference learning module trains the initial code repairer M1 through the preference learning dataset D2 to obtain a final code repairer M2 for code optimization and repair. The code repair generation module uses the final code repairer M2 to generate repair codes with a small modification range for the code data to be repaired and displays them on the display.

[0048] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented using various computer languages. This application is described with flowcharts of methods, systems, and computer program products according to embodiments of this application.

Claims

1. A code repair method based on preference learning, characterized in that, include: 1) Construct a code dataset D1 for code repair learning. Each piece of code data in the code dataset D1 contains a programming task Q and its error code C, a unit test example T, and a correct code S; 2) Fine-tune the training of the large language model LLM using the code dataset D1 to obtain the initial code repairer M1 for code repair; 3) The programming task Q and error code C of each code data in the code dataset D1 are processed by the initial code repairer M1 to obtain several candidate repair codes for error code C. After selecting each candidate repair code through the unit test sample T of error code C, the preference learning dataset D2 is constructed. 4) Train the initial code fixer M1 using the preference learning dataset D2 to obtain the final code fixer M2 for code optimization and repair; 5) The code data to be repaired is processed by the final code repairer M2. After processing, the final code repairer M2 generates repair code with a small scope of modification and displays it on the screen to achieve code repair.

2. The code repair method based on preference learning according to claim 1, characterized in that: In step 1), for each piece of code data in the code dataset D1, there is a programming task Q, which corresponds to several different error codes C. Each error code C corresponds to several unit test cases T and a correct code S. Each unit test case T of each error code C is run, and one of the unit test cases T that fails to run is added to the code dataset.

3. The code repair method based on preference learning according to claim 1, characterized in that: In step 2), firstly, several first training texts of code dataset D1 are constructed. For each piece of code data in code dataset D1, the first training text of the code data includes explanatory text and textual representations of programming task Q, error code C, and correct code S. Then, code dataset D1 and its various first training texts are input into the Large Language Model (LLM), and the LoRa efficient fine-tuning method is used to train the weights only on the additional parameter matrix of the Large Language Model (LLM). At the same time, the log-likelihood function is used as the final loss function. Then, error backpropagation is used to adjust the weights and biases of the Large Language Model (LLM). During the training process, the set of model parameters corresponding to the minimum final loss function is used as the final model parameters of the Large Language Model (LLM), and finally, the initial code repairer M1 for code repair is obtained.

4. The code repair method based on preference learning according to claim 1, characterized in that: In step 3), firstly, several second training texts are constructed for each programming task Q and error code C in the code dataset D1. The second training texts include explanatory text and textual representations of programming task Q and error code C. Then, each programming task Q and error code C in the code dataset D1 and its second training texts are input into the initial code repairer M1 for processing. After processing, several candidate repair codes for error code C are output.

5. The code repair method based on preference learning according to claim 1, characterized in that: In step 3), for each unit test case T of error code C, the unit test case T is input into each candidate repair code of error code C and run to obtain several candidate repair codes R that pass and fail. Then, the Git tool is used to operate on each candidate repair code R and its corresponding error code C to obtain the total number L of unmodified and undeleted code in candidate repair code R and the total number of lines of candidate repair code R. And the total number of lines L of error codes in the candidate repair code R. c Thus, the retention rate (Retention) after fixing the i-th error code C is obtained. i , Select a set of preference pairs from each candidate repair code R. <R w ,R l >, R w R is the candidate repair code R with the highest retention rate among all candidate repair codes R that have passed the test. l To select the candidate repair code R with the lowest retention rate among all failed candidate repair codes R, and to determine it as the best and worst repair code respectively, the programming task Q, error code C, and preference pairs are constructed as preference data {Q,C,R}. w ,R l Finally, the preference data for each error code C was selected and constructed into a preference learning dataset D2.

6. The code repair method based on preference learning according to claim 5, characterized in that: In step 4), several third training texts of the preference learning dataset D2 are first constructed, and the optimal repair code R is then developed for each preference data in the preference learning dataset D2. w And worst-case fix code R l Optimal fix code R w The third training text includes explanatory text, programming task Q, error code C, and optimal fix code R. w The text indicates that the worst-case fix code is R. l The third training text includes explanatory text, programming task Q, error code C, and worst-case fix code R. l The textual representation; The preference learning dataset D2 and its various third training texts are input into the initial code repairer M1, and trained using the Direct Preference Optimization (DPO-Positive) algorithm. The loss function of the DPO-Positive algorithm is used as the final loss function. Then, backpropagation is used to adjust the weights and biases of the additional matrix of the initial code repairer M1. During the training process, the set of model parameters corresponding to the minimum final loss function is used as the final model parameters of the initial code repairer M1, thus obtaining the final code repairer M2 for code optimization and repair.

7. The code repair method based on preference learning according to claim 1, characterized in that: In step 5), the code data to be repaired includes the programming task Q′ to be repaired and its error code C′. A repair text is constructed for the code data to be repaired, which includes explanatory text and textual representations of the programming task Q and the error code C. The code data to be repaired and its repair text are input into the final code repairer M2 for processing to generate a repair code with a small modification range, which is then displayed on the monitor.

8. A code repair system applicable to the method described in any one of claims 1-7, characterized in that, include: The code repair enhancement module trains the large language model LLM based on the code dataset D1 used for code repair learning, and constructs the initial code repairer M1 for code repair. The preference data generation module is used to generate several candidate repair codes from the programming task Q and error code C of each code data in the code dataset D1 through the initial code repairer M1, and then construct the preference learning dataset D2 after selecting the preference data. The code preference learning module trains the initial code repairer M1 on the preference learning dataset D2 to obtain the final code repairer M2 for code optimization and repair. The code repair generation module, through the final code repairer M2, generates minor repair code for the code data to be repaired and displays it on the monitor.

9. An electronic device, characterized in that, include: A memory and a processor are coupled to each other, wherein the memory stores program data, and the processor invokes the program data to perform the method as described in any one of claims 1-8.

10. A computer-readable storage medium storing program data thereon, characterized in that, When the program data is executed by the processor, the method as described in any one of claims 1-8 is implemented.