Method, device, equipment and product for generating code patch

By generating and screening candidate patches and leveraging a combination of regression testing and multilingual models, we address the problem of insufficient global understanding and verification capabilities of large language models in warehouse-level software issues, achieving efficient and accurate patch generation and selection.

CN120803894APending Publication Date: 2025-10-17BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511062385.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

When dealing with complex repository-level software issues, existing technologies make it difficult for large language models to perform global understanding, cross-file reasoning, and verification, resulting in insufficient efficiency and accuracy in patch generation. Existing methods lack persistent memory of code repositories and tool integration capabilities, making it difficult to track dependencies and perform verification steps.

Method used

By generating multiple candidate patches, using regression testing and static analysis to remove invalid patches, and combining multilingual models and tool integration, we build cross-file reasoning and global understanding capabilities to select the optimal solution.

Benefits of technology

It improves the efficiency and accuracy of searching for the optimal solution among candidate patches, enhances the understanding of the code repository, reduces redundancy and errors, and improves the automation and accuracy of patch generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803894A_ABST
    Figure CN120803894A_ABST
Patent Text Reader

Abstract

The invention relates to a method, a device, equipment and a product for generating a code patch. The method includes generating a plurality of candidate patches for repairing a software problem of a target code. The method also includes determining a plurality of screened candidate patches by removing invalid candidate patches from the plurality of candidate patches, where the invalid candidate patches are patches that do not pass the regression test. In addition, the method also includes selecting a target patch from the plurality of screened candidate patches.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of computer programming, and more specifically, to methods, apparatuses, devices, and products for generating a code patch. BACKGROUND

[0002] In the field of software development and maintenance, software issue resolution tasks are attracting more and more attention. Software issue resolution generally covers the identification, classification, and subsequent handling and repair of newly reported bugs or functional requirements. In some related technologies, the process of software issue resolution highly relies on manual operations, which not only consumes time and effort, but also is prone to delay or error caused by human factors. As the complexity of software systems continues to rise, manual handling of issues has been difficult to meet the needs of efficient development and rapid iteration. Therefore, developers are exploring the use of automated techniques to handle software issues to improve issue response efficiency, ensure system behavior correctness, and reduce development costs.

[0003] In recent years, research on automated software issue resolution technology has made significant progress. With the development of technologies such as machine learning and natural language processing, some related technologies have been able to achieve automatic repair of software issues. These technologies not only reduce the burden on developers, but also improve the performance and stability of software projects. SUMMARY

[0004] In a first aspect of embodiments of the present disclosure, a method for generating a code patch is provided. The method includes generating a plurality of candidate patches for fixing a software issue of a target code. The method further includes determining a plurality of filtered candidate patches by removing invalid candidate patches from the plurality of candidate patches, wherein an invalid candidate patch is a patch that fails a regression test. In addition, the method further includes selecting a target patch from the plurality of filtered candidate patches.

[0005] In a second aspect of embodiments of the present disclosure, an apparatus for generating a code patch is provided. The apparatus includes a candidate patch generation module configured to generate a plurality of candidate patches for fixing a software issue of a target code. The apparatus further includes a candidate patch screening module configured to determine a plurality of filtered candidate patches by removing invalid candidate patches from the plurality of candidate patches, wherein an invalid candidate patch is a patch that fails a regression test. In addition, the apparatus further includes a target patch selection module configured to select a target patch from the plurality of filtered candidate patches.

[0006] In a third aspect of embodiments of the present disclosure, an electronic device is provided. The electronic device includes one or more processors; and a storage device storing one or more programs, when executed by the one or more processors, cause the one or more processors to implement a method for generating a code patch. The method includes generating a plurality of candidate patches for fixing a software issue of a target code. The method further includes determining a plurality of filtered candidate patches by removing invalid candidate patches from the plurality of candidate patches, wherein an invalid candidate patch is a patch that fails a regression test. In addition, the method further includes selecting a target patch from the plurality of filtered candidate patches.

[0007] In a fourth aspect of embodiments of the present disclosure, a computer program product is provided. The computer program product is tangibly stored on a non-transitory computer readable medium and includes machine executable instructions that, when executed, cause a machine to implement a method for generating a code patch. The method includes generating a plurality of candidate patches for fixing a software issue of a target code. The method further includes determining a plurality of filtered candidate patches by removing invalid candidate patches from the plurality of candidate patches, wherein an invalid candidate patch is a patch that fails a regression test. In addition, the method further includes selecting a target patch from the plurality of filtered candidate patches.

[0008] The summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. The summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS

[0009] The above and other features, aspects and advantages of various embodiments of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. In the drawings similar elements are denoted by similar reference numerals, wherein:

[0010] Figure 1 A schematic diagram illustrating an example environment in which various embodiments of the present disclosure can be implemented is shown;

[0011] Figure 2 A flowchart illustrating a method for generating a code patch according to some embodiments of the present disclosure is shown;

[0012] Figure 3 A schematic diagram illustrating an example system for generating a code patch according to some embodiments of the present disclosure is shown;

[0013] Figure 4 A schematic diagram illustrating an example patch generation module according to some embodiments of the present disclosure is shown;

[0014] Figure 5 A schematic diagram showing an example patch pruning module according to some embodiments of the present disclosure is shown;

[0015] Figure 6 A schematic diagram showing an example patch selection module according to some embodiments of the present disclosure is shown;

[0016] Figure 7 A schematic diagram showing an example of constructing a prompt word for a selection agent according to some embodiments of the present disclosure is shown;

[0017] Figure 8 A schematic diagram showing an example of a selection agent utilizing a toolset to generate and execute new test cases according to some embodiments of the present disclosure is shown;

[0018] Figure 9 A block diagram of an apparatus for generating code patches according to some embodiments of the present disclosure is shown;

[0019] Figure 10 A block diagram of an apparatus capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0020] It can be understood that all user-related data involved in the present technical solution should be obtained and used after authorization by the user. This means that in the present technical solution, if the user's personal information needs to be used, the user's explicit consent and authorization are required before obtaining these data, otherwise the relevant data collection and use will not be carried out. It should also be understood that in the implementation of the present technical solution, relevant laws and regulations should be strictly followed in the collection, use and storage of data, and necessary technical and measures should be taken to protect the security of the user's data and ensure the safe use of the data.

[0021] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.

[0022] In the description of embodiments of the present disclosure, the term "comprising" and similar terms are to be interpreted as open-ended including, i.e., "including but not limited to". The term "based on" is to be interpreted as "based, at least in part, on". The term "one embodiment" or "the embodiment" is to be interpreted as "at least one embodiment". The terms "first", "second", etc. can refer to different or identical objects unless explicitly stated otherwise. Other explicit and implicit definitions can also be included below.

[0023] Currently, large language models (LLMs) exhibit strong capabilities in function-level code tasks and are widely applied in code generation and automated software bug fixing. Some related techniques achieve high bug solving rates in function-level benchmarks. However, when the task is extended to more challenging repository-level software bugs, the performance of LLMs drops significantly. This performance difference reflects the complexity of real-world software bugs. In solving software bugs, LLMs need to have a global understanding of large code repositories, be able to perform cross-file reasoning, and be able to detect subtle defects between multiple components. These challenges limit the practical application of LLMs in real-world software engineering scenarios. Therefore, building a robust, scalable, and highly reliable automated software bug fixing system has become one of the important research directions in the industry.

[0024] To narrow the gap between function-level and repository-level bug solving capabilities, some related techniques optimize the patch generation process by carefully designing agent architectures and integrating external tools to assist LLMs in generating correct solutions. For example, some related techniques introduce a problem description-based planning module and coordinate the collaboration between multiple tools in the agent system to guide patch generation. However, although the overall solving rate of these techniques remains stable across multiple runs, the problems successfully solved in each run differ significantly. This difference is due to the large-scale action space and complex reasoning trajectories in repository-level bugs, which cause LLMs to explore different solution paths in each reasoning.

[0025] In addition, some related techniques use prompt-based mechanisms to verify generated patches. These techniques use prompts to make LLMs compare generated patches with descriptions of software bugs to predict whether the patch can solve the software bug. However, prompt-based methods struggle to perform optimal solution search in large-scale integration spaces. As the number of generated patches increases, it becomes difficult for LLMs to identify subtle behavioral differences between patches. Since patch verification is usually done through a single prompt, superficial syntactic comparisons often fail to capture deep semantic differences. Furthermore, these methods lack the ability to understand the context at the repository level. In real-world software systems, bug fixing may involve multiple files and modules, requiring cross-file reasoning, global dependency tracking, and context-aware verification. However, current methods operate in a stateless, single-round interaction mode, which lacks persistent memory and tool integration capabilities, making it difficult to track dependencies, perform verification steps, or build a global understanding of code repositories.

[0026] To this end, embodiments of the present disclosure provide a scheme for generating a code patch. In the scheme, a processing device can generate a plurality of candidate patches for fixing a software issue of a target code in a code repository. Then, the processing device can determine a plurality of filtered candidate patches by removing invalid candidate patches from the plurality of candidate patches, where an invalid candidate patch is a patch that fails a regression test. Then, the processing device can select a target patch from the plurality of filtered candidate patches.

[0027] In this way, the processing device is able to reduce the size of the candidate patches by removing invalid candidate patches that fail the regression test while preserving high-quality candidate patches, thereby being able to improve the efficiency of searching for an optimal solution in the filtered candidate patches and improve the accuracy of the selected target patch.

[0028] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. As shown in FIG. 1, the environment 100 includes a code repository 110, a regression test system 120, and a processing device 130. Figure 1As shown, the environment 100 includes a processing device 102, which can be any device having processing or computing capabilities. For example, the processing device 102 can be a cloud server, a local server, a virtual server, a desktop computer, a laptop computer, a tablet computer, a smart phone, a personal assistant, or a smart wearable device, etc. The environment 100 also includes a code repository 104, which can be a storage space for storing project code of various projects. In some embodiments, the code repository 104 can be a storage space on the processing device 102. In some embodiments, the code repository 104 can be a storage space on other storage devices. Stored in the code repository 104 is project code 106, which can be stored in a single file or in multiple files. For example, taking a video platform project as an example, the code repository 104 can include multiple files, in which the project code 106 of the project is stored. In the environment 100, a software issue 108 is associated with the code repository 104. The software issue 108 can be a software bug or a functional requirement for the project code 106. The software issue 108 can include a problem description text in natural language form. In some implementations, the software issue 108 can also include code associated with the issue. For example, on a code management platform, a developer can create his own code repository 104 and upload project code 106 into the code repository 104. The developer can authorize users of the code management platform to view, download, and use the project code 106 in the code repository 104. These users can encounter bugs when using the project code 106 or the functionality implemented by the project code 106 cannot meet their requirements. Therefore, these users can describe the software issue 108 for the project code 106 on the code management platform and attach the associated code in the hope that the creator of the code repository 104 can update and iterate the project code 106 accordingly after seeing the software issue 108, so as to fix the software issue 108.

[0029] In the environment 100, the processing device 102 can generate a plurality of candidate patches 112-1, 112-2, …, and 112-N (collectively referred to as candidate patches 112) for fixing the software issue 108 of the project code 106. For example, the processing device 102 can utilize the large language model 110 (also referred to as a first large language model) and generate the plurality of candidate patches 112 based on the project code 106 and the software issue 108. For example, the processing device 102 can generate at least a portion of the candidate patches 112 by running the large language model 110 multiple times. The processing device 102 can also generate at least a portion of the plurality of candidate patches 112 by setting a particular parameter of the large language model 110 to a value that satisfies a predetermined condition (e.g., a value that is greater than a predetermined threshold) and running the large language model 110 multiple times. The processing device 102 can also generate at least a portion of the plurality of candidate patches 112 by setting the particular parameter of the large language model 110 to different values. The processing device 102 can also utilize a plurality of large language models that include the large language model 110 to respectively generate at least a portion of the plurality of candidate patches 112.

[0030] In the environment 100, the processing device 102 can determine a plurality of filtered candidate patches 116-1, 116-2, …, and 116-N (collectively referred to as candidate patches 116) by removing invalid candidate patches from the plurality of candidate patches 112. To increase the likelihood that the generated candidate patches 112 include an optimal candidate patch that can resolve the software issue 108, the processing device 102 can generate as many candidate patches 112 as possible. However, a large number of candidate patches 112 can result in a decrease in efficiency and accuracy of the processing device 102 to directly search for the optimal solution among the plurality of candidate patches 112. Accordingly, the processing device 102 can filter (also referred to as prune) the plurality of candidate patches 112 to remove invalid candidate patches from the plurality of candidate patches 112. After removing the invalid candidate patches, the remaining candidate patches from the plurality of candidate patches 112 are referred to as the filtered candidate patches 116. For example, the processing device 102 can utilize the large language model 114 (also referred to as a second large language model) to invoke a testing tool to perform a regression test on each candidate patch from the plurality of candidate patches 112. If a candidate patch 112 fails the regression test, it can be determined to be an invalid candidate patch.

[0031] In the environment 100, the processing device 102 can select the target patch 120 from the plurality of filtered candidate patches 116. For example, the processing device 102 can utilize the large language model 118 (also referred to as the third large language model) to select the target patch 120 from the plurality of filtered candidate patches 116. For example, the processing device 102 can utilize the large language model 118 to perform static analysis on each filtered candidate patch 116 to analyze whether it is able to fix the software issue 108 and whether its performance is superior to other filtered candidate patches 116. For example, the processing device 102 can also utilize the large language model 118 to perform dynamic testing on each filtered candidate patch 116 to evaluate whether it is able to fix the software issue 108 and whether its performance is superior to other filtered candidate patches 116. Then, the processing device 102 can select the candidate patch 116 whose performance is superior to other candidate patches as the target patch 120. Then, the processing device 102 (or the owner of the code repository 104) can apply the target patch 120 to the project code 106 to resolve the software issue 108.

[0032] In this way, the processing device 102 is able to reduce the size of the candidate patches by removing invalid candidate patches that fail to pass the regression test while preserving high-quality candidate patches in the plurality of candidate patches 112, thereby improving the efficiency of searching for the optimal solution in the filtered candidate patches 116 and improving the accuracy of the selected target patch 120.

[0033] Figure 2 A flowchart of a method 200 for generating a code patch is shown according to some embodiments of the present disclosure. Figure 2 A flowchart of a method 200 for generating a code patch is shown according to some embodiments of the present disclosure. The method 200 can be performed by a processing device. For example, the method 200 can be performed by the processing device 102 in Figure 1 .

[0034] As shown in Figure 2 , at block 202, the processing device can generate a plurality of candidate patches for fixing a software issue of a target code. For example, at block 202, the processing device 102 can generate the plurality of candidate patches 112 for fixing the software issue 108 of the project code 106. Figure 1In the illustrated environment 100, the processing device 102 can generate a plurality of candidate patches 112 for fixing a software issue 108 of the project code 106. For example, the processing device 102 can utilize the large language model 110 and generate the plurality of candidate patches 112 based on the project code 106 and the software issue 108. For example, the processing device 102 can generate at least a portion of the candidate patches 112 by running the large language model 110 multiple times. The processing device 102 can also generate at least a portion of the plurality of candidate patches 112 by setting a particular parameter of the large language model 110 to a value that satisfies a predetermined condition and running the large language model 110 multiple times. The processing device 102 can also generate at least a portion of the plurality of candidate patches 112 by setting the particular parameter of the large language model 110 to different values. The processing device 102 can also utilize a plurality of large language models including the large language model 110 to respectively generate at least a portion of the plurality of candidate patches 112.

[0035] At block 204, the processing device can determine a plurality of filtered candidate patches by removing invalid candidate patches from the plurality of candidate patches, where an invalid candidate patch is a patch that fails a regression test. For example, in the illustrated environment 100, the processing device 102 can determine the plurality of filtered candidate patches 116 by removing invalid candidate patches from the plurality of candidate patches 112. The processing device 102 can filter the plurality of candidate patches 112 to remove invalid candidate patches from the plurality of candidate patches 112. After removing the invalid candidate patches, the remaining candidate patches in the plurality of candidate patches 112 are referred to as the filtered candidate patches 116. For example, the processing device 102 can utilize the large language model 114 to invoke a test tool to perform a regression test on each candidate patch in the plurality of candidate patches 112. If a candidate patch 112 fails the regression test, it can be determined to be an invalid candidate patch. Figure 1 At block 206, the processing device can select a target patch from the plurality of filtered candidate patches. For example, in the illustrated environment 100, the processing device 102 can select the target patch 118 from the plurality of filtered candidate patches 116. The processing device 102 can select the target patch 118 from the plurality of filtered candidate patches 116 based on a plurality of selection criteria. For example, the processing device 102 can select the target patch 118 from the plurality of filtered candidate patches 116 based on a plurality of selection criteria including a selection criterion that is based on a number of times the large language model 114 is run to generate the target patch 118. The processing device 102 can select the target patch 118 from the plurality of filtered candidate patches 116 based on a plurality of selection criteria including a selection criterion that is based on a number of times the large language model 114 is run to generate the target patch 118.

[0036] Figure 1 ​In the illustrated environment 100, the processing device 102 may select a target patch 120 from a plurality of filtered candidate patches 116. For example, the processing device 102 may utilize a large language model 118 to select the target patch 120 from the plurality of filtered candidate patches 116. For example, the processing device 102 may utilize the large language model 118 to perform static analysis on each filtered candidate patch 116 to analyze whether it can fix the software issue 108 and whether its performance is better than other filtered candidate patches 116. For example, the processing device 102 may also utilize the large language model 118 to perform dynamic testing on each filtered candidate patch 116 to evaluate whether it can fix the software issue 108 and whether its performance is better than other filtered candidate patches 116. The processing device 102 may then select the candidate patch 116 with better performance than the other candidate patches as the target patch 120.

[0037] In this way, the processing device can reduce the size of candidate patches by removing invalid candidate patches that fail regression testing while retaining high-quality candidate patches among multiple candidate patches, thereby improving the efficiency of searching for the optimal solution among the screened candidate patches and improving the accuracy of the selected target patch.

[0038] Figure 3 FIG. 3 is a schematic diagram of an example system 300 for generating code patches according to some embodiments of the present disclosure. Figure 3 As shown, the system 300 includes a patch generation module 302, a patch pruning module 304, and a patch selection module 306. The patch generation module 302 includes an encoding agent 308 and a large language model 310 (e.g., Figure 1 An artificial intelligence (AI) agent is an AI system that simulates human intelligent behavior, and its core engine can be a large language model. The AI ​​agent is capable of perceiving its environment, making decisions, and executing actions to complete specific tasks. In the patch generation module 302, the coding agent 308 can use the large language model 310 to generate a diverse set of candidate patches, increasing the likelihood that a high-quality patch that can resolve the software issue will be found among the generated candidate patches.

[0039] The patch pruning module 304 includes a patch deduplication module 312 and a regression testing module 314. The patch pruning module 304 can build a two-layer pruning architecture by combining patch deduplication and regression testing to reduce the size of candidate patches and improve the efficiency and accuracy of searching for the optimal solution. The patch deduplication module 312 can deduplicate the multiple candidate patches generated by the patch generation module 302 to remove redundant patches from the multiple candidate patches. In the regression testing module 314, the test agent 316 can use the large language model 318 (e.g., Figure 1The large language model 114 in the image is used to perform regression testing on the candidate patches after deduplication, and the candidate patches that fail the regression test are determined as invalid patches. Then, the patch pruning module 304 can remove these invalid patches from the candidate patches after deduplication to further reduce the size of the candidate patches.

[0040] The patch selection module 306 includes a selection agent 320, a tool set 322, an operating environment 324, and a large language model 326 (e.g., Figure 1 The patch selection module 306 can identify the correct patch from the pruned candidate patches. Software problem repair in the real world usually involves multiple files and modules, which requires the system to have cross-file reasoning capabilities, context awareness capabilities, and the ability to verify the correctness of the patch across the entire code repository. Therefore, the patch selection process requires accurate code repository-level understanding. However, existing hint-based reasoning methods are usually stateless and single-round, lacking persistent memory and tool integration capabilities. These limitations hinder the ability of existing systems to track dependencies, perform verification steps, and establish a coherent and global understanding of the code repository. In the patch selection module 306, the selection agent 320 can iteratively collect and analyze relevant code snippets in the project code by calling the toolset 322, the runtime environment 324, and the large language model 326 to build a static program understanding of the project code and candidate patches. In addition, the selection agent 320 can also build a dynamic program understanding of the project code and candidate patches by automatically generating new test cases and collecting the execution results of the candidate patches on the new test cases. By combining static program understanding and dynamic program understanding, the selection agent 320 is able to select the optimal solution from the pruned candidate patches as the final target patch.

[0041] By utilizing the patch generation module 302, the patch pruning module 304, and the patch selection module 306, the processing device can improve the efficiency and accuracy of searching for the optimal solution among a large number of candidate patches. In addition, the processing device can achieve a code repository-level understanding of the project code, thereby further improving the accuracy of the selected target patch.

[0042] In some embodiments, when generating multiple candidate patches for fixing software issues in target code, the processing device can utilize a first language model based on the target code to generate multiple candidate patches for the software issues. By utilizing the first language model, the processing device can automatically generate multiple candidate patches for fixing the software issues based on the target code, thereby increasing the automation and diversity of patch generation, improving the efficiency and accuracy of software issue repairs, and reducing the burden of manual analysis.

[0043] In some embodiments, the processing device can generate multiple candidate patches by adjusting a temperature parameter of the first language model, where the temperature parameter indicates the diversity of the results generated by the language model. By adjusting the temperature parameter of the language model to control the randomness in the patch generation process, multiple candidate patches with a higher diversity can be generated, thereby covering more potential correct repair solutions and improving the robustness and success rate of the overall problem solving.

[0044] In some embodiments, the first language model is one of multiple language models, and the processing device can utilize the multiple language models to generate at least a portion of the multiple candidate patches. By utilizing multiple language models to collaboratively generate candidate patches, the advantages of different models can be leveraged to improve the diversity and coverage of patch generation, thereby enhancing adaptability to complex software issues, improving the probability of generating correct patches, and improving the overall repair effect.

[0045] Figure 4 FIG. 4 is a schematic diagram illustrating an example patch generation module 400 according to some embodiments of the present disclosure. Figure 4 As shown, the patch generation module 400 (eg, Figure 3 The patch generation module 302 in the code repository 402 can obtain project code 404 in the code repository 402 and software issues 406 for the project code 404. The coding agent 408 can use multiple large language models 410-1, 410-2, ..., and 410-N (collectively referred to as large language models 410) in a polling manner to generate multiple candidate patches 414-1, 414-2, ..., and 414-N (collectively referred to as large language models 414). The multiple large language models 410 can be different large language models. In this way, the multiple candidate patches 414 generated by the multiple different large language models 410 can have a higher diversity.

[0046] In addition, each large language model 410 may have a corresponding temperature parameter. Figure 4As shown, the large language model 410-1 has a temperature parameter 412-1, the large language model 410-2 has a temperature parameter 412-2, and so on, and the large language model 412-N has a temperature parameter 412-N (collectively referred to as temperature parameters 412). The temperature parameters 412 are hyperparameters used to control the randomness and diversity of the generated text. The values of the temperature parameters 412 can be between 0 and 1. When the values of the temperature parameters 412 are low (e.g., less than 0.3), the large language model 410 can conservatively generate results with high certainty. At this time, the large language model 410 tends to select the words with the highest probability to output more stable but less varied results. When the values of the temperature parameters 412 are high, the large language model 410 can more randomly generate diverse results that are more creative and can have lower accuracy. The encoding agent 408 can adjust the temperature parameters 412 of the large language model 410 to make the large language model 410 generate more diverse candidate patches 414. In addition, the encoding agent 408 can employ a high-temperature sampling strategy to set the temperature parameters 412 to values greater than a predetermined temperature threshold to further increase the diversity of the generated candidate patches 414. When the number of generated candidate patches 414 reaches a predetermined scale, the generation process can be terminated.

[0047] In some embodiments, the plurality of candidate patches includes a first candidate patch and a second candidate patch, and in determining the plurality of filtered candidate patches by removing invalid candidate patches from the plurality of candidate patches, the processing device can determine that a semantics of the first candidate patch is identical to a semantics of the second candidate patch. Then, the processing device can remove one of the first candidate patch and the second candidate patch from the plurality of candidate patches. By identifying and removing semantically equivalent candidate patches, the redundancy in the candidate patches can be reduced, the structure of the candidate patch set can be optimized, and thus the efficiency of the subsequent filtering and verification can be improved, the computational overhead can be reduced, and the accuracy and robustness of the patch selection process can be enhanced.

[0048] In some embodiments, the processing device can generate a structured representation of the first candidate patch and a structured representation of the second candidate patch. The processing device can generate a normalized representation of the first candidate patch by removing first semantically irrelevant elements in the structured representation of the first candidate patch that do not affect program behavior. The processing device can generate a normalized representation of the second candidate patch by removing second semantically irrelevant elements in the structured representation of the second candidate patch that do not affect program behavior. Then, the processing device can determine that the semantics of the first candidate patch are the same as the semantics of the second candidate patch based on the normalized representation of the first candidate patch and the normalized representation of the second candidate patch. By structurally parsing and semantically normalizing the candidate patches, semantically irrelevant elements that do not affect program behavior can be removed, so that the semantics equivalent patches can be more accurately identified. In this way, the processing device can improve the efficiency and accuracy of identifying redundant candidate patches, and improve the overall performance and resource utilization of the system.

[0049] In some embodiments, the first semantically irrelevant elements and the second semantically irrelevant elements include at least one of the following: redundant spaces, redundant line breaks, or comments. By removing semantically irrelevant elements such as redundant spaces, line breaks, and comments in the candidate patches, the processing device can improve the accuracy of semantic equivalence identification, which helps to more effectively identify redundant patches and simplify the candidate patch set.

[0050] In some embodiments, the plurality of candidate patches includes a third candidate patch. The processing device can obtain a regression test set associated with the target code, the regression test set including a plurality of test cases. The processing device can generate an updated code by applying the third candidate patch to the target code. In response to the updated code failing any test case in the regression test set, the processing device can determine the third candidate patch as an invalid candidate patch. Then, the processing device can remove the third candidate patch from the plurality of candidate patches. By applying the candidate patches to the target code and performing regression testing, the processing device can automatically identify invalid patches that fail the test, so that patches that may introduce new errors or damage the original functionality can be removed, improving the accuracy and reliability of patch screening, and improving the stability and credibility of the system in actual software problem repair.

[0051] In some embodiments, the processing device can obtain a candidate regression test set associated with the target code, the candidate regression test set comprising a plurality of candidate test cases. The processing device can determine, with the second language model, a first candidate test case of the plurality of candidate test cases, wherein the software issue and the first candidate test case are both associated with a same functionality of the code. Then, the processing device can determine the regression test set by removing the first candidate test case from the candidate regression test set. By utilizing the second language model to identify the test case associated with the software issue from the candidate regression test set, the processing device is able to filter out the regression test set that accurately reflects the key functionality, thereby improving the pertinence and effectiveness of the test, and enhancing the precision of the patch validation process and the overall robustness of the system.

[0052] In some embodiments, in response to a portion of the plurality of candidate patches failing any test case in the regression test set and the proportion of the portion of the plurality of candidate patches being greater than or equal to a predetermined proportion threshold, the processing device can cancel removing the portion of the candidate patches. By introducing the proportion threshold judgment mechanism when the majority of the candidate patches fail the regression test, it is possible to avoid discarding potentially valid patches prematurely due to test misjudgment, thereby improving the fault tolerance and robustness of the system, and ensuring that the possible correct solution is still retained when there is a test error or a boundary condition.

[0053] Figure 5 A schematic diagram of an example patch pruning module 500 according to some embodiments of the present disclosure is shown. The patch pruning module 500 can perform de-duplication and regression testing on the candidate patches generated by the patch generation module, thereby implementing a hierarchical pruning scheme and reducing the size of the candidate patches. As shown, the patch pruning module 500 (e.g., the patch pruning module 304 in Figure 5 Figure 3 The patch pruning module 500 includes a patch de-duplication module 504 and a regression testing module 508. The patch de-duplication module 504 can obtain a plurality of candidate patches 502-1, 502-2, …, and 502-N (collectively referred to as candidate patches 502) generated by the patch generation module (e.g., the patch generation module 400 in Figure 4 The patch de-duplication module 504 can then identify a pair of candidate patches in the plurality of candidate patches 502 that have the same semantics, and remove either candidate patch in the pair of candidate patches from the plurality of candidate patches 502. The patch de-duplication module 504 can repeat this process until there are no redundant candidate patches among the remaining candidate patches. The remaining candidate patches are denoted as candidate patches 506-1, 506-2, …, and 506-N (collectively referred to as candidate patches 506) in Figure 5

[0054] ​​In the process of deduplication, the patch deduplication module 504 can utilize a patch parser to convert the original candidate patches 502 into a structured representation. Then, the patch deduplication module 504 can generate a normalized representation of the candidate patches 502 by removing semantically irrelevant elements in the structured representation that do not affect the behavior of the program. For example, the patch deduplication module 504 can generate a normalized representation of the candidate patches 502 by removing redundant spaces, redundant line breaks, or comments in the structured representation. In this process, candidate patches 502 that cannot be parsed due to syntax errors can be determined as invalid patches and removed. The patch deduplication module 504 can determine candidate patches with the same normalized representation as semantically equivalent patches and only keep one of these candidate patches.

[0055] After obtaining the deduplicated candidate patches 506, the regression test module 508 can perform regression testing on the candidate patches 506. The goal of regression testing is to remove candidate patches with defects to further reduce the size of the candidate patches. The regression test module 508 can utilize the test agent 512 to retrieve test cases from the original code repository and perform regression testing on the candidate patches 506 to remove candidate patches that break existing functionalities. The test agent 512 can execute all test cases in the original code repository and keep the passed tests as the initial test set 510. However, not all passed tests can be categorized as regression test cases because software bug fixes can reasonably change some existing functionalities in the project code, causing some tests to fail. Therefore, the test agent 512 can utilize the large language model to identify a regression test set 514 from the initial test set 510 that is most likely representative of true regression tests. Then, each candidate patch in the deduplicated candidate patches 506 can be applied to the original project code separately to generate updated project code. The test agent 512 can utilize the updated project code to execute test cases in the regression test set 514. For any candidate patch 506, if the corresponding updated project code fails any test case in the regression test set 514, the candidate patch 506 can be determined as an invalid candidate patch and removed. Only when the updated project code passes all test cases in the regression test set 514, the corresponding candidate patch 506 can be determined as a valid candidate patch and kept. The kept candidate patches are shown as filtered candidate patches 516-1, 516-2, …, and 516-N (collectively referred to as candidate patches 516) in FIG. 5B. Figure 5

[0056] ​However, if a portion of the plurality of candidate patches 506 fails all the test cases in the regression test set 514, and the proportion of the portion of the plurality of candidate patches 506 to the plurality of candidate patches 506 is greater than or equal to a predetermined proportion threshold (e.g., 90%, 95%, or 100%), the test agent 512 can conservatively retain all the candidate patches 506, i.e., cancel the deletion of the candidate patches 506 that fail all the test cases in the regression test set 514. Since the regression test cases obtained from the original code repository can have errors, or there can be misjudgments in the pruning process of the patch pruning module 500, this way can avoid the premature removal of possible effective patches.

[0057] In some embodiments, when selecting the target patch from the plurality of screened candidate patches, the processing device can select the target patch by analyzing the target code, the software problem, and the plurality of screened candidate patches using a third language model. By comprehensively analyzing the target code, the software problem, and the screened candidate patches using the third language model, a code repository-level understanding of the semantics, context, and software problem of the patch can be built, thereby improving the accuracy of the selected target patch and improving the effectiveness of fixing the software problem.

[0058] In some embodiments, the processing device can select a plurality of analyzed candidate patches from the plurality of screened candidate patches by analyzing the target code, the software problem, and the plurality of screened candidate patches using a third language model. The processing device can use the third language model to generate a plurality of new test cases for the plurality of analyzed candidate patches. Then, the processing device can select the target patch from the plurality of analyzed candidate patches by verifying the functionality of the plurality of analyzed candidate patches using the plurality of new test cases. By using the third language model to deeply analyze the target code, the software problem, and the candidate patches and automatically generate new test cases for functionality verification, the accuracy of the patch evaluation mechanism for selecting the target patch can be improved. In addition, this way can also improve the accuracy and reliability of the selected target patch and improve the robustness of the system when facing complex or boundary software problems.

[0059] In some embodiments, the processing device can select a plurality of candidate target patches from the plurality of screened candidate patches using a plurality of selection models. Then, the processing device can select the candidate target patch with the most votes as the target patch from the plurality of candidate target patches using a majority voting strategy. By introducing a plurality of selection models to independently evaluate the candidate patches and adopting a majority voting strategy to determine the final target patch, the selection results of the plurality of selection models can be effectively integrated, the possibility of deviation caused by a single model can be reduced, thereby improving the accuracy, stability, and robustness of the overall decision of patch selection.

[0060] Figure 6 A schematic diagram of an example patch selection module 600 is shown in accordance with some embodiments of the present disclosure. As shown, the patch selection module 600 (e.g., the patch selection module 306 in FIG. 6B) can receive a plurality of filtered candidate patches 602-1, 602-2, …, and 602-N (collectively, candidate patches 602) from the patch pruning module. The patch selection module 600 includes a selection agent 604. The selection agent 604 can utilize a large language model to perform static analysis on the project code in the code repository, the description of the software issue, and the candidate patches to build a static program understanding 610 of this information. In addition, the selection agent 604 can collect code associated with the software issue, e.g., code mentioned in the software issue, project code modified by the candidate patches 602, code associated through dependency relationships, etc. The selection agent 604 can then augment the selection agent’s 604 static program understanding of the project code, the software issue, and the candidate patches by analyzing this associated code using the large language model. Figure 6 Figure 3 In addition to building the static program understanding, the selection agent 604 can also utilize the large language model to automatically generate new test cases. The selection agent 604 can then execute the new test cases against the candidate patches 602 in a runtime environment 608 by invoking a toolset 606 to build a dynamic program understanding 612 of the project code and the candidate patches 602. For example, the toolset 606 can include an agent tool to list directory contents or view file contents, an agent tool to create new files, an agent tool to edit files by inserting text, an agent tool to edit files by replacing strings, an agent tool to revert modifications to files, an agent tool to execute shell commands, etc. The selection agent 604 can interact with the agent tools in the toolset 606 through function calls. These function calls can be converted to shell commands that can be executed in a customized runtime environment 608 (e.g., a Docker environment). The results of the execution can be returned to the selection agent 604 in a structured format (e.g., JSON format) to optimize the selection agent’s 604 dynamic program understanding of the code repository and the candidate patches. The selection agent 604 can iteratively use the toolset 606 until the selection agent 604 has obtained sufficient code repository-level understanding. The selection agent 604 can then select one of the plurality of candidate patches 602 as a candidate target patch.

[0061] In addition to building the static program understanding, the selection agent 604 can also utilize the large language model to automatically generate new test cases. The selection agent 604 can then execute the new test cases against the candidate patches 602 in a runtime environment 608 by invoking a toolset 606 to build a dynamic program understanding 612 of the project code and the candidate patches 602. For example, the toolset 606 can include an agent tool to list directory contents or view file contents, an agent tool to create new files, an agent tool to edit files by inserting text, an agent tool to edit files by replacing strings, an agent tool to revert modifications to files, an agent tool to execute shell commands, etc. The selection agent 604 can interact with the agent tools in the toolset 606 through function calls. These function calls can be converted to shell commands that can be executed in a customized runtime environment 608 (e.g., a Docker environment). The results of the execution can be returned to the selection agent 604 in a structured format (e.g., JSON format) to optimize the selection agent’s 604 dynamic program understanding of the code repository and the candidate patches. The selection agent 604 can iteratively use the toolset 606 until the selection agent 604 has obtained sufficient code repository-level understanding. The selection agent 604 can then select one of the plurality of candidate patches 602 as a candidate target patch.

[0062] ​In this way, the selection agent 604 can respectively use multiple selection models to select multiple candidate target patches 614-1, 614-2, ..., and 614-N (collectively referred to as candidate target patches 614) from multiple candidate patches 602. Then, the selection agent 604 can use a majority voting strategy to select a final target patch 616 from multiple candidate target patches 614. In this way, it is possible to reduce errors caused by the hallucinations generated by the large language model and improve the consistency of the selected final target patch. For example, when N candidate target patches are given, the selection agent 604 can perform N rounds of iterations. For each round of iteration, the selection agent can use the large language model to vote for one of the N candidate target patches. The selection agent 604 can select the patch with the highest number of votes as the output result. To improve efficiency, if the first N / 2 votes reach a consensus, the selection agent 604 can immediately determine the consensus patch as the final result and skip the remaining (NN / 2) executions. When a tie occurs among multiple candidate patches after N rounds of voting, the selection agent 604 can randomly select one from the candidate target patches with the highest number of votes. In this way, the patch selection module 600 can screen out the most credible patches, thereby optimizing the software problem-solving capabilities based on the large language model.

[0063] Figure 7 FIG. 7 is a schematic diagram showing an example 700 of constructing prompt words for selecting an agent according to some embodiments of the present disclosure. Figure 7 As shown, the processing device can use the prompt generation module 702 to generate a prompt for selecting an agent (e.g., Figure 6 726 of the selection agent 604 in the example. Prompt 726 may include code repository path 704, so that the selection agent can read the project code in the code repository from code repository path 704. Prompt 726 may also include software problem 706, which is a description text in natural language and may include related code. Prompt 726 may also include candidate patch list 708, so that the selection agent can obtain the code of the candidate patch through candidate patch list 708. Prompt 726 may also include tool set 710, so that the selection agent can call various agent tools through tool set 710. Prompt 726 may also include prompt template 712.

[0064] The prompt template 712 may include a role description 714 , an example of which may be: “Serve as a code selection expert. Given a code repository, a software problem, and multiple candidate patches proposed by colleagues, select the correct patch to resolve the software problem.”

[0065] The prompt template 712 can also include an understanding task 716. An example of the understanding task 716 can be: “Understand the software issue and code repository: Carefully read the description of the software issue to understand the software issue. You need to inspect the code repository for context, including: (1) the code referenced in the description of the software issue; (2) the original code that was modified by applying the candidate patches; (3) the unchanged parts of the same file; (4) related files, functions, or modules that interact with the affected code.”

[0066] The prompt template 726 can also include a static analysis task 718. An example of the static analysis task 718 can be: “Analyze the candidate patches: For each candidate patch, analyze its logic and what it is intended to fix. Consider whether these changes align with the description of the software issue and coding standards.”

[0067] The prompt template 726 can also include a dynamic verification task 720. An example of the dynamic verification task 720 can be: “Verify functionality (optional but recommended): If needed, write and run unit tests to assess the correctness and potential side effects of each candidate patch.”

[0068] The prompt template 726 can also include a selection task 722. An example of the selection task 722 can be: “Select the best patch: Choose the patch that best addresses the software issue and minimizes the risk of introducing new software issues.”

[0069] Furthermore, the prompt template 726 can also include an example output 724. An example of the example output 724 can be: “If the correct patch is successfully selected, please submit the answer in the following format: ### Status: Success ### Result: Patch-X ### Analysis: [Explain why Patch-X is the correct choice.]”

[0070] Figure 8A diagram illustrating an example 800 of a selection agent utilizing a toolset to generate and execute new test cases is shown, in accordance with some embodiments of the present disclosure. As shown in 800, in the process of performing dynamic verification on candidate patches to build dynamic program understanding, a selection agent 802 can invoke an agent tool in a toolset 804 to generate a command for a new unit test. Then, the agent tool in the toolset 804 can generate a new test case and return it to the selection agent 802. Then, the selection agent 802 can invoke another agent tool in the toolset 804 to execute the generated new test case on the candidate patch. Then, the agent tool in the toolset 804 can execute the new test case on the candidate patch in a running environment and return the test result to the selection agent 802. In this way, the selection agent 802 is able to invoke the agent tool in the toolset 804 to automatically generate and execute new test cases, implement dynamic verification on candidate patches, and thus be able to understand the behavior of the patches in actual running more deeply. This way can enhance the comprehensiveness and accuracy of patch evaluation, help to identify potential defects or functional problems, and thus be able to improve the reliability of the selected target patch.

[0071] Figure 9 A block diagram of an apparatus 900 for generating a code patch is shown, in accordance with some embodiments of the present disclosure. As shown, the apparatus 900 includes a candidate patch generation module 902 configured to generate a plurality of candidate patches for fixing a software issue of a target code. The apparatus 900 further includes a candidate patch screening module 904 configured to determine a plurality of screened candidate patches by removing invalid candidate patches in the plurality of candidate patches, wherein an invalid candidate patch is a patch that fails a regression test. In addition, the apparatus 900 further includes a target patch selection module 906 configured to select a target patch from the plurality of screened candidate patches. Figure 9

[0072] In some embodiments, the candidate patch generation module 902 includes a first language model usage module configured to generate the plurality of candidate patches for the software issue by utilizing a first language model based on the target code.

[0073] In some embodiments, the first language model usage module includes a temperature parameter adjustment module configured to generate the plurality of candidate patches by adjusting a temperature parameter of the first language model, wherein the temperature parameter indicates a diversity of results generated by the language model.

[0074] ​In some embodiments, wherein the first language model is one of a plurality of language models, and the first language model using module comprises: a second candidate patch generating module configured to utilize the plurality of language models to respectively generate at least a portion of the plurality of candidate patches.

[0075] In some embodiments, wherein the plurality of candidate patches comprises a first candidate patch and a second candidate patch, and the candidate patch screening module 904 comprises: a semantic comparison module configured to determine that a semantic of the first candidate patch is identical to a semantic of the second candidate patch; and a first candidate patch removing module configured to remove one of the first candidate patch and the second candidate patch from the plurality of candidate patches.

[0076] In some embodiments, wherein the semantic comparison module comprises: a structured representation generating module configured to generate a structured representation of the first candidate patch and a structured representation of the second candidate patch; a normalized representation generating module configured to generate a normalized representation of the first candidate patch by removing first semantic-irrelevant elements of the structured representation of the first candidate patch that do not affect program behavior; an element removing module configured to generate a normalized representation of the second candidate patch by removing second semantic-irrelevant elements of the structured representation of the second candidate patch that do not affect program behavior; and a normalized representation using module configured to determine that the semantic of the first candidate patch is identical to the semantic of the second candidate patch based on the normalized representation of the first candidate patch and the normalized representation of the second candidate patch.

[0077] In some embodiments, wherein the first semantic-irrelevant elements and the second semantic-irrelevant elements comprise at least one of: redundant spaces, redundant line breaks, or comments.

[0078] In some embodiments, wherein the plurality of candidate patches comprises a third candidate patch, and the candidate patch screening module 904 comprises: a regression test set obtaining module configured to obtain a regression test set associated with the target code, the regression test set comprising a plurality of test cases; an updated code generating module configured to generate an updated code by applying the third candidate patch to the target code; a regression test set using module configured to determine the third candidate patch as the invalid candidate patch in response to the updated code failing any test case in the regression test set; and a second candidate patch removing module configured to remove the third candidate patch from the plurality of candidate patches.

[0079] In some embodiments, the apparatus 900 further includes a candidate regression test set obtaining module configured to obtain a candidate regression test set associated with the target code, the candidate regression test set including a plurality of candidate test cases; a second language model using module configured to determine, using a second language model, a first candidate test case of the plurality of candidate test cases, wherein the software issue and the first candidate test case are both associated with a same functionality of the code; and a regression test set determining module configured to determine the regression test set by removing the first candidate test case from the candidate regression test set.

[0080] In some embodiments, the apparatus 900 further includes a candidate patch retrieving module configured to, in response to a portion of the plurality of candidate patches failing any test case in the regression test set and a ratio of the portion of the plurality of candidate patches being greater than or equal to a predetermined ratio threshold, cancel removing the portion of the plurality of candidate patches.

[0081] In some embodiments, wherein the target patch selecting module 906 includes a third language model using module configured to select the target patch by analyzing the target code, the software issue, and the plurality of filtered candidate patches using a third language model.

[0082] In some embodiments, wherein the third language model using module includes a candidate patch analyzing module configured to select a plurality of analyzed candidate patches from the plurality of filtered candidate patches by analyzing the target code, the software issue, and the plurality of filtered candidate patches using the third language model; a new test case generating module configured to generate a plurality of new test cases for the plurality of analyzed candidate patches using the third language model; and a new test case using module configured to select the target patch from the plurality of analyzed candidate patches by verifying functionality of the plurality of analyzed candidate patches using the plurality of new test cases.

[0083] In some embodiments, wherein the target patch selecting module 906 includes a candidate target patch selecting module configured to select a plurality of candidate target patches from the plurality of filtered candidate patches using a plurality of selection models; and a voting module configured to select a candidate target patch with a most number of votes from the plurality of candidate target patches as the target patch using a plurality voting strategy.

[0084] It can be understood that the apparatus 700 of the present disclosure can achieve at least one of the advantages achievable by the methods or processes described above. For example, in this manner, the processing device can reduce the size of the candidate patches by removing invalid candidate patches that fail regression testing while retaining high-quality candidate patches from a plurality of candidate patches, thereby improving the efficiency of searching for the optimal solution among the screened candidate patches and increasing the accuracy of the selected target patch.

[0085] Figure 10 1 is a block diagram of a device 1000 capable of implementing various embodiments of the present disclosure. Figure 1 The processing device 102 is shown. Figure 10 As shown, the device 1000 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 1001, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 1002 or computer program instructions loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The CPU / GPU 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004. Although not shown in FIG. Figure 10 As shown in FIG, device 1000 may further include a co-processor.

[0086] Various components in device 1000 are connected to I / O interface 1005, including: an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0087] The various methods or processes described above may be performed by the CPU / GPU 1001. For example, in some embodiments, the methods may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the CPU / GPU 1001, one or more steps or actions in the methods or processes described above may be performed.

[0088] In some embodiments, the methods and processes described above can be tied to a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions thereon for performing various aspects of the present disclosure.

[0089] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0090] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0091] Computer readable program instructions for carrying out operations of the present disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including object oriented programming languages and conventional procedural programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0092] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0093] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0094] The computer program product of the second aspect can include a computer readable storage medium. The computer readable storage medium can include instructions. The instructions can include one or both of: instructions for causing a computer to enable a user equipment device to receive a configuration message from a base station, the configuration message comprising an indication of a set of one or more parameters for a first type of hybrid automatic repeat request process, the first type of hybrid automatic repeat request process being associated with a first type of data; and instructions for causing a computer to enable a user equipment device to receive a configuration message from a base station, the configuration message comprising an indication of a set of one or more parameters for a first type of hybrid automatic repeat request process, the first type of hybrid automatic repeat request process being associated with a first type of data.

[0095] Embodiments of the present disclosure have been described above, with the understanding that these embodiments are exemplary only, and are not restrictive, and are not limited to the disclosed embodiments. Many modifications and changes to the described embodiments are possible, without departing from the scope and spirit of the described embodiments. The selection of terms to be used herein is intended to best explain the principles of the embodiments, practical application, or technical improvement over the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A method for generating a code patch, comprising: generating a plurality of candidate patches for fixing software problems in the target code; Determine a plurality of filtered candidate patches by removing invalid candidate patches from the plurality of candidate patches, wherein the invalid candidate patches are patches that fail the regression test; as well as A target patch is selected from the plurality of filtered candidate patches.

2. The method of claim 1 , wherein generating the plurality of candidate patches for fixing the software problem of the target code comprises: Based on the target code, a first language model is used to generate the plurality of candidate patches for the software problem.

3. The method according to claim 2, wherein generating the plurality of candidate patches for the software problem based on the target code using the first language model comprises: The plurality of candidate patches are generated by adjusting a temperature parameter of a first language model, wherein the temperature parameter indicates diversity of results generated by the language model.

4. The method according to claim 2, wherein the first language model is one of a plurality of language models, and generating the plurality of candidate patches for the software problem using the first language model based on the target code comprises: The multiple language models are used to respectively generate at least a portion of the multiple candidate patches.

5. The method according to claim 1 , wherein the plurality of candidate patches include a first candidate patch and a second candidate patch, and determining the plurality of filtered candidate patches by removing the invalid candidate patch from the plurality of candidate patches comprises: Determining that the semantics of the first candidate patch are the same as the semantics of the second candidate patch; as well as One of the first candidate patch and the second candidate patch is removed from the plurality of candidate patches.

6. The method of claim 5 , wherein determining that the semantics of the first candidate patch are the same as the semantics of the second candidate patch comprises: generating a structured representation of the first candidate patch and a structured representation of the second candidate patch; generating a normalized representation of the first candidate patch by removing first semantically irrelevant elements that do not affect program behavior from the structured representation of the first candidate patch; generating a normalized representation of the second candidate patch by removing second semantically irrelevant elements that do not affect program behavior from the structured representation of the second candidate patch; as well as It is determined based on the normalized representation of the first candidate patch and the normalized representation of the second candidate patch that the semantics of the first candidate patch are the same as the semantics of the second candidate patch. 7 . The method of claim 6 , wherein the first semantically irrelevant element and the second semantically irrelevant element comprise at least one of the following: redundant spaces, redundant line breaks, or comments.

8. The method of claim 1 , wherein the plurality of candidate patches include a third candidate patch, and determining the plurality of filtered candidate patches by removing the invalid candidate patch from the plurality of candidate patches comprises: Obtaining a regression test set associated with the target code, the regression test set including a plurality of test cases; generating updated code by applying the third candidate patch to the target code; In response to the updated code failing any test case in the regression test set, determining the third candidate patch as the invalid candidate patch; as well as The third candidate patch is removed from the plurality of candidate patches.

9. The method according to claim 8, further comprising: Obtaining a candidate regression test set associated with the target code, the candidate regression test set including a plurality of candidate test cases; determining a first candidate test case from the plurality of candidate test cases using a second language model, wherein the software problem and the first candidate test case are both associated with a same function of the code; as well as The regression test set is determined by removing the first candidate test case from the candidate regression test set.

10. The method according to claim 8, further comprising: In response to some candidate patches among the plurality of candidate patches failing any test case in the regression test set and a ratio of the some candidate patches to the plurality of candidate patches being greater than or equal to a predetermined ratio threshold, removing the some candidate patches is canceled.

11. The method according to claim 1 , wherein selecting the target patch from the plurality of screened candidate patches comprises: The target patch is selected by analyzing the target code, the software problem, and the plurality of screened candidate patches using a third language model.

12. The method according to claim 11, wherein selecting the target patch by analyzing the target code, the software problem, and the plurality of screened candidate patches using the third language model comprises: selecting a plurality of analyzed candidate patches from the plurality of filtered candidate patches by analyzing the target code, the software problem, and the plurality of filtered candidate patches using the third language model; generating a plurality of new test cases for the plurality of analyzed candidate patches using the third language model; as well as The target patch is selected from the plurality of analyzed candidate patches by verifying functionality of the plurality of analyzed candidate patches using the plurality of new test cases.

13. The method according to claim 1, wherein selecting the target patch from the plurality of screened candidate patches comprises: selecting a plurality of candidate target patches from the plurality of screened candidate patches using a plurality of selection models; as well as A majority voting strategy is used to select a candidate target patch with the largest number of votes from the multiple candidate target patches as the target patch.

14. An apparatus for generating a code patch, comprising: a candidate patch generation module configured to generate a plurality of candidate patches for fixing software problems of a target code; a candidate patch screening module configured to determine a plurality of screened candidate patches by removing invalid candidate patches from the plurality of candidate patches, wherein the invalid candidate patches are patches that fail the regression test; as well as The target patch selection module is configured to select a target patch from the plurality of filtered candidate patches.

15. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, wherein when the instructions are executed by the processor, the electronic device performs the method according to any one of claims 1 to 13.

16. A computer program product tangibly stored on a non-transitory computer-readable medium and comprising machine-executable instructions that, when executed, cause a machine to implement the method according to any one of claims 1 to 13.