Test case generation enhancement method and automatic program repair method
Generate test cases through a large language model and combine multiple fault location methods to automatically generate and filter suspicious code snippets, and use the large language model to generate and sort patches, solving the accuracy problem of the lack of automatic program repair under triggering tests, and improving the applicability and effectiveness of the repair tool.
Patent Information
- Application Number
- CN202510474463.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-01
AI Technical Summary
In the absence of triggered tests, the fault location and patch verification methods of existing automatic program repair tools are significantly affected, especially the spectrum-based fault location tools cannot be used, and the quality of automatically generated test cases is random, affecting the repair effect.
The automatic test case generation tool based on the large language model is adopted to automatically generate multiple test cases, combine the fault location method based on spectrum and information retrieval to filter out suspicious code snippets, and use the large language model to generate patches. The automatically generated test cases are sorted and verified, and finally the repair solution is verified using the original test suite.
In the absence of trigger testing, the accuracy of fault location and the effectiveness of patch generation are improved, the impact of algorithm quality randomness is reduced, and the applicability and effectiveness of automatic program repair tools are improved.
Smart Images

Figure CN120407405A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of automatic program repair and software engineering, and in particular to a method-level automatic program repair method with enhanced test case generation. Background Art
[0002] As modern software systems become increasingly complex, involving millions of lines of code, manually detecting and fixing bugs has become extremely time-consuming and difficult. The regular rapid iteration and release of new products also require development teams to shorten the development cycle, which necessitates more efficient bug detection and repair mechanisms. Moreover, as software is iterated, the maintenance cost is gradually increasing. The need for automatic program repair has emerged. Automatic program repair is a method that uses artificial intelligence and machine learning technologies to automatically detect and fix bugs or faults in software programs. The development and application of automatic program repair technology are of great significance for improving software development efficiency, reducing costs, and ensuring software quality and security.
[0003] Automatic program repair usually adopts a "generate and verify" workflow, which consists of three parts: fault localization, patch generation, and patch verification. However, most current automatic program repair tools use trigger tests written by developers, mainly for fault localization and patch verification. A trigger test refers to a unit test case that throws an exception when executing to the program segment containing the bug.
[0004] However, relevant research has pointed out that such test cases are often written by developers after the bug is discovered or even fixed. In the initial stage when the bug is discovered, if an automatic program repair tool is used, there may be a situation of lacking trigger tests. In the case of lacking trigger tests, the existing fault localization and patch verification methods will be significantly affected. In particular, the widely used spectrum-based fault localization tools may be completely unusable. Therefore, some tools propose to use large language models to generate trigger tests to complete automatic program repair. However, these tools either only use automatically generated test cases to complete fault localization; or only use them for patch verification, ignoring the impact of test cases on the entire process. Currently, there is no method that applies the test cases generated by large language models to all stages of automatic program repair and completely does not use manually written trigger tests. The present invention is expected to improve the applicability and effectiveness of automatic program repair tools in real software development work by using automatically generated test cases for fault localization, patch generation, and patch sorting, and designing relevant algorithms to reduce the impact of the randomness of their quality. Summary of the Invention
[0005] The purpose of the present invention is to provide a method-level automatic program repair method with enhanced test case generation in view of the deficiencies of the prior art.
[0006] The object of the present invention is achieved by the following technical solutions: A method-level automatic program repair method for enhancing test case generation, comprising the following steps:
[0007] (1) Use an automatic test case generation tool to automatically generate a set of test cases with the fault report as input information, denoted as sequence T;
[0008] (2) Execute all test cases in sequence T, and filter them according to their results, and filter out the test cases that fail to execute, denoted as sequence T FIB ;
[0009] (3) Based on the test case sequence T FIB obtained in step (2), combine the spectrum-based fault localization method and the information retrieval-based fault localization method to calculate the method-level suspiciousness scores of all code fragments, and obtain the fault localization sequence L F = [(m1, s1), (m2, s2), …, (m i , s i ), …, (m n , s n )], where m i and s i are the code method name and suspiciousness score corresponding to the i-th sorted localization respectively, and n is the length of the fault localization sequence;
[0010] (4) Based on the fault localization sequence L F obtained in step (3), obtain the corresponding suspicious code fragments; among them, the suspicious code fragment corresponding to each localization is specifically the code and annotation text from the method name to the end of the method body;
[0011] (5) Combine the fault report, the suspicious code fragments obtained in step (4), a random test case text obtained in step (2), and explanatory text into a prompt, input it into a large language model, and use the large language model to generate patches to obtain a patch sequence P = [p1, p2, …, p i , …, p n′ , where p i represents the i-th patch, and n′ is the length of the patch sequence;
[0012] (6) Replace each patch in the patch sequence obtained in step (5) with the corresponding original code text one by one, and execute all test cases obtained in step (2), and sort the patches according to the execution results to obtain the final patch sequence;
[0013] (7) Use the original test suite to verify the patches in the final patch sequence, and filter out the patches that fail to execute; submit the first N patches in the filtered final patch sequence as the final repair solution.
[0014] Furthermore, the automatic test case generation tool is a method for automatically generating test cases based on a large language model. Its input information is a fault report. By using the large language model's ability to understand natural language and code text, trigger tests are generated, and multiple test cases are automatically generated at one time as a set of test cases.
[0015] Furthermore, step (3) specifically includes the following sub-steps:
[0016] (3.1) Use the information retrieval-based fault localization method to obtain the suspiciousness scores of all code snippets in the software repository, and sort all code snippets in descending order of the suspiciousness scores to obtain the sorted suspicious code sequence, denoted as where m i and are the code method name corresponding to the i-th sorted location and the suspiciousness score at the IRFL method level, respectively;
[0017] (3.2) Use the information retrieval-based fault localization method to insert all the test case sequences T obtained in step (2) FIB into the original test suite, execute all test cases and calculate the suspiciousness scores of all code snippets in the software repository, and sort all code snippets in descending order of the suspiciousness scores to obtain the sorted suspicious code sequence, denoted as where m i and are the code method name corresponding to the i-th sorted location and the suspiciousness score at the SBFL method level, respectively;
[0018] (3.3) According to the following formula, normalize the suspiciousness scores of each code method in the sorted suspicious code sequences obtained in steps (3.1) and (3.2) respectively to obtain the normalized suspiciousness scores:
[0019]
[0020] In the formula, e is the natural exponent, represents the suspiciousness score at the IRFL method level or the SBFL method level corresponding to the normalized suspiciousness score;
[0021] [[ID=4)1](3.4) Add the normalized suspiciousness scores at the IRFL method level and the SBFL method level obtained in step (3.3) to obtain the final suspiciousness score;
[0022] (3.5) Sort all code methods in descending order according to the final suspiciousness scores obtained in step (3.4) to obtain the final fault localization sequence L F = [(m1, s1), (m2, s2), …, (m i , s i ), …, (m n , s n )], where m i and s i are the code method name and suspiciousness score corresponding to the localization ranked i respectively, and n is the length of the fault localization sequence.
[0023] Further, step (6) specifically includes the following sub-steps:
[0024] (6.1) Replace each patch in the patch sequence obtained in step (5) with the corresponding original code text one by one, and execute all test cases in the sequence T FIB obtained in step (2), record the successful execution cases, denoted as (p i , N i ), where i is the patch number, p i is the i-th patch text, and N i is the number of test cases passed by the i-th patch; directly remove the patches that cause compilation failure or other abnormal reasons for execution failure;
[0025] (6.2) Based on the execution situations recorded in step (6.1), sort each patch in the patch sequence obtained in step (5), and sort them in descending order according to the number of test cases N i passed by the patches, that is, the patch passing the most test cases has the highest priority, and the patch passing the fewest test cases has the lowest priority, and so on, to obtain the sorted patch sequence;
[0026] (6.3) Based on the sorted patch sequence obtained in step (6.2), for the cases where the number of test cases passed is the same, a sorting method combining the coding similarity-based sorting method and the lemmatized entropy-based sorting method is used for sub-sorting to obtain the finally sorted patch sequence; the priorities during sorting are from high to low in turn: the number of test cases passed, the lemmatized entropy value, and the coding similarity.
[0027] The beneficial effects of the present invention are as follows: the present invention can perform automatic program repair in the case of lacking trigger tests and only having fault report information; the present invention uses large language models to generate test cases to reproduce faults, and then locates faults, generates patches and sorts patches through the automatically generated test cases; the present invention is applicable to method-level fault repair in the case of having fault reports but lacking trigger tests, and can also be used alone for certain steps such as fault location and patch sorting, reducing the influence brought by the randomness of algorithm quality, and helping to improve the applicability and effect of automatic program repair tools. Brief Description of the Drawings
[0028] Figure 1 It is a flowchart of the method-level automatic program repair method for enhancing test case generation of the present invention. Detailed Embodiments
[0029] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and cannot limit the present application.
[0030] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0031] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination". Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method including a series of elements not only includes those elements but also other elements not expressly listed, or also includes elements inherent to such process, method. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, method, article or apparatus including the said element.
[0032] The present invention will be described in detail below with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.
[0033] The present invention uses an Automated Test Case Generation Tools to generate trigger tests using a fault report and perform a preliminary screening based on the execution results; subsequently, a spectrum-based fault localization method and an information retrieval-based fault localization method are respectively used to determine a suspicious code sequence according to the execution results of the test cases and the similarity between the fault report and the code snippet; a patch is generated according to the suspicious code sequence, and after replacing the original code text with the patch one by one, all automatically generated test cases are executed, and the patches are sorted according to the execution results of the test cases; finally, a test suite is used to complete patch verification, and a final repair solution is formed and submitted to the user.
[0034] See Figure 1 , the method for enhancing test case generation of the present invention, a method-level automatic program repair method, specifically includes the following steps:
[0035] (1) Use an automated test case generation tool to automatically generate a set of test cases, denoted as sequence T, with the fault report as the input information.
[0036] Further, the automated test case generation tool is a method for automatically generating test cases based on a large language model. Its input information is the fault report, and the large language model is used to generate trigger tests based on its understanding of natural language and code text, and multiple test cases are automatically generated at one time as a set of test cases.
[0037] It should be understood that in addition to the method for automatically generating test cases based on large language models, the automatic test case generation tool can also use other automatic test case generation tools.
[0038] Specifically, the automatic test case generation tool is used to generate trigger tests according to the fault report. The present invention first uses the automatic test case generation tool to generate trigger tests according to the fault report. A trigger test refers to a unit test case that throws an exception when executing to the program segment containing the fault. Therefore, the existing fault localization tool can utilize the characteristic that the trigger test throws an exception at the fault to locate the program segment that causes the exception and complete the fault localization. However, the manually written trigger tests do not always exist prior to the discovery of the fault. Therefore, an automatic generation method is needed to ensure its existence. When implementing the present invention, an automatic test case generation method based on large language models is adopted. Its input information is the fault report, and the large language model is used to generate trigger tests by virtue of its understanding ability of natural language and code text. The success rate on the existing data set is approximately 27%. Due to the randomness of the quality of the automatically generated test cases, the present invention adopts the method of generating multiple test cases at one time to ensure the generation of available trigger tests as much as possible. All the generated test case sets are denoted as T.
[0039] (2) Execute all the test cases in sequence T and screen them according to their results. Screen out the test cases that execute failed, denoted as sequence T FIB .
[0040] Specifically, execute all the test cases in the automatically generated sequence T and screen them according to their results. The correct trigger test should throw an exception on the software version containing the fault. Since test cases are usually written during the software development and maintenance stage to ensure the normal execution of each program segment, and the test cases usually contain several assertion sentences, whose meaning is usually: if the input is XXX, then the expected output is XXX, and the actual output and the expected output are compared; if the two are not the same, a program exception will be thrown, indicating that the execution result is different from the expectation. The present invention utilizes this characteristic to execute all the automatically generated test cases on the program of the fault version, retain the test cases that throw exceptions, that is, retain the test cases that execute failed, and remove other test cases to reduce the influence brought by the randomness of their quality. The finally obtained test case set is denoted as T FIB .
[0041] (3) Based on the test case sequence T obtained in step (2) FIB, calculate the method-level suspiciousness scores of all code snippets by combining the spectrum-based fault localization (SBFL) method and the information-retrieval based fault localization (IRFL) method, and obtain the fault localization sequence T after sorting F =[(m1, s1), (m2, s2), …, (m i , s i ), …, (m n , s n )], where m i and s i are the code method name and the suspiciousness score corresponding to the i-th sorted localization respectively, and n is the length of the fault localization sequence.
[0042] It should be noted that for automatic fault repair, fault localization is required, and then it is necessary to know the code method name, method path and its line number where the fault occurs, which can be specifically completed through a fault localization tool; among them, the method path refers to the file path where the faulty method is located in the code repository, such as org / java / Bug.java, where org and java are both folder names, and Bug.java is the file name containing the faulty method. In the present invention, the two fault localization tools, namely the spectrum-based fault localization method and the information-retrieval based fault localization method, are combined to reduce the influence brought by the randomness of automatically generated test cases. The specific steps are as follows:
[0043] (3.1) Use the information-retrieval based fault localization method to obtain the suspiciousness scores of all code snippets in the software repository, and sort all code snippets in descending order of the suspiciousness scores to obtain the sorted suspicious code sequence, denoted as where m i and are the code method name and the IRFL method-level suspiciousness score corresponding to the i-th sorted localization respectively.
[0044] Specifically, use the information-retrieval based fault localization method to perform information retrieval on all code snippets in the software repository according to the fault report, where each code snippet corresponds to a code method; according to the text similarity between the fault report and the code snippet, as the suspiciousness score corresponding to the code snippet, and sort all code snippets in descending order of the suspiciousness scores, and output the sorted suspicious code sequence, denoted as
[0045] It should be understood that the fault localization method based on information retrieval is a method that transforms software engineering problems into text retrieval tasks. It locates the faulty code by analyzing the semantic similarity between the code and error reports (such as Bug descriptions, stack traces). Its core idea is that the code semantically similar to the Bug report is more likely to be the fault point, which will not be elaborated here. The accuracy of the fault localization method based on information retrieval is not high, but it does not rely on any test cases and can be used as one of the bases for judgment.
[0046] (3.2) Use the fault localization method based on information retrieval to insert all the test case sequences T obtained in step (2) FIB into the original test suite, execute all the test cases and calculate the suspiciousness scores of all code snippets in the software repository, and sort all the code snippets in descending order of the suspiciousness scores to obtain the sorted suspicious code sequence, denoted as where m i and are the name of the code method corresponding to the i-th ranked localization and the suspiciousness score at the SBFL method level respectively.
[0047] Specifically, when using the fault localization method based on information retrieval for fault localization, at least one test case that throws an exception and one test case that executes successfully to completion are required to complete the calculation. The accuracy of the calculation mainly depends on whether the test case that throws an exception covers the faulty code. Therefore, usually a manually written trigger test is required for accurate localization. In the present invention, automatically generated test cases are used, and their quality is random. Therefore, when calculating the suspiciousness scores of all code snippets using the fault localization method based on information retrieval, all the test case sequences T obtained in step (2) FIB need to be inserted into the original test suite and all the test cases are executed. In this embodiment, the preset working scenario is a software code repository equipped with a test suite, and then a fault report appears, but there is no associated trigger test, which is a test case that throws an exception at the fault location. Generally speaking, test cases that execute successfully to completion are more likely to exist; based on such logic, it may be easier to understand: if not all test cases can execute successfully to completion, then developers may be aware of the fault earlier instead of waiting until someone submits a fault report to start the repair. Thus, by the overall tendency, false localization caused by a certain incorrect test case is avoided. After obtaining the suspiciousness scores of all code snippets, all the code snippets are sorted in descending order of the suspiciousness scores to obtain the sorted suspicious code sequence, denoted as where, when calculating the suspiciousness scores of all code snippets using the fault localization method based on information retrieval, its calculation formula can be expressed as:
[0048]
[0049] In the formula, score is the suspiciousness score of the code snippet calculated using the information retrieval-based fault localization method. A is the number of test cases that are successfully executed and executed by this SBFL method. B is the number of test cases that are successfully executed but not executed by this SBFL method. C is the number of test cases that throw exceptions and are executed by this SBFL method. D is the number of test cases that throw exceptions but are not executed by this SBFL method. It should be understood that the information retrieval-based fault localization method is a method of locating software defects by analyzing the correlation between the coverage spectrum (such as line coverage, branch coverage) during program execution and the execution results (success / failure) of test cases. Its core idea is: the code that is frequently executed in failed tests but rarely executed in successful tests is more likely to be the fault point, which will not be elaborated here.
[0050] (3.3) Normalize the suspiciousness scores of each code method in the sorted suspicious code sequence obtained in steps (3.1) and (3.2) respectively according to the following formula to obtain the normalized suspiciousness scores:
[0051]
[0052] In the formula, e is the natural exponent, represents the suspiciousness score at the IRFL method level or SBFL method level the corresponding normalized suspiciousness score, When, it is to normalize the suspiciousness scores of each code method in the sorted suspicious code sequence obtained in step (3.1). When, it is to normalize the suspiciousness scores of each code method in the sorted suspicious code sequence obtained in step (3.2).
[0053] (3.4) Add the normalized suspiciousness scores at the IRFL method level and SBFL method level obtained in step (3.3) to obtain the final suspiciousness score s i , which is expressed as:
[0054]
[0055] In the formula, represents the normalized suspiciousness score at the IRFL method level, represents the normalized suspiciousness score at the SBFL method level.
[0056] (3.5) Sort all code methods in descending order according to the final suspiciousness score obtained in step (3.4) to obtain the final fault localization sequence LF = [(m1, s1), (m2, s2), …, (m i , s i ), …, (m n , s n ), ]), where m i and s i are the code method name and the suspiciousness score corresponding to the i-th sorted location respectively, and n is the length of the fault location sequence.
[0057] (4) Based on the fault location sequence L F obtained in step (3), the corresponding suspicious code segments are obtained; among them, the suspicious code segment corresponding to each location is specifically the code and comment text from the method name to the end of the method body.
[0058] Specifically, based on the fault location sequence L F obtained in step (3), the corresponding suspicious code segments can be obtained; among them, the suspicious code segment corresponding to each location is specifically a complete method code, usually including the method name, parameter names, method body, etc.; in addition, in the current software development process, developers are used to adding header comments in several lines before the method name to explain the function of the function. Therefore, the header comments are regarded as part of the method and included in the suspicious code segments.
[0059] (5) Combine the fault report, the suspicious code segments obtained in step (4), a randomly selected test case text obtained in step (2), and the explanatory text into a prompt, input it into the large language model, and use the large language model to generate patches to obtain a patch sequence P = [p1, p2, …, p i , …, p n′ , where p i represents the i-th patch, and n' is the length of the patch sequence.
[0060] Specifically, combine the fault report, the suspicious code segments obtained in step (4), a randomly selected test case text obtained in step (2), and the explanatory text into a prompt, input it into the large language model, and use the large language model to generate patches; among them, the test case sequence T FIBThe test cases in it exist in text form and are transformed into executable computer programs through compilation; in addition, some explanatory texts are required in the prompts of large language models, such as "You are an expert in Java programs. Please help me modify the following method containing faults.", "The following is the fault report:", and so on. Using large language models to generate patches is a common practice at present, but there is no definite conclusion on the construction of prompts. The present invention designs to use fault reports, suspicious code segments, explanatory texts, and a randomly generated test case as the main content of the prompt, and generate patches through large language models; due to the text length limit in the input of large language models, it is impossible to add all the test cases obtained in step (2) to the prompt. The present invention randomly selects a test case to add to the prompt during implementation; at the same time, in order to use as many test cases as possible to avoid the negative impact of the randomness of their quality, the present invention generates multiple batches of patches at one time, that is, when a prompt is input into the large language model, multiple patches can be obtained. Each batch of patches uses a test case, and the number of patches generated between each batch is the same.
[0061] (6) For each patch in the patch sequence obtained in step (5), replace the corresponding original code text one by one, and execute all the test cases obtained in step (2), and sort the patches according to the execution results to obtain the final patch sequence. Specifically, it includes the following sub-steps:
[0062] (6.1) For each patch in the patch sequence obtained in step (5), replace the corresponding original code text one by one, and execute the sequence T FIB in all the test cases obtained in step (2), obtain the execution results, record the successful execution cases, denoted as (p i , N i ), where i is the serial number of the patch, p i is the i-th patch text, and N i is the number of test cases passed by the i-th patch; the patches that fail to compile or fail to execute due to other abnormal reasons are considered incorrect patches and are directly removed from the patch sequence.
[0063] It should be noted that each patch in the patch sequence is generated for the same faulty method, and they can replace the source code (the original faulty method). Although different batches of patches use different test cases in the prompt during generation, this will not affect their consistency with the source code in form. Further explanation, it is stated in the prompt that "The following is the faulty code", and this part remains unchanged, so the large language model will also clearly output the result that should be a repair patch for the faulty code when generating patches.
[0064] It should be understood that both the patch and the fault code are a method with the same method name, the same parameters, and the same output type. Several lines of code are included in the method body, and only a few lines of code in the method body are changed between the patch and the fault code. Since the method name of the code where the fault is located, the file path containing the code method name, and its line number can be obtained in the fault location stage of step (3), therefore, the complete code of the method where the fault is located can be obtained according to the file path and the line number, and the patch can be replaced back accordingly.
[0065] (6.2) Based on the execution situation recorded in step (6.1), sort each patch in the patch sequence obtained in step (5) according to the number N of test cases passed by the patch i in descending order, that is, the patch passing the most test cases has the highest priority (corresponding to the case of N i being the largest), the patch passing the fewest test cases has the lowest priority, and so on, to obtain the sorted patch sequence.
[0066] (6.3) Based on the sorted patch sequence obtained in step (6.2), for the case where the number of test cases passed is the same, that is, N i = N j where i and j are the initial sequence numbers of two patches respectively, then a sorting method based on coding similarity and a sorting method based on token entropy are combined for sub - sorting to obtain the final sorted patch sequence; the priorities during sorting are from high to low in turn: the number of test cases passed, the token entropy value, and the coding similarity.
[0067] Specifically, for the case where the number of test cases passed is the same, that is, N i = N j where i and j are the initial sequence numbers of two patches respectively, then a sorting method based on coding similarity and a sorting method based on token entropy are combined for sub - sorting, that is, the coding similarity is obtained using the sorting method based on coding similarity, and the token entropy value is calculated using the sorting method based on token entropy. When sorting each patch in the patch sequence obtained in step (5), first sort each patch in descending order according to the number of test cases passed; secondly, for the case where the number of test cases passed is the same, that is, N i = N j sub - sort them in descending order according to the token entropy value; finally, for the case where the token entropy values are the same, sub - sort them in descending order according to the coding similarity, and so on, and sort the patches in the order of priority from high to low during sorting to obtain the final sorted patch sequence.
[0068] It should be understood that in this embodiment, the embedding model provided by OPENAI is used. A piece of code is input into the large language model, and the output is a feature vector. The cosine similarity of the feature vectors of two pieces of code is the coding similarity. Furthermore, a sorting method based on the coding similarity is used for fine-grained sorting. The token entropy can be obtained in the patch generation stage. The large language model completes the output by generating tokens one by one, and will consider the probability value of each token to determine which token should be generated next. The value obtained by taking the logarithm of this probability value is the entropy of a single token, that is, the token entropy value. Furthermore, a sorting method based on the token entropy is used for fine-grained sorting.
[0069] (7) Use the original test suite to verify the patches in the final patch sequence obtained in step (6), and filter out the patches that fail to execute; submit the first N patches in the filtered final patch sequence as the final repair solution, and N can be freely set by the user.
[0070] Specifically, use the original test suite to verify the patches. Through the original test suite, it can be confirmed whether the generated patches introduce other errors, maximizing the accuracy of the final recommended solution. Just execute all the test cases in the original test suite. If no exception is thrown, it means that the patch does not introduce other problems. Replace the patch and compile and execute it. If the compiler reports an error during the compilation stage, it means that the patch has an error; during the execution stage, if any test case throws an exception due to the patch, it means that it has modified the code logic unrelated to the fault, and since the test case throws an exception, it means that its modification is incorrect. Submit the first N patches in the final patch sequence as the final repair solution, and N can be freely set by the user; since the correctness of the final patch needs to be manually verified, N takes the values of 1, 3, and 5 when evaluating the effect of the present invention.
[0071] Exemplarily, the effect of the automatic program repair method proposed by the present invention is evaluated through ablation experiments and comparative experiments, and the effects of each module and the final overall effect are respectively investigated: by comparing the results of using two fault localization tools separately and the combined results, the effect of fault localization using automatically generated test cases is verified; by comparing the patch generation results with or without adding test cases to the prompt words, the effect of patch generation using automatically generated test cases is verified; by comparing the results of the patch sorting method based on automatically generated test cases and the existing patch sorting method to verify the effect; finally, the overall results are compared with the results of existing automatic program repair tools applicable to the same scenario to verify its effect.
[0072] The following respectively elaborates on the specific experimental methods and settings.
[0073] Experimental dataset: The Defects4J dataset was used in the experiment, which contains 835 fault repair records taken from real open-source projects. In this experiment, 374 faults with a change range within a single method were selected as the experimental dataset, and the manually written trigger tests in the fault version code repository were removed, while the remaining original test suites were retained.
[0074] Test case generation: In terms of the automatic test case generation tool, the present invention uses the tool Libro based on the large language model to generate a total of 50 test cases for each fault, and executes all the test cases on the version containing the faulty code, and retains the test cases that throw exceptions, denoted as T FIB 。
[0075] Fault localization effect: The two fault localization tools used in the experiment are BoostNSift based on information retrieval and GZoltar based on spectrum. The evaluation metric of the experiment is Top@N, which means the number of faults whose correct fault localization is included in the top N suspicious localizations. The evaluation results of the fault localization module are shown in Table 1, where IRFL represents the fault localization method that only uses information retrieval-based, SBFL represents the spectrum-based fault localization method, and ATFL represents the fault localization method proposed in the present invention that combines the two. The comparison of the three shows that the fault localization method proposed in the present invention can locate 67% and 35% more faults respectively than the information retrieval-based method and the spectrum-based method alone when there is no manually written trigger test. At the same time, compared with using manually written trigger tests, it can achieve about 82% of its effect, which means that the method described in the present invention can meet the general fault localization requirements in the absence of manually written trigger tests.
[0076] Table 1: Evaluation result table of fault localization in experimental evaluation
[0077] Indicator IRFL SBFL ATFL ATFL (using manual trigger test) Top@1 71 88 119 145 Top@3 115 139 172 193 Top@5 129 156 187 213
[0078] Patch generation effect: The present invention uses the gpt-3.5-turbo model to evaluate the patch generation effect. A total of 100 patches are generated for each fault. When evaluating the patch generation effect alone, the suspicious code selected in the prompt is the faulty code, that is, it is assumed that the fault localization is completely correct. The experiment compared the influence of adding test cases and not adding test cases in the prompt on the patch generation effect. The evaluation metric is the number of faults that can be correctly repaired considering all patches. During the experiment, the correctness of the patches was confirmed by manual inspection. When adding test cases, the number of patches in each batch is X, where X is the value obtained by dividing 100 by N, and N is T FIBThe number of test cases in it. This ensures that the number of patches corresponding to each test case is the same. If the total number of patches is less than 100, a test case is randomly selected and added to the prompt to generate patches to make up the difference. The experimental results are as follows: Without adding test cases, 117 faults were correctly repaired; with the addition of test cases, 119 faults were correctly repaired, showing a slight improvement compared to the former.
[0079] Patch sorting effect: The present invention uses automatically generated test cases for patch sorting. The patches used in this experiment are the patches generated by gpt-3.5-turbo in the patch generation experiment. First, all patches are replaced in the original code one by one, and all test cases in T FIB are executed, and the results are recorded. Subsequently, the patches are sorted from high to low according to the number of test cases passed by each patch; for patches with the same number of passed test cases, a sorting algorithm based on coding similarity and token entropy is used for sub-sorting. Among them, the sorting method of coding similarity uses the embedding model developed by OpenAI to encode the patch and the suspicious code snippet to obtain the encoding vector. Subsequently, the cosine similarity between the two is calculated. On the premise that the similarity is greater than 0.95, the qualified patches are sorted from low to high according to the similarity; the sorting method of token entropy is to record the probability value (probability) of each token during the patch generation stage, and then calculate the sum (Sumentropy) or average value (Mean entropy) of the probability values of all tokens in the patch as the basis for its sorting, and sort from high to low. The present invention compares the patch sorting methods based on coding similarity and token entropy, and judges the pros and cons of the sorting methods through the Top@N index; the meaning of Top@N is the number of faults containing the correct patch among the top N patches; since the two sorting methods used for comparison do not need to go through compilation and other links, and compilation and execution and other links can help judge the correctness of some patches. For fair comparison, both algorithms participating in the comparison add a compiler check link to exclude the influence of irrelevant variables.
[0080] Table 2 shows the results of the patch sorting experiment. It can be seen from it that the sorting method based on automatically generated test cases has a certain improvement compared to the existing two sorting methods; compared with the token entropy average value sorting method with the best effect in the Top@1 index, it has increased by about 6%, and in the Top@5 index, it has increased by about 27.8% compared with the best-performing coding similarity-based sorting method.
[0081] Table 2: Evaluation results of patch sorting in experimental evaluation
[0082] Sorting method Top@1 Top@3 Top@5 Token entropy (summation) 59 63 65 Token entropy (average value) 60 66 67 Coding similarity 49 67 72 Test case sorting 64 83 92
[0083] Overall experimental results: To evaluate the overall effectiveness of the present invention, a comparison was made with existing automatic program repair methods applicable to the same scenario, and the iFixR tool was selected as the comparison object. Both it and the present invention are applicable to triggerless testing and use only the fault report as the input information for automatic program repair. Since the latter conducts experiments on different versions of the Defects4J dataset, which is smaller than the dataset used in the present invention, only the intersection of the two datasets was considered when evaluating the results. The evaluation metric is still Top@N, that is, the number of faults for which the correct patch is included in the top N patches finally recommended. The overall effectiveness evaluation results are shown in Table 3. It can be seen from this that, compared with the existing methods, the present invention has a relatively significant improvement in overall effectiveness.
[0084] Table 3: Evaluation results of overall effectiveness in experimental evaluation
[0085] Sorting method Top@1 Top@3 Top@5 Top@all Benchmark method 11 17 21 23 This method 33 38 41 51
[0086] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method-level automatic program repair method for enhancing test case generation, characterized in that It includes the following steps: (1) Use an automatic test case generation tool, take the fault report as the input information, and automatically generate a set of test cases, denoted as sequence T; (2) Execute all test cases in execution sequence T and perform screening based on their results. Screen out the test cases that failed to execute and denote them as sequence T FIB ; (3)Based on the test case sequence T obtained in step (2) FIB , calculate the method-level suspiciousness scores of all code fragments by combining the spectrum-based fault localization method and the information retrieval-based fault localization method, and obtain the fault localization sequence L after sorting F = [(m1, s1), (m2, s2), …, (m i , s i ), …, (m n , s n )], where m i and s i are the code method name and the suspiciousness score corresponding to the i-th sorted location respectively, and n is the length of the fault localization sequence; (4)Based on the fault location sequence L obtained in step (3) F , the corresponding suspicious code segments are obtained; among them, the suspicious code segment corresponding to each location is specifically the code and annotation text from the method name to the end of the method body; (5) Combine the fault report, the suspicious code snippet obtained in step (4), a randomly selected test case text obtained in step (2), and the explanatory text into a prompt, and input it into the large language model. Use the large language model to generate a patch to obtain a patch sequence P = [p1, p2, …, p i , …, p n′ , where p i represents the i-th patch, and n′ is the length of the patch sequence; (6) For each patch in the patch sequence obtained in step (5), replace the corresponding original code text one by one, and execute all the test cases obtained in step (2). Sort the patches according to the execution results to obtain the final patch sequence; (7) Use the original test suite to verify the patches in the final patch sequence, and filter out the patches that fail to execute; Submit the first N patches in the filtered final patch sequence as the final repair solution.
2. The method for generating an enhanced method-level automatic program repair method for test cases according to claim 1, wherein The automatic test case generation tool is an automatic test case generation method based on a large language model. Its input information is the fault report. By using the large language model's ability to understand natural language and code text, it generates trigger tests and automatically generates multiple test cases at one time as a set of test cases.
3. The method for generating enhanced method-level automatic program repair for test cases according to claim 1, characterized in that, The specific steps of step (3) include the following sub-steps: (3.1) Use the fault localization method based on information retrieval to obtain the suspiciousness scores of all code snippets in the software repository, and sort all code snippets in descending order of the suspiciousness scores to obtain the sorted suspicious code sequence, denoted as where m i and are the code method name corresponding to the i-th ranked location and the suspiciousness score at the IRFL method level, respectively; (3.2) Use the fault localization method based on information retrieval to insert all the test case sequences T obtained in step (2) FIB into the original test suite, execute all the test cases and calculate the suspiciousness scores of all code fragments in the software repository, and sort all the code fragments in descending order of the suspiciousness scores to obtain the sorted suspicious code sequence, denoted as where m i and are the code method name corresponding to the i-th ranked location and the suspiciousness score at the SBFL method level, respectively; (3.3) Normalize the suspiciousness scores of each code method in the sorted suspicious code sequences obtained in steps (3.1) and (3.2) respectively according to the following formula to obtain the normalized suspiciousness scores: where e is the natural exponent, represents the suspiciousness score at the IRFL method level or the SBFL method level and the corresponding suspiciousness score after normalization; (3.4) Add the normalized suspiciousness scores at the IRFL method level and the SBFL method level obtained in step (3.3) to obtain the final suspiciousness score; (3.5) Sort all code methods in descending order according to the final suspiciousness scores obtained in step (3.4) to obtain the final fault localization sequence L 1 = [(m1, s1), (m2, s2), …, (m i , s i ), …, (m n , s n )], where m i and s i are the code method name and the suspiciousness score corresponding to the localization ranked i respectively, and n is the length of the fault localization sequence.
4. The method for generating enhanced method-level automatic program repair for test cases according to claim 1, characterized in that The specific steps of step (6) include the following sub-steps: (6.1) Replace each patch in the patch sequence obtained in step (5) with the corresponding original code text, and execute all the test cases in the sequence T obtained in step (2), record the successful execution cases, denoted as (p FIB , N i ), where i is the patch number, p i is the i-th patch text, and N i is the number of test cases passed by the i-th patch; i Patches that cause execution failure due to compilation failure or other abnormal reasons are directly removed; (6.2) Based on the execution status recorded in step (6.1), sort each patch in the patch sequence obtained in step (5) according to the number N of test cases passed by the patch i Sort them from largest to smallest, that is, the patch passing the most test cases has the highest priority, and the patch passing the fewest test cases has the lowest priority, and so on, to obtain the sorted patch sequence; (6.3) Based on the sorted patch sequence obtained in step (6.2), for the case where the number of passing test cases is the same, combine the sorting methods based on coding similarity and token entropy to perform a detailed sorting to obtain the finally sorted patch sequence; The priorities during sorting are, from high to low: the number of passing test cases, token entropy value, coding similarity.
Citation Information
Cited By
Automatic program repairing method combining executable invariant and differential signal
CN121255253A
An automatic program repair method combining executable invariants and differential signals
CN121255253B