A method and device for inspecting website backend after bug repair

Through PatchID technology, dynamic behavior expression snapshot and test case enhancement, overfitting patches are identified and classified, solving the problem that overfitting patches occupy a large number of in the existing technology, and improving the efficiency and accuracy of automatic program repair.

CN116257447BActive Publication Date: 2025-05-13XIAN ZHOUYUAN DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310210845.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-07
Publication Date
2025-05-13
Estimated Expiration
2043-03-07

AI Technical Summary

Technical Problem

In existing automatic program repair technologies, overfitting patches occupy a large number of them, resulting in developers requiring manual verification, consuming a lot of resources, and it is difficult for the existing technology to effectively identify and subdivide overfitting patches.

Method used

A PatchID technology is proposed, which creates a new test case by constructing the dynamic behavior expression snapshot, enhances the test set, and judges whether it is overfitted based on the changes in the snapshot value of the patch, and then classifies the overfit patch.

Benefits of technology

It effectively reduces the number of overfitting patches during the automatic software repair process, increases the speed of programmers to fix bugs, reduces the cost of software development, and is better than existing similar methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116257447B_ABST
    Figure CN116257447B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for inspecting a website backend after a bug is fixed. The maximum suspicious snapshot is found for the website backend code after the bug is fixed, and overfitting patches are identified and classified. The program dynamic behaviors of the passed test cases are the same in the bug program and the correct patch program, and the program dynamic behaviors of the failed test cases are different in the bug program and the correct patch program. A snapshot that causes a program error is constructed from the bug program and the test set, and the same snapshot is read from the patch program. Whether the patch is overfitting is determined based on whether the value of the snapshot changes with the use of the patch. The present invention reinterprets patch similarity from the perspective of program invariants and program expressions, and proposes a five-tuple representation method for calculating patch similarity, which is used for identifying and segmenting overfitting patches in automatic patch generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bug repair, and relates to a method and a device for inspecting a website backend after bug repair. The method is a method for identifying overfitting patches in patches generated by a software automatic repair tool and classifying the overfitting patches. The purpose is to reduce the number of overfitting patches in the software automatic repair process after using the technology, thereby improving the speed of programmers repairing bugs and reducing the cost of software development. Background Art

[0002] Automatic program repair (APR) has triggered extensive research in the past decade, and a large number of repair techniques have been proposed, among which test set-based repair techniques account for the majority. Test set-based repair tools use a given test set as an oracle. If the generated patch can pass the test set, the patch will be considered correct. However, in practice, the test set is weak and cannot fully express the expected functions of the program, resulting in patches that pass all test cases not being completely correct. There are patches that pass all test cases but are still wrong, that is, overfitting patches. This leads to a large number of invalid patches generated by APR technology. Current repair technologies are far from mature, and most repair technologies will simply accept patches that pass the test set. In order to filter out these patches, developers often need to manually verify the patches, which consumes too many resources. Due to the low performance of repair technologies, developers must manually verify a large number of wrong patches. Therefore, solving the overfitting of patches has become an issue that urgently needs to be studied and solved.

[0003] If a patch is considered to be a correct patch if it passes the test set, then the number of overfitted patches can be reduced by simply enhancing the test set. However, automatic test generation tools can only generate test inputs, and appropriate test outputs still need to be manually determined by humans. Even this approach still cannot express a complete oracle. Especially for large projects, it is very difficult to get a complete oracle. At present, being able to identify whether a patch is overfitted is already a great success. Because quickly identifying overfitted patches can improve the success rate of APR technology and developers in fixing bugs, but if there is technology that can subdivide overfitted patches, it can further increase the speed at which developers fix program errors.

[0004] Generally speaking, overfitting patches can be divided into the following three categories: (1) A-Overfitting Patch: the patch neither completely fixes the incorrect behavior nor destroys the original correct behavior; (2) B-Overfitting Patch: the patch fixes the original incorrect behavior but destroys the original correct behavior, which is called regression error; (3) AB-Overfitting Patch: the patch not only does not fix the incorrect behavior but also destroys the original correct behavior. At present, people have proposed different overfitting detection methods. One of the strategies is to mine the deep behavior of the test set and the program. Through the principle of patch similarity, the success rate of overfitting patch identification can reach 56%.

[0005] Explanation of terms

[0006] Program abstract state: The value abstracted from program behavior during the execution of the website backend code. Summary of the invention

[0007] The purpose of the present invention is to propose a method and device for inspecting the backend of a website after a bug is fixed in view of the deficiencies of the prior art. The present invention proposes a new technology - PatchID, the core idea of ​​which is that the dynamic behaviors of the passed test cases in the bug program and the correct patch program are the same, while the dynamic behaviors of the failed test cases in the bug program and the correct patch program are different. First, a snapshot of the dynamic behavior expression that causes the program error is constructed from the bug program and the test set, then a new test case is generated to enhance the original test set, and finally the same snapshot is read from the patch program, and whether the patch is overfitted is determined based on whether the snapshot value changes with the use of the patch.

[0008] In a first aspect, the present invention provides a method for inspecting a website backend after a bug is fixed, comprising the following steps:

[0009] Step 1: Find the most suspicious dynamic behavior expression snapshot in the website backend code after the bug is fixed;

[0010] Step 1-1: Get the dynamic behavior expression snapshot of each test case;

[0011] The backend website before running the bug fix, and the test set t corresponding to the backend website o First, construct the Boolean expression required by the snapshot and obtain the Boolean expression set B bug , collect the program abstract state during the execution of each test case; then calculate the Boolean expression set B bugThe value of each Boolean expression in the test set generates a snapshot of the dynamic behavior expression of each test case;

[0012] The dynamic behavior expression snapshot is expressed using a five-tuple based on the patch similarity principle:

[0013] snapshot= <l,b,?,i,v i > Formula (1)

[0014] Where l is the unique position identifier of each statement, b is a Boolean expression, ? is the value of b (true or false), i is the unique serial number of each test case in the test set, and v is the i Represents test case t i The actual value of b during the execution of the buggy program;

[0015] Step 1-2: Calculate the suspiciousness of each dynamic behavior expression snapshot. The calculation formula is defined as follows:

[0016]

[0017] Among them s Denotes dependent variables (syntactic analysis of expression dependence), dy s Indicates dynamic analysis variables (dynamic analysis);

[0018] Each snapshot has a corresponding suspiciousness, which is determined by the following two factors: 1) s ; 2)dy s . s As the number of times b appears in the statements before and after l increases; the more times b takes the value of ? in the failed test cases, the fewer times it takes the value of ? in the passed test cases, then dy s The larger the value of .

[0019] Step 1-3: Filter the dynamic behavior expression snapshot with the maximum suspiciousness, denoted as s max ; The dynamic behavior expression snapshot with the maximum suspiciousness is the dynamic behavior expression snapshot of the Bug, and then the test set t is obtained o The corresponding snapshot set s bug ;

[0020] Step 2: For the test set t o Perform data enhancement to obtain the test set t after data enhancement e

[0021] Use Evosuite software to randomly generate multiple new test cases, replace the test set in step 1-1 with the new test cases, and then repeat step 1-1 to calculate the dynamic behavior expression snapshot of these test cases, denoted as s new ; if s new With s max If s is the same, new The corresponding test case is added to the test set t o Otherwise, discard s new The corresponding test cases finally get the test set t after data enhancement e ;

[0022] Step 3: Identify overfitting patches and classify the identified overfitting patches; the details are as follows:

[0023] Step 3-1: Get the location you need to monitor in the patch used for the bug fix patch ;

[0024] Since the location of the bug before the website backend is fixed cannot be monitored directly in the patch, you need to reselect a location in the patch. patch To monitor the Boolean expression b that is the same as the dynamic behavior expression of the bug; no matter which repair operation is performed, the program can only have correct program behavior after the repair operation is completed, so the statement that defines the first difference between the bug and the patch is recorded as start s , the last different statement is recorded as end s , the following rules are used to select the listening position:

[0025] 1) If start s Use block statements and end s At start s Internally, then l patch The next statement after the end of a block statement; the block statement is for, while or if;

[0026] 2) If start s Do not use block statements to judge end s Is it the last statement? If not, l patch At the end s The next statement is, if then l patch =end s ;

[0027] Step 3-2: Run the backend website with bug fixes and the test set with data enhancement e , get the test set t e Each test case in lpatch The program abstract state on the t, and then get the dynamic behavior expression snapshot, and finally get the test set t e The corresponding snapshot set s patch ; where s patch With s bug The Boolean expression b is the same as ?;

[0028] Step 3-3: Put the two sets s bug 、s patch According to the test set t e Compare the serial numbers of the test cases in to obtain the set s bug 、s patch The number of test cases that have failed tests that are the same as v is N f , and the set s bug 、s patch The number of different v between the test cases that passed the test N p ;

[0029] According to the following formula (3), the type of patch is identified:

[0030]

[0031] Among them, correct represents the correct patch, A represents an overfitting patch of type A, that is, the patch neither completely fixes the incorrect behavior nor destroys the original correct behavior; B represents an overfitting patch of type B, that is, the patch fixes the original incorrect behavior but destroys the original correct behavior, which is called regression error; AB represents an overfitting patch of type AB, that is, the patch not only fails to fix the incorrect behavior but also destroys the original correct behavior.

[0032] In a second aspect, a testing device is provided, comprising:

[0033] The maximum suspicious snapshot search module is used to find the dynamic behavior expression snapshot with the maximum suspiciousness in the website backend code after the bug is fixed;

[0034] Test data enhancement module, used to enhance the test set t o Perform data augmentation;

[0035] Identification and classification module of overfitted patches.

[0036] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described.

[0037] According to a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described is implemented.

[0038] The beneficial results of the present invention are specifically:

[0039] 1. This paper reinterprets patch similarity from the perspective of program invariants and program expressions, and proposes a five-tuple representation method for calculating patch similarity, which is used for overfitting patch identification and segmentation in automatic patch generation.

[0040] 2. The present invention identifies 63 overfitting patches and 15 correct patches in the classic Java dataset Defects4j. Experimental data show that it is superior to existing similar methods. Since the technology can subdivide patches, developers can modify overfitting patches into correct patches more quickly. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 The figure is an overall flow chart of the method of the present invention. DETAILED DESCRIPTION

[0042] The present invention is described in detail below in combination with the software automatic repair technology according to the accompanying drawings. Figure 1 As shown, the specific steps are as follows:

[0043] Step 1: Find the most suspicious dynamic behavior expression snapshot in the website backend code after the bug is fixed;

[0044] Step 1-1: Get the dynamic behavior expression snapshot of each test case;

[0045] The backend website before running the bug fix, and the test set t corresponding to the backend website o First, construct the Boolean expression required by the snapshot and obtain the Boolean expression set B bug , collect the program abstract state during the execution of each test case; then calculate the Boolean expression set B bug The value of each Boolean expression in the test set generates a snapshot of the dynamic behavior expression of each test case;

[0046] A Boolean expression is a combination of variables of the same type using logical symbols (<, ≤, ≥, >, ≠, =, !).

[0047] The dynamic behavior expression snapshot is expressed using a five-tuple based on the patch similarity principle:

[0048] snapshot= <l,b,?,i,v i > Formula (1)

[0049] Where l is the unique position identifier of each statement, b is a Boolean expression, ? is the value of b (true or false), i is the unique serial number of each test case in the test set, and v is the i Represents test case t i The actual value of b during the execution of the buggy program;

[0050] After using the correct patch, the passed test cases are the same as the previous Boolean expressions and their values, while the failed tests should be different. For example, there is a buggy program, all the passed test cases make a Boolean expression b have the value false, and all the failed test cases make b have the value true. Then judging whether the patch is overfitting is not just a single way of observing the output of the program, but can be done by comparing the value of b in a certain statement before and after using the patch. When the passed test case tests the patch, the value of b should be consistent with the buggy program; when the failed test case tests the patch, the value of b should be different from the buggy program.

[0051] Step 1-2: Calculate the suspiciousness of each dynamic behavior expression snapshot. The calculation formula is defined as follows:

[0052]

[0053] Among them s Denotes dependent variables (syntactic analysis of expression dependence), dy s Indicates dynamic analysis variables (dynamic analysis);

[0054] Step 1-3: Filter the dynamic behavior expression snapshot with the maximum suspiciousness, denoted as s max ; The dynamic behavior expression snapshot with the maximum suspiciousness is the dynamic behavior expression snapshot of the Bug, and then the test set t is obtained o The corresponding snapshot set s bug ;

[0055] Step 2: For the test set t o Perform data enhancement to obtain the test set t after data enhancement e

[0056] Use Evosuite software to randomly generate multiple new test cases, replace the test set in step 1-1 with the new test cases, and then repeat step 1-1 to calculate the dynamic behavior expression snapshot of these test cases, denoted as s new ; if s new With s max If s is the same, new The corresponding test case is added to the test set t o Otherwise, discard s new The corresponding test cases finally get the test set t after data enhancement e ;

[0057] Step 3: Identify overfitting patches and classify the identified overfitting patches; the details are as follows:

[0058] Step 3-1: Get the location you need to monitor in the patch used for the bug fix patch ;

[0059] Since the location of the bug before the website backend is fixed cannot be monitored directly in the patch, you need to reselect a location in the patch. patch The Boolean expression b that monitors the dynamic behavior expression of the bug is the same as the expression b; for the bug program, the patch generally includes insert, delete, replace and update. No matter which repair operation is performed, the program can only have correct program behavior after the repair operation is completed, so the statement that defines the first difference between the bug and the patch is recorded as start s , the last different statement is recorded as end s , the following rules are used to select the listening position:

[0060] 1) If start s Use block statements and end s At start s Internally, then l patch The next statement after the end of a block statement; the block statement is for, while or if;

[0061] 2) If start s Do not use block statements to judge end s Is it the last statement? If not, l patch At the end s The next statement is, if then l patch =end s ;

[0062] Step 3-2: Run the backend website with bug fixes and the test set with data enhancement e , get the test set te Each test case in l patch The program abstract state on the t, and then get the dynamic behavior expression snapshot, and finally get the test set t e The corresponding snapshot set s patch ; where s patch With s bug The Boolean expression b is the same as ?;

[0063] Step 3-3: Put the two sets s bug 、s patch According to the test set t e Compare the serial numbers of the test cases in to obtain the set s bug 、s patch The number of test cases that have failed tests that are the same as v is N f , and the set s bug 、s patch The number of different v between the test cases that passed the test N p ;

[0064] According to the following formula (3), the type of patch is identified:

[0065]

[0066] Among them, correct represents the correct patch, A represents an overfitting patch of type A, that is, the patch neither completely fixes the incorrect behavior nor destroys the original correct behavior; B represents an overfitting patch of type B, that is, the patch fixes the original incorrect behavior but destroys the original correct behavior, which is called regression error; AB represents an overfitting patch of type AB, that is, the patch not only fails to fix the incorrect behavior but also destroys the original correct behavior.

[0067] The present invention is experimentally verified on two datasets, wherein the first dataset is Dfects4J, which is composed of patches generated by 6 APR tools on Defects4J. The second dataset is Java+JML dataset created by Nilizadeh et al.

[0068] Defects4J. Currently, Defecets4J proposed by Just is the most widely used Java program dataset in the field of automatic program repair. Defects4J has 17 projects so far, which contain 835 defects. Each program defect in this dataset contains at least one test case that can trigger it. This method uses the six most commonly used projects in this dataset, namely: Chart, Time, Math, Lang, Closure and Mockito, among which Chart is a project dedicated to displaying icons; Time is a project for date and time processing; Math is a project for scientific computing; Lang is a set of additional methods for operating JDK classes; Closure is an optimizing compiler for Javascript; Mockito is a simulation framework for unit testing. The number of bugs contained in each project is shown in Table 1 below.

[0069] Table 1: Defects4j Project

[0070] Project Name Number of bugs Chart 26 Time 26 Math 106 Lang 64 Closure 174 Mockito 38 Total 434

[0071] This method uses six existing repair tools to repair on the Defects4J dataset and obtain candidate patches. The six automatic program defect repair tools are jGenProg, Nopol 2015, Nopol 2017, ACS, HDRepair and jKali. Among them, jGenProg is the Java version of GenProg, which is a heuristic search repair tool based on genetic algorithm; Nopol is a repair technology for conditional statement errors in Java programs. This technology gives different repair strategies for the type of error statements: if the code location where the error is located is a conditional statement, the repair patch usually generated by Nopol is to modify the original conditional statement; if the code location where the error is located is a non-conditional statement, a new condition is added to skip the execution of the current statement to achieve repair. This dataset includes two versions, Nopol 2015 and Nopol 2017; ACS is a high-precision conditional statement synthesis tool, which extracts patch templates for repair based on statistical analysis; HDRepair is also a repair tool based on statistical analysis; jKali is a re-implementation of Kali on Java, which is a repair tool with only deletion function.

[0072] Java+JML dataset. This dataset proposed by Nilizadeh is the first verified and publicly available Java program dataset. It consists of the following four parts: correct programs, mutated error programs, test suites, and APR-based patches. The programs in this dataset have JML specifications for experimental evaluation. This dataset implements various classic algorithms and data structures, such as bubble sort, factorial, queue, etc. They are all small programs with formal specifications written in JML, so they can be considered as programs with oracles. The test suite is created using an AFL-based fuzzing tool, and the test suites are divided into Small and Medium according to the number of test cases generated. The error program is created by injecting a single error into each java program through PITest, a Java program mutation tool. PITest generates errors by changing control conditions, changing assignment expressions, removing method calls, and changing return values. The APR-based repair patches are obtained using the following repair tools, namely ARJAE, Cardumen, jGenProg, jKali, jMutRepair, Kali-A, and Nopol.

[0073] Experimental results:

[0074] Performance on Defects4J. A total of 220 patches were generated on the Defects4J dataset using the APR tool. This method conducted experiments on these 220 patches to determine whether they were overfitting patches. A total of 166 patches were found to be overfitting patches. The remaining patches were terminated due to exceeding the set execution time limit and no final results were given. Among these 166 patches, this method gave the determination results of whether they were overfitting patches for the remaining 157 patches except for 9 patches. The specific patch determination results are shown in Table 2.

[0075] Table 2: Defects4j Dataset

[0076]

[0077] Tables 3 and 4 show the running results of this method on related defect repair tools and different projects, respectively. As shown in the table, PatchID successfully filtered out 78 patches from 157 patches, including 63 overfitting patches and 15 correct patches. And for the 63 overfitting patches, PatchID successfully divided them into three categories, among which A-Overfitting Patch accounted for the largest proportion, reaching 50; followed by B-Overfitting Patch with 8; AB-Overfitting Patch number was 5.

[0078] Table 3: Results By APR Tools

[0079] Tool Correct Overfitting Correct detected Overfitting detected A B AB Nopol2015 5 20 2(40%) 10(50%) 9 0 1 Nopol2017 3 68 2(66.66%) 36(52.94%) 25 8 3 HDRepair 4 5 3(75%) 1(20%) 1 0 0 ACS 11 6 7(63.63%) 1(16.66%) 1 0 0 jKali 1 14 0 8(57.14%) 8 0 0 jGenprog 6 14 1(16.67%) 7(50%) 6 0 1 Total 30 127 15(50%) 63(49.61%) 50 8 5

[0080] "Correct / overfitting detected" indicates the number of patches correctly classified by the proposed method from the "Correct / overfitting" patches.

[0081] A=A-Overfitting Patch, B=B-Overfitting Patch, AB=AB-OverfittingPatch

[0082] Table 4: Result By Project

[0083] Project Correct Overfitting Correct detected Overfitting detected A B AB Lang 6 10 2(33.33%) 3(50%) 3 0 0 Math 16 49 8(50%) 22(44.90%) 20 1 1 Chart 3 21 1(33.33%) 12(57.14%) 10 0 2 Time 2 10 2(100%) 6(60%) 5 1 0 Closure 2 37 1(50%) 20(54.05%) 12 6 2 Mockito 1 0 1(100%) 0 0 0 0 Total 30 127 15(51.85%) 63(49.61%) 50 8 5

[0084] Overfitting patch. From Table 4, we can find that PatchID has better effects on the four repair tools Nopol2015, Nopol2017, jKali, and jGenprog (the worst success rate is 50%), but it has poorer effects on ACS and HDRepair (the best success rate is only 20%). We also found that among the overfitting patches generated by these six tools, the patches that did not fix the original errors of the program were the most, and the patches that destroyed the original correct behavior of the program were relatively few. However, Nopol2015 and Nopol2017 (these two tools are tools that modify program conditional statements to fix bugs) have a total of 12 patches that destroy the correct behavior of the program, and only jGenprog among the other tools generates an AB-Overfitting Patch. We guess that modifying program conditional statements is more likely to introduce new errors.

[0085] According to the Project, the success rate of overfitting patch identification is relatively stable, ranging from 43% to 60%. PatchID has the highest success rate in the Time Project, reaching 60%. The lowest success rate in the Math Project is only 44.90%. Here, the number of patches that destroy the original correct behavior of the program is the largest, with a total of 8.

[0086] Correct patch. Among the 157 patches, there are 30 correct patches, and PatchID can correctly judge 15 patches, with a success rate of 50%. This is exciting news. As far as we know, no other tool can achieve such a high success rate. Among the patches generated by Nopol2017, HDRepair and ACS, PatchID's success rate exceeds 60%, and the highest is 75%. From the perspective of Project, except Lang and Chart, the success rate of other Projects is not low. It is particularly noteworthy that the success rate on Mockito and Time projects is 100%.

[0087] Compared with Xiong's results on this dataset, we identified one more overfitting patch than Xiong's method, but Xiong's method did not identify any correct patches while PatchID identified 15. For 220 patches, his method identified 62 in total and PatchID identified 78 patches. However, Xiong increased the recognition success rate to 56.3% through the strategy of trimming the average value, while PatchID was 49.7%. In addition, Xiong's method can only identify patches for four projects: Chart, Lang, Math, and Time, while PatchID covers patches for all six projects. In terms of versatility, PatchID is more extensive.

[0088] Performance on Java+JML dataset. We selected 236 overfitting patches based on the Medium test suite and 336 overfitting patches based on the Small test suite from the Java+JML dataset. These overfitting patches were judged by the JML specification and determined to be overfitting patches. In addition, there are 21 FalseNegatives patches (JML specification mistakenly considers a correctrepairedprogram as overfitted). The PatchID algorithm was run on a total of 593 patches and the running results of 380 patches were obtained. The specific results are shown in Table 5.

[0089] Table 5: Java+JMLDataset

[0090] Patch Type Collected Validated Medium 236 144 Small 336 221 False Negatives 21 15 Total 593 380

[0091] From the data in Table 6, we can see that based on the Medium type patches, PatchID can correctly identify 72 overfitting patches with a success rate of 50%; based on the Small type patches, PatchID can correctly identify 92 overfitting patches with a success rate of 41.62%; and among the FalseNegatives patches, only 5 correct patches are identified with a success rate of only 33.33%.

[0092] From the perspective of overfitting classification, PatchID does not identify any B-Overfitting patches on this dataset. In addition, except for the four AB-Overfitting patches, all other overfitting patches are of A-Overfitting type.

[0093] It is obvious from the success rates of Medium and Small that as the number of test cases in the test suite decreases, the success rate also decreases. This data shows that weak test suites affect the success rate of PatchID.

[0094] Table 6: Result By PatchType

[0095] Patch Type Correct detected Overfitting detected A B AB Medium 72(50%) 72(50%) 72 0 0 Small 129(58.37%) 92(41.62%) 88 0 4 False Negatives 5(33.33%) 10(66.67%) 10 0 0 Total 206 174 170 0 4

Claims

1. A method for inspecting the backend of a website after a bug is fixed, characterized in that The method comprises the following steps: Step 1: Find the most suspicious dynamic behavior expression snapshot in the backend code of the website after the bug is fixed; Step 1-1: Get the dynamic behavior expression snapshot of each test case; The backend website before running the bug fix, and the test set t corresponding to the backend website o First, construct the Boolean expression required by the snapshot and obtain the Boolean expression set B bug , collect the program abstract state during the execution of each test case; then calculate the Boolean expression set B bug The value of each Boolean expression in the test set generates a snapshot of the dynamic behavior expression of each test case; The dynamic behavior expression snapshot is expressed using a five-tuple based on the patch similarity principle: in represents the unique position identifier of each statement, b represents a Boolean expression, ? represents the value of b, i represents the unique serial number of each test case in the test set, and v i Represents test case t i The actual value of b during the execution of the buggy program; Step 1-2: Calculate the suspiciousness of each dynamic behavior expression snapshot. The calculation formula is defined as follows: Among them s Denotes the dependent variable, dy s Represents dynamic analysis variables; Step 1-3: Filter the dynamic behavior expression snapshot with the maximum suspiciousness, denoted as s max ; The dynamic behavior expression snapshot with the maximum suspiciousness is the dynamic behavior expression snapshot of the Bug, and then the test set t is obtained o The corresponding snapshot set s bug ; Step 2: For the test set t o Perform data enhancement to obtain the test set t after data enhancement e : Use Evosuite software to randomly generate multiple new test cases, replace the test set in step 1-1 with the new test cases, and then repeat step 1-1 to calculate the dynamic behavior expression snapshot of these test cases, denoted as s new ; if s new With s max If s is the same, new The corresponding test cases are added to the test set t o Otherwise, discard s new The corresponding test cases finally get the test set t after data enhancement e ; Step 3: Identify overfitting patches and classify the identified overfitting patches; the details are as follows: Step 3-1: Get the location to be monitored in the patch used for the bug fix Since the location of the bug before the website backend is fixed cannot be monitored directly in the patch, you need to reselect a location in the patch To monitor the Boolean expression b that is the same as the dynamic behavior expression of the bug; no matter which repair operation is performed, the program can only have correct program behavior after the repair operation is completed, so the statement that defines the first difference between the bug and the patch is recorded as start s , the last different statement is recorded as end s , the following rules are used to select the listening position: 1) If start s Use block statements and end s At start s Internally, then The next statement after the end of a block statement; the block statement is for, while or if; 2) If start s Do not use block statements to judge end s Is it the last statement? If not, At the end s The next statement is, if Step 3-2: Run the backend website with bug fixes and the test set with data enhancement e , get the test set t e Each test case in The program abstract state on the t, and then get the dynamic behavior expression snapshot, and finally get the test set t e The corresponding snapshot set s patch ; where s patch With s bug The Boolean expression b is the same as ?; Step 3-3: Put the two sets s bug 、s patch According to the test set t e Compare the serial numbers of the test cases in to obtain the set s bug 、s patch The number of test cases that have failed tests that are the same as v is N f , and the set s bug 、s patch The number of different v between the test cases that passed the test N p ; According to the following formula (3), the type of patch is identified: Where correct represents the correct patch, A represents an overfitting patch of type A, i.e., the patch neither completely fixes the incorrect behavior nor destroys the original correct behavior; B represents an overfitting patch of type B, i.e., the patch fixes the original incorrect behavior but destroys the original correct behavior, which is called regression error; AB represents an overfitting patch of type AB, i.e., the patch not only fails to fix the incorrect behavior but also destroys the original correct behavior.

2. A testing device for implementing the method of claim 1, characterized in that include: The maximum suspicious snapshot search module is used to find the dynamic behavior expression snapshot with the maximum suspicious degree in the website backend code after the bug is fixed; Test data enhancement module, used to enhance the test set t o Perform data augmentation; Identification and classification module of overfitted patches.

3. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method of claim 1.

4. A computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method of claim 1 is implemented.

Citation Information

Patent Citations

  • Systems and applications of lighter-than-air (LTA) platforms

    CN101415602A

  • Fault bypassing method and device based on quintuple hash path

    CN113300873A