Test case evolution generation method based on multi-coverage index guidance
By building an abstract syntax tree and control flow diagram, combining depth-first search and multi-objective optimization algorithms, high coverage test cases are generated, which solves the problem of inefficiency of existing methods in complex program structures, and realizes efficient and automated test case generation.
Patent Information
- Application Number
- CN202510348076.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The existing test case generation methods are inefficient and have insufficient coverage when dealing with complex program structures, making it difficult to adapt to the diverse testing needs of modern complex software systems.
Based on the multi-coverage index-guided test cases evolution generation method, by constructing an abstract syntax tree and control flow diagram, combining depth-first search and multi-objective optimization algorithm NSGA-II, we automatically generate high-coverage test cases, comprehensively considering statement coverage, branch coverage and output coverage.
It significantly improves the efficiency and quality of test case generation, can quickly generate test cases that meet different coverage requirements, reduces test costs, and improves the automation level and effectiveness of software testing.
Smart Images

Figure CN120276987A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of software testing, and particularly relates to a method for evolving test case generation guided by multiple coverage metrics, which can be used to automatically generate and optimize software test cases, and improve test coverage and test efficiency. Background Art
[0002] In the software development process, software testing, as a key link to ensure software quality, is of great importance. Among them, the generation of test cases is the core task in software testing, and its main goal is to verify the correctness and integrity of software functions through systematic input data design. Traditional test case generation methods mainly rely on manual design or random generation strategies. These methods are not only inefficient, but also have significant problems such as insufficient coverage and low defect discovery rate, making it difficult to meet the increasingly complex test requirements of modern software systems.
[0003] In recent years, with the rapid development of automated testing technology, the test case generation method based on path coverage has gradually become a research hotspot. The path coverage method designs test cases to cover all possible execution paths in the program control flow graph (Control Flow Graph, CFG), and theoretically can achieve comprehensive test coverage. However, this method faces severe challenges when dealing with complex program structures: there are often a large number of paths in the program control flow graph, including many unreachable paths and redundant paths, which not only increases the complexity of test case generation, but also significantly increases the test cost.
[0004] To overcome the limitations of traditional test case generation methods, researchers have proposed an innovative test case generation method based on the branch evolution theory. This method accurately identifies key branch nodes through in-depth analysis of the program control flow graph (CFG), and constructs test cases using static analysis techniques for branch paths. However, the existing methods still have obvious deficiencies in practical applications, mainly manifested in the overly single coverage target, making it difficult to adapt to the diverse test requirements of modern complex software systems. Summary of the Invention
[0005] To solve the above problems, starting from multiple key coverage metrics, the present invention is committed to constructing a more comprehensive and efficient test case generation mechanism, and proposes an innovative test case evolution method, aiming to significantly improve test efficiency and quality while ensuring test coverage. The method of the present invention takes multi-dimensional coverage metrics as the core guidance, aiming to significantly improve the generation efficiency and coverage of test cases, and effectively reduce the test cost. This method can better adapt to the diverse requirements of the actual test environment, and ensure the accuracy and comprehensiveness of test activities.
[0006] The technical solution of the present invention:
[0007] A method for evolutionary generation of test cases guided by multiple coverage metrics is as follows:
[0008] Step (1) For the program Code under test, perform syntax parsing on it to generate its corresponding AST abstract syntax tree. Construct its corresponding CFG control flow graph by analyzing the nodes of the AST abstract syntax tree, which is a graphical representation method for describing the program execution flow, and describe the basic blocks and their execution paths through nodes and edges. And explore the target paths based on DFS (depth - first search), and automatically filter out the target path set Path according to the Z - path coverage criterion. target . By performing branch sorting on the target path set Path target to obtain the sorted target path set Path order ; Identify the branch nodes set Branches, statement nodes Statements, and output type Outputs information of the AST abstract syntax tree, and construct the instrumentation function set F based on rules insert , and combine it with the program Code under test to generate the instrumented program Code insert ; Feed the obtained instrumented program Code insert and the sorted target path set Path order into Step (3);
[0009] Step (2) Analyze the program Code under test, and use the random strategy to complete the generation of random test cases to obtain the test case set T list , and obtain the execution path set Path by executing the test case set T list and filtering the executable test cases; Encode the corresponding test cases of the execution path set Path exec through a binary conversion program to obtain the initial population Pop exec , and feed the obtained initial population Pop origin and the execution path set Path origin into Step (3); exec
[0010] Step (3) Set the maximum number of iterations Max; The sorted target path set Path order , the instrumented program Code insert , the initial population Pop origin and the execution path set Path exec As input, it is input into the optimization algorithm program OP and optimized by the multi-objective optimization algorithm NSGA-II; the result of the instrumentation function generated by the optimization algorithm program OP is obtained, and it enters the coverage analysis program CA. By collecting the result of the instrumentation function, the statement coverage SC, branch coverage BC, and output coverage OC information are obtained for comprehensive analysis to determine whether the analysis result passes the verification of the set coverage rules. If it passes the verification, the current population Pop current The added result test case set T result ; Otherwise, it is judged whether the iteration times of the current population reach the maximum iteration times Max. If so, the iteration is terminated, and the current population Pop current is added to the result test case set T result ; Otherwise, the current population continues to be iteratively optimized.
[0011] Furthermore, step (1) specifically includes the following steps:
[0012] 1-1) For the program under test Code, perform syntax parsing on it to generate its corresponding AST abstract syntax tree; perform node recognition on the AST abstract syntax tree, recognize the output node and the method return value type information to obtain the possible types of output, and construct the output types Outputs; construct output instrumentation functions according to the output types Outputs to generate the output function set F outputs Proceed to step 1-4); if the method return value type of the program under test Code is Boolean, the keys of the dictionary data of the output function are false and true; otherwise, the keys of the dictionary data corresponding to the output function are the specific statement nodes. The output instrumentation function is to insert the statement execution record function CountOutput before each return statement;
[0013] 1-2) Perform node recognition on the AST abstract syntax tree, recognize the branch nodes, construct the branch node set Branches, and generate the branch function set F by constructing instrumentation functions based on the branch predicate rules insert Proceed to step 1-4); the branch predicate rules are as follows:
[0014] Branch weight: w = k * n (k = 0.1, n is the number of program branch nodes);
[0015] Branch predicate (judgment expression): Boolean; branch function expression: 1;
[0016] Branch predicate (judgment expression): a <= b; branch function expression: w * (a - b) + 1;
[0017] Branch predicate (judgment expression): a >= b; branch function expression: w * (b - a) + 1;
[0018] Branch predicate (judgment expression): a == b; Branch function expression: w * abs(a - b) + 1;
[0019] Branch predicate (judgment expression): a != b; Branch function expression: 1;
[0020] Branch predicate (judgment connection predicate): Branch function expression X and Branch function expression Y; Branch function expression: max(Branch function expression X, Branch function expression Y);
[0021] Branch predicate (judgment connection predicate): Branch function expression X or Branch function expression Y; Branch function expression: min(Branch function expression X, Branch function expression Y);
[0022] 1 - 3) Identify nodes in the AST (Abstract Syntax Tree), identify statement nodes, construct a statement node set Statements, and generate a statement function set F statements Proceed to step 1 - 4); The statement instrumentation function is to insert a statement execution record function CountState before each statement;
[0023] 1 - 4) The branch function set F branches 、output function set F outputs and statement function set F statements are integrated to form an instrumentation function set F insert and combined with the program under test Code to generate an instrumented program Code insert ;
[0024] 1 - 5) Analyze the nodes of the abstract syntax tree AST, extract the edge relationships between its nodes, generate an initial edge set Edge origin , and obtain its corresponding CFG (Control Flow Graph); Use DFS (Depth - First Search) to construct paths, where: During the path construction process, according to the Z - path coverage criterion, loop nodes are traversed only once; Identify relevant edges through depth - first traversal DFS to form complete paths; Finally, generate a target path set Path target ;
[0025] 1 - 6) Identify the branch nodes in the target path set Path target , obtain the branch complexity of the paths through a branch complexity analysis program, and sort the paths according to the branch complexity to obtain a sorted target path set Path order .
[0026] Furthermore, step (3) specifically includes the following steps:
[0027] 3 - 1) Set the maximum number of iterations Max for the current population;
[0028] 3-2) Integrate the sorted target path set Path order , the instrumented program Code insert , the initial population Pop origin and the execution path set Path exec information into step 3-3);
[0029] 3-3) Execute the test case set for population construction through the optimization algorithm program OP, and obtain the instrumented function results of the population to enter step 3-4);
[0030] 3-4) Through the coverage analysis program CA, extract the CountState function, CountOutput function, and branch distance information of individuals and the population as a whole in the current population Pop current , and then calculate the branch coverage BC, output coverage OC, and statement coverage SC of individuals and the whole, and enter step 3-5);
[0031] 3-5) Through comprehensive analysis, calculate the Pareto front of the current population Pop current and obtain the overall coverage information Total current of the current population Pop coverage . Verify the overall coverage information Total current and the Pareto front of the current population Pop coverage . If the overall coverage information Total coverage passes the verification by the set coverage rules or the Pareto front passes the verification by the set coverage rules, add the current population Pop current to the result test case set T result and enter step 3-7); Otherwise, enter step 3-6); Among them, the coverage rules for the overall coverage information Total coverage are: for the overall branch coverage BC, output coverage OC, and statement coverage SC, calculate the weighted sum according to the preset parameters, and judge whether it reaches the set threshold. If it reaches, the verification passes; otherwise, the verification fails; The coverage rules for the Pareto front are: if the branch coverage BC, output coverage OC, and statement coverage SC of individuals in the population are all greater than or equal to the preset threshold, the verification passes; otherwise, the verification fails;
[0032] 3-6) Judge the iteration times of the current population Pop current . If the iteration times of the current population Pop current reach the maximum iteration times Max, add the current population Pop current to the result test case set T result, terminate the iteration. Otherwise, optimize the current population Pop through the population crossover, selection, and mutation procedures current to generate a new population Pop for the next generation new and proceed to step 3-3);
[0033] 3-7) Finally, output the result test case set T result .
[0034] Compared with the prior art, the present invention has the following advantages and effects:
[0035] The present invention provides an efficient method for generating software test cases. It deeply analyzes the syntax structure of the program under test through static analysis technology, constructs an abstract syntax tree and a control flow graph, and then obtains a target path set and key node information, providing accurate data support for the generation of test cases. At the same time, the present invention comprehensively considers multiple key indicators such as statement coverage, branch coverage, and output coverage rate. By constructing specific instrumentation functions and adopting the advanced multi-objective optimization algorithm NSGA-II, it automatically generates test cases with high coverage rates. This method not only has a high degree of automation, reducing manual intervention, but also has pertinence, comprehensiveness, and practicality. It can quickly generate test cases that meet different coverage rate requirements, effectively improving the efficiency and quality of software testing. In addition, the scalability of the present invention enables it to adapt to test requirements in different scenarios, having a wide range of application prospects. Compared with the prior art, it significantly improves the automation level and test effect in the field of software testing. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a schematic flowchart of the test case evolutionary generation method based on multi-coverage metrics guidance of the present invention.
[0037] Figure 2 is a sub-flowchart of the static analysis-based target path recognition and instrumentation of the test case evolutionary generation method based on multi-coverage metrics guidance of the present invention.
[0038] Figure 3 is a sub-flowchart of the test case evolution based on multi-coverage metrics of the test case evolutionary generation method based on multi-coverage metrics guidance of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The method of the present invention will be described in detail below in conjunction with the drawings, technical solutions, and embodiments.
[0040] As Figure 1 shown, the test case evolutionary generation method based on multi-coverage metrics guidance of the present invention proceeds as follows: First, analyze the program under test Code to generate its corresponding AST abstract syntax tree and generate a test case set T using a random strategy listFor the obtained AST (Abstract Syntax Tree), its CFG (Control Flow Graph) is constructed by identifying nodes. With the association of nodes and edges, DFS (Depth-First Search) is used to extract directory paths, and according to the Z-path coverage criterion, the target path set Path is constructed. target The target path set is sorted according to the branch complexity to obtain the sorted target path set Path. order Meanwhile, the nodes in the AST are identified to obtain three key coverage metrics, namely the branch node set Branches, the statement node set Statements, and the output type Outputs information, and the instrumentation function set F is constructed based on these three key coverage metrics. insert It is combined with the program under test Code to generate the instrumented program Code. insert Then, for the test case set T list It is encoded into the initial population Pop through a binary conversion program. origin Meanwhile, the test cases in the test case set T are executed. list The executable test cases are selected as the execution path set Path. exec Then, the sorted target path set Path order , the instrumented program Code insert , the initial population Pop origin and the execution path Path exec are input into the optimization algorithm program OP together. The instrumentation function results are obtained through the multi-objective optimization algorithm NSGA-II, and the results are input into the coverage analysis program CA for analysis to obtain the results of three key coverages, namely statement coverage SC, branch coverage BC, and output coverage OC. After comprehensive analysis, it is judged whether the current population Pop current passes the coverage verification. If it passes, the current population Pop current information is added to the result test case set T result , otherwise it is judged whether the iteration times of the current population Pop current reach the maximum iteration times Max. If so, the current population Pop current is added to the result test case set T result , otherwise the current population Pop current is input into the optimization algorithm program OP for further optimization. Finally, the result test case set T result is output.
[0041] Taking a Python program as an example below, the implementation details of each process are described in detail. The specific implementation is as follows:
[0042] (1) As Figure 2As shown in the figure, it is the target path recognition and instrumentation based on static analysis. First, the syntax of the program to be tested Code is parsed to generate its corresponding AST (Abstract Syntax Tree). Node analysis is performed based on the AST, and a CFG (Control Flow Graph) is constructed. The target path is obtained according to DFS (Depth-First Search), and combined with the Z path coverage criterion, a target path set Path is generated. target Furthermore, branch nodes are identified. The program calculates and sorts the branch distances of the paths through branch complexity analysis, and obtains the sorted target path set Path. order Identify branch nodes, statement nodes, and output nodes in the AST. Respectively establish a branch node set Branches, a statement node set Statements, and an output type set Outputs. For the output type set Outputs, use the CountOutput(line number) function to generate output instrumentation functions, forming an output function set F. outputs For the branch node set Branches, create instrumentation functions based on branch predicate rules to obtain a branch function set F. branches For the statement node set Statements, construct code line instrumentation functions through the CountState(line number) function to form a statement function set F. statements Finally, integrate these three function sets to form a complete instrumentation function set F. insert And combine it with the program to be tested Code to generate the instrumented program Code. insert .
[0043] 1.1 AST Node Recognition and Instrumentation Function Construction
[0044] a) For the program to be tested, take the program that determines whether three input values are increasing or decreasing as an example. The program to be tested Code:
[0045]
[0046]
[0047] b) Generate the AST corresponding to the program to be tested Code:
[0048]
[0049] c) Generate the CFG (Control Flow Graph) of the program Code under test. By identifying the nodes and edges of the program Code under test, construct its string representation. Characterize the branch logic and execution path of the program through nested parentheses and the symbols "i" (representing the "if" conditional statement) and "r" (representing the "return" return statement): two branches follow "i" (representing the paths when the condition is true and false respectively), "r" represents the termination point of the program, and the nested parentheses reflect the hierarchical and logical structure of the program, thus clearly describing the control flow and branch relationship of the program. The string form of the generated CFG is as follows:
[0050] “
[0051] (i(1,(r2),(i(3,(r4),(r5)))))
[0052] ”
[0053] d) Obtain the set of statement nodes Statements of this AST (Abstract Syntax Tree) as:
[0054]
[0055] Instrument the code lines. Among them, the role of the statement instrumentation function is to insert the CountState(line number) function after the statement. It is stipulated that for the return statement, instrumentation is performed before the statement. This function is used to count the number of times each line of code is executed. During the execution of the test case, when passing through a certain path, the instrumentation function of the corresponding line number's branch path will be called. By counting and recording all the statement instrumentation functions called during the execution process, all the executed code lines can be obtained, thereby analyzing the statement coverage of the program. The statement instrumentation function is as follows:
[0056]
[0057] The set of statement functions F statements is:
[0058]
[0059] e) Identify the set of branch nodes Branches of this AST (Abstract Syntax Tree) as:
[0060]
[0061]
[0062] Build the instrumentation function for branch nodes based on the branch predicate rule:
[0063]
[0064] f) Identify the incoming edge nodes of the AST abstract syntax tree to obtain the initial edge set Edge origin It is:
[0065] “
[0066] [('Start','1'),('1','r2'),('1','3'),('3','r4'),('3','r5'),('r2','End'),('r4','End'),('r5','End')]
[0067] ”
[0068] Among them, the node prefix "r" is used to identify that the node is a return statement. After accurately extracting the target path through the DFS (Depth-First Search) strategy and operating in combination with the Z-path coverage criterion, the target path set Path is successfully constructed target .:
[0069]
[0070] For the branch complexity of the path, it is necessary to pre-compute the node complexity in its path. The node complexity formula is: NC = (n1 / N1 × N2 / n2), where n1 represents the number of different operators in the program node, n2 represents the number of different operands in the program node. N1 represents the total number of operators appearing in the program node, and N2 represents the total number of operands appearing in the program node. For example, in the node "1", the operands are "num1", "num2", and "num3", and the operator is "<". The branch complexity formula is: Among them, if the path path_j is single-branched, then Otherwise BC′(path_j) = max α {θ∣α = 1}B_j(path). Among them, for any path path_j, the variable α (1 ≤ α < n, n is the total number of operators in the branch node) represents the α-th operator of its branch node; the variable θ (1 ≤ i < n) represents the number of operators in the branch node. B_j(path_j) is the calculated value of the branch node operator of the path path_j, and its value calculates the branch value BC′(path_j) according to the conditional judgment category of the branch node, and finally obtains the branch complexity BC(path_j) of the path path_j. Calculate the branch complexity through the branch node set Branches to obtain the sorted target path set Path order :
[0071]
[0072]
[0073] For each piece of data, the value of the "branch_complex" field is the branch complexity of the corresponding path, and the value of the "sort_node_complex" field is the node complexity of the corresponding path.
[0074] g) The output type set Outputs of the AST abstract syntax tree determines the return value type of the function Code under test. It is found that the program does not specify a return type. Therefore, node recognition is performed on the return statements in the program under test, and the output type set Outputs is as follows:
[0075]
[0076] The output function instrumentation counts the execution times by adding the CountOutput(line number) function after each output statement. During test execution, the line number instrumentation function for a specific path is triggered to record the execution situation for analyzing the output coverage:
[0077]
[0078] Output function set F outputs is as follows:
[0079]
[0080] h) Obtain the instrumented program Code insert . By integrating the output function set F outputs , the branch function set F branches and the statement function set F statements to form the instrumentation function set F insert , and integrating F insert with the program Code under test to obtain the instrumented program Code insert :
[0081]
[0082]
[0083] (2) As Figure 1 shown, analyze the program Code under test, use the random strategy to generate random test cases to obtain the test case set T list , obtain the execution path set Path list by executing the test case set T exec and filtering the executable test cases; convert the test cases corresponding to the execution path set Path exec through the binary conversion program, and obtain the initial population P origin by encoding. Finally, the obtained initial population P originWith the execution path set Path exec 。
[0084] 2.1. Random test case generation and encoding
[0085] a) Use the random number filling strategy to generate an initial set of test cases T list :
[0086] “
[0087] [(1, 2, 3), (4, 5, 6), (1, 0, 1), (9, 8, 6), (0, 0, 0), (-1, -3, 0)]
[0088] ”
[0089] By executing the test case set T list , the executable path Path is obtained exec as follows:
[0090]
[0091]
[0092] b) Obtain the binary population. By inputting the test cases corresponding to the obtained execution path set Path exec into the binary conversion program, encoding them into binary format, using the 32-bit IEEE 754 standard binary format, that is, encoding each input parameter as a 32-bit binary value, to obtain the initial population Pop origin :
[0093]
[0094] (3) As Figure 3 shown, by integrating the sorted target path set Path order , the instrumented program Code insert , the initial population Pop origin Execution path set Path exec Execute the test case set constructed by the population through the optimization algorithm program OP to obtain the instrumented function results of the population. Specifically, analyze the CountState function information, CountOutput function information, and branch distance through the coverage analysis program CA to obtain the branch coverage BC, output coverage OC, and statement coverage BC information. Through comprehensive analysis, the three types of coverage rate results are weighted and calculated in turn according to the ratio of the preset parameters. Obtain the overall coverage rate information Total current of the current population Pop coverage , verify this information according to the set coverage rate rules. If the verification is passed, then the current population Popcurrent Add the result test case set T result ; Otherwise, judge the current population Pop current to see if the number of iterations has reached the maximum number of iterations Max. If so, add the current population Pop current to the result test case set T result ; Otherwise, optimize the current population Pop through the population crossover, selection, and mutation procedures current to generate the next-generation population Pop new Continue the population iteration optimization; finally, output the result test case set T result .
[0095] 3.1. Calculation of individual and population coverage
[0096] a) Analyze the information of the CountState function. For the statement coverage of the population, calculate the statement coverage result of the population by obtaining the statement coverage dictionary data maintained by the CountState function. The statement coverage SC of the individuals in the current population Pop current is as follows:
[0097]
[0098]
[0099] The overall statement coverage of the current population:
[0100]
[0101] b) Analyze the information of the CountOutput function. For the output coverage of the population, calculate the output coverage result of the individuals by obtaining the output coverage dictionary data maintained by the CountOutput function. The output coverage OC of the current population Pop current is as follows:
[0102]
[0103] The overall current output coverage of the population:
[0104]
[0105] c) Analyze the branch distance. For the current population Pop currentThe branch coverage BC of an individual is obtained by dividing the number of target paths covered by the current individual by the total number of target paths. By matching the execution paths of the individuals in the population with the target paths, calculating the degree of closeness of each execution path to each target path, the branch fitness function of the individual is obtained as F(x_j) = [approach_level(x_j, path_target) + 1.001^branch_distance(x_j, path_target)] * banch_complex(x_j, path_j), where banch_complex(x_j, path_j) is the result of normalizing the branch complexity: banch_complex(x_j, path_j) = log(1 + branch_complex(x_j, path_j)), that is, the branch complexity of the current execution path; approach_level(x_j, path_j) is the layer proximity, which represents the number of nodes of the execution path that are the same as the nodes of the target path divided by the number of nodes of the target path; branch_distance(x_j, path_target) is the branch distance, which is the result of the branch instrumentation function. Finally, the branch fitness result of the individual is f(x_j) = max(F(x_j)). If the population needs to be iteratively optimized, the fitness function result of the individual is used as one of the bases for individual selection. The branch coverage BC of the individuals in the current population is as follows:
[0106]
[0107] The branch coverage rate of the population is the number of currently covered branch paths divided by the total number of target paths. By supplementing the branch coverage BC results, the overall coverage rate of the current population:
[0108]
[0109] d) Calculate the Pareto front. Given the current population Pop current None of the solutions can dominate each other (that is, no solution is better than another in all objectives), so all individuals in the current population are regarded as equivalent excellent solutions and are completely retained as the result of the Pareto front. Therefore, the Pareto front of the current population Pop current is as follows:
[0110]
[0111] 3.2. Coverage verification
[0112] a) Determine whether the coverage rate of the current population Pop current passes the verification of the set coverage rate rule. For the overall coverage rate information Totalcoverage , by performing a weighted sum calculation on the overall branch coverage BC, output coverage OC, and statement coverage SC in the current population Pop current according to the preset parameters, that is, setting the parameters as 0.7:0.2:0.1, the result is 0.7×1.0 + 0.2×1.0 + 0.1×1.0 = 1.0. Since the result is greater than the set threshold of 0.86, the coverage rate of the current population Pop current is verified by the overall coverage rate information Total coverage set coverage rate rules. For the Pareto front, since the branch coverage BC, output coverage OC, and statement coverage SC metrics of the individuals are all greater than or equal to the preset threshold, that is, setting the threshold as 0.33, the coverage rate of the current population Pop current is verified by the coverage rate rules set by the Pareto front.
[0113] b) Output result test case set T result . Since the coverage rate of the current population Pop current is verified by the set coverage rate rules, the current population Pop current is added to the result test case set T result , and the output result test case set T result :
[0114]
[0115] c) If the current population Pop current fails to meet the preset coverage rate standard, it needs to be optimized. The optimization process includes the following steps: First, adopt the single-point crossover strategy, randomly select a crossover point in the parental binary individuals, and exchange the gene information after this point to generate new offspring individuals. Second, perform individual screening through the roulette wheel selection mechanism, which selects based on the proportion of individual fitness, ensuring that individuals with higher fitness have a greater chance of entering the next generation. Finally, introduce a mutation program, set the mutation probability of the binary individuals as 0.2, and randomly change some bits in the individual gene encoding to enhance the population diversity. This mutation process may involve random modification or exchange of genes. After these steps, a new generation population Pop new is obtained, and the iterative optimization continues. During this process, single-point crossover ensures the effective recombination of gene information, roulette wheel selection guarantees the inheritance of excellent individuals, and mutation operations increase the diversity of population evolution.
Claims
1. A method for evolutionary generation of test cases guided by multiple coverage metrics, characterized in that The specific steps are as follows: Step (1): For the program Code under test, perform syntax parsing on it to generate its corresponding AST abstract syntax tree. By analyzing the nodes of the AST abstract syntax tree, construct its corresponding CFG control flow graph, which is a graphical representation method for describing the program execution flow, and describe basic blocks and their execution paths through nodes and edges; and explore the target path based on depth-first search (DFS), and automatically screen out the target path set Path according to the Z-path coverage criterion. target ; By performing branch sorting on the target path set Path target , obtain the sorted target path set Path order ; Identify the branch node set Branches, statement node set Statements, and output type Outputs information of the AST abstract syntax tree, and construct the instrumentation function set F based on rules insert , and combine it with the program Code under test to generate the instrumented program Code insert ; Feed the obtained instrumented program Code insert and the sorted target path set Path order into Step (3); Step (2) analyzes the program Code under test and uses a random strategy to generate random test cases to obtain a test case set T list , and obtains an execution path set Path by executing the test case set T list and filtering executable test cases exec ; The execution path set Path is encoded by a binary conversion program exec to obtain the initial population Pop corresponding to the test cases origin , and the obtained initial population Pop origin and the execution path set Path exec Proceed to step (3); Step (3) sets the maximum number of iterations Max; the sorted target path set Path order , the instrumented program Code insert , the initial population Pop origin and the execution path set Path exec are used as inputs and input into the optimization algorithm program OP for optimization through the multi-objective optimization algorithm NSGA-II; obtain the instrumented function results generated by the optimization algorithm program OP, enter the coverage analysis program CA, collect the instrumented function results to obtain statement coverage SC, branch coverage BC, and output coverage OC information for comprehensive analysis, and determine whether the analysis results pass the verification of the set coverage rules. If the verification is passed, the current population Pop current is added to the result test case set T result ; otherwise, determine whether the iteration times of the current population reach the maximum number of iterations Max. If so, terminate the iteration and add the current population Pop current to the result test case set T result ; Otherwise, continue to iteratively optimize the current population.
2. The test case evolutionary generation method guided by multiple coverage metrics according to claim 1, wherein Step (1) specifically includes the following steps: 1-1) For the program Code under test, perform syntax parsing on it to generate its corresponding AST (Abstract Syntax Tree); perform node recognition on the AST to recognize the output nodes and method return value type information, so as to obtain the possible types of outputs, and construct the output types Outputs; construct output instrumentation functions according to the output types Outputs to generate the output function set F outputs Proceed to step 1-4); if the method return value type of the program Code under test is Boolean, the keys of the dictionary data of the output function are false and true; otherwise, the key of the dictionary data corresponding to the output function is the specific statement node; the output instrumentation function is to insert the statement execution record function CountOutput before each return statement; 1-2) Identify nodes in the AST (Abstract Syntax Tree), identify branch nodes, construct a set of branch nodes Branches, and generate a set of branch functions F by constructing instrumentation functions based on branch predicate rules insert Proceed to step 1-4); 1-3) Identify nodes in the AST abstract syntax tree, identify statement nodes, construct a statement node set Statements, and generate a statement function set F statements Proceed to step 1-4); the statement instrumentation function is to insert a statement execution record function CountState before each statement; 1 - 4) Combine the branch function set F branches , the output function set F outputs , and the statement function set F statements to form the instrumentation function set F insert , and combine it with the program under test Code to generate the instrumented program Code insert ; (1-5) Perform node analysis on the abstract syntax tree (AST), extract the edge relationships between its nodes, and generate an initial edge set Edge origin , and obtain its corresponding control flow graph (CFG); construct paths using depth-first search (DFS), where: during the path construction process, according to the Z-path coverage criterion, loop nodes are traversed only once; identify relevant edges through depth-first traversal (DFS) to form complete paths; finally, generate a target path set Path target ; 1 - 6) Identify the target path set Path target For the branch nodes in it, obtain the branch complexity of the path through the branch complexity analysis program, sort the paths according to the branch complexity, and obtain the sorted target path set Path order .
3. The test case evolutionary generation method guided by multiple coverage metrics according to claim 1, wherein In step 1-2), the branch predicate rules are as follows: Branch weight: w = k * n, k = 0.1, where n is the number of program branch nodes; Branch predicate: Boolean; Branch function expression: 1; Branch predicate: a <= b; Branch function expression: w * (a - b) + 1; Branch predicate: a >= b; Branch function expression: w * (b - a) + 1; Branch predicate: a == b; Branch function expression: w * abs(a - b) + 1; Branch predicate: a!= b; Branch function expression: 1; Branch predicate: Branch function expression X and Branch function expression Y; Branch function expression: max(Branch function expression X, Branch function expression Y); Branch predicate (judgment connection predicate): Branch function expression X or Branch function expression Y; Branch function expression: min(Branch function expression X, Branch function expression Y).
4. The test case evolutionary generation method guided by multiple coverage metrics according to claim 1, characterized in that Step (3) specifically includes the following steps: 3-1) Set the maximum number of iterations Max for the current population; 3-2) Integrate the sorted target path set Path order , the instrumented program Code insert , the initial population Pop origin and the execution path set Path exec The information enters step 3-3); 3-3) Execute the test case set for population construction through the optimization algorithm program OP, and obtain the results of the instrumentation functions of the population to enter step 3-4); 3-4) Through the coverage analysis program CA, extract the CountState function, CountOutput function, and branch distance information of the individuals in the current population Pop current from the population as a whole, and then calculate the branch coverage BC, output coverage OC, and statement coverage SC of the individuals with respect to the whole, and proceed to step 3-5); 3 - 5) By comprehensive analysis, calculate the Pareto front of the current population Pop current and obtain the overall coverage information Total current of the current population Pop coverage ; Verify the overall coverage information Total current and the Pareto front of the current population Pop coverage . If the overall coverage information Total coverage passes the verification by the set coverage rule or the Pareto front passes the verification by the set coverage rule, add the current population Pop current to the result test case set T result and enter step 3 - 7); Otherwise, enter step 3 - 6); Among them, the coverage rule of the overall coverage information Total coverage is: For the overall branch coverage BC, output coverage OC, and statement coverage SC, judge whether the result of weighted summation calculation according to the preset parameters reaches the set threshold. If it reaches, the verification passes; Otherwise, the verification fails; The coverage rule of the Pareto front is: If the branch coverage BC, output coverage OC, and statement coverage SC of the individuals in the population are all greater than or equal to the preset threshold, the verification passes; Otherwise, the verification fails; 3 - 6) Judge the number of iterations of the current population Pop current ; if the number of iterations of the current population Pop current reaches the maximum number of iterations Max, then add the current population Pop current to the result test case set T result , and terminate the iteration; otherwise, optimize the current population Pop current through the population crossover, selection, and mutation procedures to generate a new next-generation population Pop new and enter step 3 - 3); 3-7) Final output result test case set T result .