Test case sorting method based on multi-arm bandit algorithm
By combining coverage information and fault probability prediction, the priority of test cases is dynamically adjusted, the problem of inefficient testing in complex industrial software systems is solved, and efficient fault detection and resource utilization are achieved.
Patent Information
- Application Number
- CN202510241202.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
Existing test case sorting methods cannot effectively deal with complex industrial software systems, resulting in inefficient testing.
The multi-arm gambling machine algorithm combining coverage information and failure probability prediction is used to dynamically adjust the priority of test cases to ensure that test cases with high failure probability and high coverage are performed first.
It significantly improves fault detection efficiency, reduces waste of testing resources, and adapts to industrial software systems of different sizes and complexities.
Smart Images

Figure CN120179552A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of software testing in software engineering, and particularly relates to a test case sorting method based on the multi-armed bandit (MAB) algorithm, which can be used to optimize the execution order of test cases and improve the fault detection efficiency. Background Art
[0002] In modern software engineering, software testing is a key link to ensure software quality. With the continuous increase in the complexity of software systems, the number and execution time of test cases also increase accordingly. Especially in the field of industrial software, such as computer-aided engineering (CAE) tools, the execution cost of test cases is very high. Traditional test case sorting methods usually rely on static priority rules and cannot dynamically adapt to the execution results and coverage information of test cases, resulting in low test efficiency. To improve test efficiency, researchers have been exploring more intelligent and flexible test case sorting methods to cope with increasingly complex software systems.
[0003] However, existing test case sorting methods mainly rely on heuristic algorithms or statistical methods based on historical data, and these methods often perform poorly when dealing with large-scale and complex industrial software. For example, benchmark datasets such as Siemens, SIR, and Defects4J are widely used to evaluate test case sorting methods, but these datasets are usually small in scale and cannot fully reflect the true complexity of industrial software. Test cases of industrial software usually involve a large number of physical phenomenon simulations and complex dependencies, which makes it difficult for traditional sorting methods to effectively cope with. In addition, the specific domain behavior characteristics of industrial software (such as highly dependent test cases and complex physical phenomenon simulations) further increase the difficulty of test case sorting. These characteristics make the dependencies between test cases more complex, and traditional static sorting methods are difficult to capture these dynamic changes, resulting in low test efficiency.
[0004] In recent years, reinforcement learning algorithms have made significant progress in optimization problems. Especially the multi-armed bandit (MAB) algorithm, because it can balance exploration and exploitation in a dynamic environment, has gradually been applied to the test case ranking problem. The core idea of the MAB algorithm is to continuously try different test cases and adjust the strategy according to the feedback, so as to find a balance between exploring new test cases and exploiting known effective test cases. This method is particularly suitable for dynamically changing test environments, and can dynamically adjust the ranking strategy according to the failure scores and coverage information of test cases, thereby improving the test efficiency. Coverage information is an important indicator for evaluating the effectiveness of test cases, usually including code coverage, branch coverage, etc. By combining coverage information, the priority of test cases can be evaluated more accurately, thereby improving the test efficiency. Failure probability prediction is to analyze historical data and the execution results of test cases to predict which test cases are more likely to discover new failures. Combining coverage information and the results of failure probability prediction can further improve the accuracy of test case ranking. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to dynamically optimize the execution order of test cases based on the multi-armed bandit algorithm to improve the fault detection efficiency. To solve this problem, the present invention proposes a multi-armed bandit algorithm that combines coverage information and failure probability prediction, and ensures that test cases with high failure probability and high coverage can be executed preferentially by dynamically adjusting the priority of test cases.
[0006] The technical solution of the present invention:
[0007] A test case ranking method based on the multi-armed bandit algorithm, the specific steps are as follows:
[0008] Step (1), Collect the test case code files of all CAE open source projects from Github to obtain the initial data set, and obtain the enhanced data set D through data augmentation technology aug . Traverse the enhanced data set D in units of open source projects aug , Extract syntactic features and structural features from several test case code files contained in each project, including variable declarations, control flow statements, operators, etc., and thus obtain the text feature data set D text_feature . Through the TF-IDF vectorization method, the text features contained in the text feature data set D text_featur are sequentially converted into digital features, and thus the corresponding digital feature data set D is obtained digital_feature . Normalize all the features contained in the digital feature data set D digital_feature respectively to ensure that the value of each feature after normalization is within the interval [0,1], and finally obtain the preprocessing data set D process, so as to predict the failure probability in step (2).
[0009] Step (2): Input the preprocessed dataset D in step (1), process , divide it, and divide it into a training set D train and a validation set D val . Use the XGBoost model to train on the training set D train to predict the failure probability of each test case; on the validation set D val , evaluate the stability of the model through cross-validation, and optimize the parameter configuration of the model by combining hyperparameter tuning to ensure that it has good generalization ability and prediction accuracy on both the training set and the test set. Then, from the CAE open-source projects collected in step (1), use the gcov tool to extract the code coverage and method coverage. Finally, obtain the failure probability prediction results and coverage information of all test cases.
[0010] Step (3): Use the predicted failure probability prediction results and coverage information as input, and use the multi-armed bandit algorithm to sort the test cases. The construction of the reward value is dynamically calculated by combining the prediction probability and the coverage rate. The weight of the prediction probability is set to ω p , and the weight of the coverage rate is set to ω c , and the final reward value is generated by weighted summation. There are four solver algorithms for the multi-armed bandit algorithm, including ThompsonSampling, UCB1, BayesianUCB, and EpsilonGreedy. These solver algorithms use the reward value to dynamically adjust the priority of the test cases to ensure that the test cases with high failure probability and high coverage rate can be executed first. By executing the sorted test cases, calculate the average position failure detection rate (APFD) value to evaluate the sorting effect.
[0011] Furthermore, step (1) specifically includes the following steps:
[0012] 1-1) Collect the test case code files of all CAE open-source projects from Github to obtain the initial dataset. To improve the diversity of the experiment and enhance the generalization ability of the model, an error injection method is used for data augmentation. Based on the code files in the initial dataset, randomly insert code statements that may cause project errors to generate diverse code variants and simulate different programming scenarios. Specifically, the system first randomly selects a file from a group of predefined code files, and then randomly selects a statement from a carefully designed statement pool for insertion. The selected statements cover multiple common code operations, such as mesh refinement, definition of the finite element space, common operations related to mathematical models, and some program logic error statements such as uninitialized variables and array out-of-bounds. Finally, obtain the dataset D after data augmentationaug。
[0013] 1 - 2) Traverse the enhanced dataset D in units of open - source projects aug , for each project, generate the abstract syntax tree (AST) corresponding to the code of each test case code file contained therein, and extract the key structured information contained in the code, such as variable declarations, control - flow statements (such as if, for, while), operators (such as addition, subtraction, multiplication, division, logical operators), etc., to obtain the text feature data of all code files in the project. Finally, obtain the text feature dataset D that takes the project as the unit and contains the text features corresponding to all project codes text_feature .
[0014] 1 - 3) Traverse the text feature dataset D text_feature , and convert the text features contained therein into digital features in turn through the TF - IDF vectorization method, thereby obtaining the corresponding digital feature dataset D digital_feature .
[0015] 1 - 4) Traverse the digital feature dataset D digital_feature , for each feature value, map it to the interval [0, 1] through the Min - Max normalization formula, so as to eliminate the deviation caused by different features due to different dimensions and value ranges, ensure that the features have a unified scale, and provide standardized input for the subsequent machine - learning model training. Finally, obtain the pre - processed dataset D process .
[0016] Furthermore, step (2) specifically includes the following steps:
[0017] 2 - 1) Divide the dataset D process into a training set D train and a test set D val . Usually, the training set D train accounts for 80% of the total data volume, and the test set D val accounts for 20%. Then initialize the XGBoost model and set its basic parameters. The basic parameters include the learning rate, maximum depth, number of trees, etc. The default value of the learning rate is set to 0.1, the maximum depth value is 6, and the number of trees is adjusted according to the data scale and computing resources. During the training process, the model optimizes the parameters by minimizing the loss function (cross - entropy loss) and gradually improves the prediction ability. To further verify the stability of the model, 5 - fold cross - validation is adopted to comprehensively evaluate the performance of the model
[0018] 2 - 2) Use the initialized XGBoost model to train on the training set D train , and predict the failure probability of each test case. During the training process, the model optimizes the parameters by minimizing the loss function (cross - entropy loss) and gradually improves the prediction ability. On the validation set Dval Evaluate the stability of the model through cross - validation, and optimize the parameter configuration of the model by combining hyperparameter tuning to ensure its good generalization ability and prediction accuracy on both the training set and the test set. Finally, obtain the fault prediction probability score of the test case.
[0019] 2 - 3) Compile the CAE open - source projects collected in step (1). When compiling, enable the coverage instrumentation option of the GCC compiler to generate the instrumented executable file and the associated.gcno file. Then, run the instrumented executable file and execute the test cases to dynamically generate the.gcda file to record the code execution situation. Then, use the gcov tool to parse the.gcda file to generate a detailed coverage report. Finally, obtain the coverage information of the test cases.
[0020] Furthermore, step (3) specifically includes the following steps:
[0021] 3 - 1) Input the coverage information (code coverage and method coverage) and the fault probability prediction results of the test cases as the initial data of the multi - armed bandit model. Construct the reward value by dynamically calculating the combination of the prediction probability and the coverage. The weight of the prediction probability is set as ω p , and the weight of the coverage is set as ω c , and the final reward value is generated by weighted summation.
[0022] 3 - 2) Use four solver algorithms (Thompson Sampling, UCB1, Bayesian UCB, EpsilonGreedy) to sort the test cases. By dynamically weighing exploration and exploitation, generate the optimal execution sequence. Each algorithm adjusts the test order based on different strategies (such as confidence intervals, random sampling, or greedy algorithms). Finally, by comparing their fault detection rate (APFD) metrics, determine the solver with the best performance and optimize the test efficiency. The specific implementation of the solver algorithms is as follows:
[0023] a) The EpsilonGreedy solver is based on the ε - greedy strategy, and its core idea is to balance between exploration and exploitation. In each step, the EpsilonGreedy solver randomly selects a test case with probability ε (exploration) and selects the test case with the currently estimated highest reward with probability 1 - ε (exploitation). The estimated value of the reward is updated by formula (1):
[0024]
[0025] Where: is the estimated reward value of test case i at time t; r i (t) is the actual reward value of test case i at time t; Ni (t) is the number of times test case i has been selected before time t.
[0026] To balance exploration and exploitation, the EpsilonGreedy solver can dynamically adjust the value of ε as shown in Equation (2). For example, over time, the value of ε is gradually decreased:
[0027]
[0028] where: ε0 is the initial exploration probability; α is the decay coefficient that controls the decay rate of ε.
[0029] b) The UCB1 solver is based on the Upper Confidence Bound (UCB) strategy. Its core idea is to balance exploration and exploitation through the confidence interval. The calculation formula of UCB1 is as shown in Equation (3):
[0030]
[0031] where: UCB i (t) represents the upper confidence bound value of test case i at time t; is the estimated reward value of test case i at time t; N i (t) is the number of times test case i has been selected before time t;
[0032] In each step, the UCB1 solver selects the test case i with the maximum UCB value * , as shown in Equation (4):
[0033] i * = argmax i UCB i (t)(4)
[0034] The estimated value of the reward is also updated through Equation (1).
[0035] c) The BayesianUCB solver is based on the Bayesian Upper Confidence Bound strategy, assuming that the reward of each test case follows a Beta distribution. Its core idea is to estimate the distribution of the reward through Bayesian update and calculate the upper confidence bound. The calculation of BayesianUCB is as shown in Equation (5):
[0036]
[0037] where: BayesianUCB i (t) is the Bayesian upper confidence bound value of test case i at time t; α i (t) and β i(t) is the Beta distribution parameter of test case i at time t; c is the confidence parameter that controls the intensity of exploration.
[0038] At each step, the BayesianUCB solver selects the test case i with the largest BayesianUCB value * , as shown in formula (6):
[0039] i * = argmax i BayesianUCB i (t)(6)
[0040] According to the test case i * Rewards Update Beta distribution parameters and As shown in formulas (7) and (8) respectively:
[0041]
[0042] d) The ThompsonSampling solver is based on the Thompson sampling strategy, the core idea of which is to select test cases by sampling from the Beta distribution of each test case. The specific steps are as follows:
[0043] First, a value θ is sampled from the Beta distribution for each test case i i , as shown in formula (9):
[0044] θ i ~Beta(α i (t),β i (t)) (9)
[0045] Then, select the test case i with the largest sampling value * , as shown in formula (10):
[0046] i * = argmax i θ i (10)
[0047] Finally, according to the test case i * Rewards Update Beta distribution parameters and The updating method is as shown in formulas (7) and (8).
[0048] 3-3) Execute the test cases in the sorted order generated by the multi-armed bandit algorithm and record the execution results of each test case. During the execution process, the pass or fail status of each test case is recorded. These execution results are not only used to calculate the Average Percentage of Faults Detected (APFD) value to evaluate the sorting effect, but also fed back to the multi-armed bandit algorithm to dynamically adjust the internal state of the solver (such as updating the reward estimate or Beta distribution parameters), thereby further optimizing the subsequent test case sorting strategy.
[0049] 3-4) To evaluate the sorting effect, the Average Percentage of Faults Detected (APFD) is used as a quantitative metric. The APFD value measures the efficiency of the sorted test cases in fault detection. The higher the value, the better the sorting strategy. The calculation formula of APFD is shown in Formula (11):
[0050]
[0051] where: m is the total number of faults detected in the test cases; n is the total number of test cases; TF j is the position where the j-th fault is first detected in the sorted test cases.
[0052] By calculating the APFD value, the effectiveness of the sorting strategy can be intuitively evaluated. If the APFD value is high, it indicates that the sorted test cases can detect more faults earlier, thereby improving the test efficiency. In addition, the present invention feeds back the test execution results to the MAB algorithm to further optimize the test case sorting strategy. Through continuous iterative optimization and dynamic adjustment of the balance between exploration and exploitation, the fault detection efficiency can be significantly improved and the waste of test resources can be reduced.
[0053] Compared with the prior art, the present invention has the following advantages and effects:
[0054] The method of the present invention can dynamically optimize the execution order of test cases and significantly improve the fault detection efficiency. By combining the coverage information and the fault probability prediction, the present invention can preferentially execute the test cases with high fault probability and high coverage, reducing the waste of test resources. In addition, the present invention has high scalability and can adapt to industrial software systems of different scales and complexities. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is a schematic flowchart of the test case sorting method based on the multi-armed bandit algorithm of the present invention.
[0056] Figure 2 is a sub-flowchart of the data preprocessing stage of the test case sorting method based on the multi-armed bandit algorithm of the present invention.
[0057] Figure 3 It is a process sub - graph of fault probability prediction in the test case sorting method based on the multi - armed bandit algorithm of the present invention.
[0058] Figure 4 It is a process sub - graph of the multi - armed bandit algorithm sorting in the test case sorting method based on the multi - armed bandit algorithm of the present invention. Detailed implementation manners
[0059] The following further illustrates the detailed implementation manners of the present invention in combination with the accompanying drawings and technical solutions.
[0060] As Figure 1 shown, the test case sorting method empowered by the multi - armed bandit algorithm of the present invention proceeds as follows: First, collect the test case code files of all CAE open - source projects from Github, and obtain the enhanced dataset D through data augmentation technology. aug Traverse the enhanced dataset D on a per open - source project basis. aug Extract syntactic features and structural features from several test case code files contained in each project, thereby obtaining the text feature dataset D. text_feature Through the TF - IDF vectorization method, convert the text features contained in the text feature dataset D text_feature into digital features in sequence, thereby obtaining the corresponding digital feature dataset D. digital_feature Perform normalization processing on all features contained in the digital feature dataset D digital_feature to ensure that the value of each feature after normalization is within the interval [0, 1], and finally obtain the pre - processed dataset D. process Then divide the dataset D process into a training set D train and a validation set D val . Use the XGBoost model to train on the training set D train to predict the fault probability of each test case; evaluate the stability of the model through cross - validation on the validation set D val and optimize the parameter configuration of the model by combining hyperparameter tuning to ensure that it has good generalization ability and prediction accuracy on both the training set and the test set. Then, from the collected CAE open - source projects, use the gcov tool to extract the code coverage and method coverage. Construct the reward value through dynamic calculation by combining the predicted probability and coverage. Finally, select four solver algorithms to dynamically adjust the priorities of the test cases, and finally generate the test case execution sequence. By executing the sorted test cases, calculate the average position fault detection rate (APFD) value to evaluate the sorting effect. The detailed implementation manners are as follows:
[0061] (1) AsFigure 2 As shown, collect the test case code files of all CAE open-source projects from Github to obtain the initial dataset, and obtain the enhanced dataset D through data augmentation techniques aug Traverse the enhanced dataset D in units of open-source projects aug , extract syntactic features and structural features from several test case code files contained in each project, including variable declarations, control flow statements, operators, etc., and thus obtain the text feature dataset D text_feature Through the TF-IDF vectorization method, sequentially convert the text features contained in the text feature dataset D text_feature into digital features, and thus obtain the corresponding digital feature dataset D digital_feature Perform normalization processing on all features contained in the digital feature dataset D digital_feature to ensure that the value of each feature after normalization is within the range of [0, 1], and finally obtain the preprocessed dataset D process。 The specific implementation process is as follows:
[0062] 1-1) Collect the test case code files of all CAE open-source projects from Github to obtain the initial dataset, as shown in Table 1. To improve the diversity of the experiment and enhance the generalization ability of the model, the error injection method is used for data augmentation. Based on the code files in the initial dataset, randomly insert code statements that may cause project errors to generate diverse code variants and simulate different programming scenarios. Specifically, the system first randomly selects a file from a group of predefined code files, and randomly selects a statement from a carefully designed statement pool for insertion. The selected statements cover multiple common code operations, such as mesh refinement, definition of finite element space, common operations related to mathematical models, and some program logic error statements such as uninitialized variables and array out-of-bounds. As shown in Table 2. Finally, obtain the enhanced dataset D aug。
[0063] Table 1 Summary of test case code content
[0064] File Name Summary of Code Content 4623_b.py Contains mesh refinement operations and finite element space definition test_file_logger.py Test logging function test_eigen_solvers.py Test eigenvalue solver 11231_c.py Contains operations related to mathematical models test_point.py Test geometric point operations test_flags.py Test flag setting test_linear_solvers.py Test linear solver test_vector_interface.py Test vector interface test_cad_tessellation_modeler.py Test CAD tessellation modeling … …
[0065] Table 2 Data augmentation example
[0066] File Name Enhancement Type Example of inserted code statement 4623_b.py Array out-of-bounds error array
[100] =0 (array size is 50) test_file_logger.py Uninitialized variable int uninitialized_var; test_eigen_solvers.py Mesh refinement operation mesh.refine(2); test_point.py Mathematical model operation model.solve_equation(); test_linear_solvers.py Uninitialized variable double result; test_vector_interface.py Array out-of-bounds error vector
[100] =1.0; (vector size is 50) test_cad_tessellation_modeler.py Calculate contact force operation force = node.GetSolutionStepValue() … … …
[0067] 1-2) Traverse the enhanced dataset D in units of open-source projects aug, for each project, generate the Abstract Syntax Tree (AST) corresponding to the code of each test case code file contained therein, and extract the key structural information contained in the code, such as variable declarations, control flow statements (such as if, for, while), operators (such as addition, subtraction, multiplication, division, logical operators), etc., to obtain the text feature data of all code files in the project. Finally, obtain the text feature dataset D containing the text features corresponding to all project codes on a project-by-project basis text_feature . As shown in Table 3 below:
[0068] Table 3 Text Feature Dataset
[0069]
[0070] 1-3) Traverse the text feature dataset D text_feature , and convert the text features contained therein into digital features in turn through the TF-IDF vectorization method to obtain the corresponding digital feature dataset D digital_feature . As shown in Table 4 below:
[0071] Table 4 Digital Feature Dataset
[0072]
[0073] 1-4) Traverse the digital feature dataset D digital_feature , for each feature value, map it to the interval [0,1] through the Min-Max normalization formula (Formula (15)), so as to eliminate the deviation caused by different features due to different dimensions and value ranges, ensure that the features have a unified scale, and provide standardized input for the subsequent training of machine learning models. The formula for feature normalization is shown in Formula (12) below. Finally, obtain the preprocessing dataset D process . For example, the normalized digital features are shown in Table 5 below:
[0074] Table 5 Features after Normalization of Test Cases
[0075]
[0076]
[0077] where: x′ ij is the j-th feature value of the i-th test case after normalization; x ij is the j-th feature value of the i-th test case; min(x j ) is the minimum value of the j-th feature; max(x j ) is the maximum value of the j-th feature.
[0078] (2) As Figure 3 shown, input the preprocessing dataset D process, partition it into the training set \(D\) train and the validation set \(D\) val . Use the XGBoost model to train on the training set \(D\) train to predict the failure probability of each test case; on the validation set \(D\) val evaluate the stability of the model through cross - validation, and combine hyperparameter tuning to optimize the parameter configuration of the model, ensuring that it has good generalization ability and prediction accuracy on both the training set and the test set. The hyperparameters are shown in Table 6 below. Then, from the collected CAE open - source projects, use the gcov tool to extract the code coverage and method coverage. Finally, obtain the failure probability prediction results and coverage information of all test cases. The specific implementation process is as follows:
[0079] Table 6 Hyperparameter Table
[0080]
[0081] 2 - 1) Divide the dataset \(D\) process into the training set \(D\) train and the test set \(D\) val . Usually, the training set \(D\) train accounts for 80% of the total data volume, and the test set \(D\) val accounts for 20%. Then initialize the XGBoost model and set its basic parameters. These parameters include the learning rate, maximum depth, number of trees, etc. The default value of the learning rate is set to 0.1, the maximum depth value is 6, and the number of trees is adjusted according to the data scale and computing resources. During the training process, the model optimizes the parameters by minimizing the loss function (cross - entropy loss) and gradually improves the prediction ability. To further verify the stability of the model, 5 - fold cross - validation is adopted to comprehensively evaluate the performance of the model.
[0082] 2 - 2) Use the initialized XGBoost model to train on the training set \(D\) train to predict the failure probability of each test case. During the training process, the model optimizes the parameters by minimizing the loss function (cross - entropy loss) and gradually improves the prediction ability. On the validation set \(D\) val evaluate the stability of the model through cross - validation, and combine hyperparameter tuning to optimize the parameter configuration of the model, ensuring that it has good generalization ability and prediction accuracy on both the training set and the test set. Finally, obtain the failure prediction probability scores of the test cases. For example, as shown in Table 7 below:
[0083] Table 7 Failure Probability Scores of Each Test Case Finally Obtained
[0084] File Name Failure probability score 4623_b.py 0.851 test_file_logger.py 0.851 test_eigen_solvers.py 0.742 11231_c.py 0.698 3442_b.py 0.664 10649_b.py 0.515 test_point.py 0.501 11307_b.py 0.395 test_flags.py 0.189 test_linear_solvers.py 0.185 test_vector_interface.py 0.173 test_cad_tessellation_modeler.py 0.119 …
[0085] (2 - 3) Compile the CAE open - source projects collected in step (1), enabling the code coverage instrumentation option of the GCC compiler during compilation to generate the instrumented executable file and the associated.gcno file; then, run the instrumented executable file and execute the test cases to dynamically generate the.gcda file to record the code execution situation; next, use the gcov tool to parse the.gcda file to generate a detailed code coverage report. Finally, obtain the code coverage information of the test cases. As shown in Table 8 below:
[0086] Table 8 Code coverage information of each test case finally obtained
[0087] File Name Method coverage Line coverage 4623_b.py 0.389 0.379 test_file_logger.py 0.559 0.469 test_eigen_solvers.py 0.604 0.500 11231_c.py 0.409 0.310 11307_b.py 0.411 0.356 10649_b.py 0.417 0.325 test_point.py 0.623 0.479 3442_b.py 0.439 0.434 test_flags.py 0.592 0.499 test_linear_solvers.py 0.624 0.477 test_vector_interface.py 0.555 0.462 test_cad_tessellation_modeler.py 0.554 0.462 … … …
[0088] (3) As Figure 4 shown, take the predicted failure probability prediction result and the code coverage information as inputs, and dynamically construct the reward value by combining the predicted probability and the code coverage. There are four solver algorithms for the multi - armed bandit algorithm, including Thompson Sampling, UCB1, BayesianUCB, and EpsilonGreedy. Adopt these four solver algorithms and use the reward value to dynamically adjust the priorities of the test cases to ensure that the test cases with high failure probability and high code coverage can be executed preferentially. Finally, by comparing their Average Percentage of Faults Detected (APFD) metrics, determine the solver with the best performance and optimize the test efficiency. After the execution, a test case sequence will be generated, as shown in Table 9 below.
[0089] Table 9 Test case execution sequence finally obtained
[0090]
[0091]
[0092] By calculating the APFD value, the effectiveness of the sorting strategy can be intuitively evaluated. If the APFD value is high, it indicates that the sorted test cases can detect more faults earlier, thus improving the test efficiency. In addition, the present invention feeds back the test execution results to the MAB algorithm to further optimize the sorting strategy of the test cases. Through continuous iterative optimization and dynamically adjusting the balance between exploration and exploitation, the fault detection efficiency can be significantly improved, and the waste of test resources can be reduced.
[0093] For example, use the Thompson Sampling solver algorithm for sorting, and then by executing the sorted test cases, calculate the Average Position of Fault Detection (APFD) value to be 0.8667, indicating that the method enables the sorted test cases to detect more faults earlier.
Claims
1. A test case sorting method based on a multi-armed bandit algorithm, characterized in that: The specific steps are as follows: Step (1) Collect the test case code files of all CAE open source projects from Github to obtain the initial data set, and obtain the enhanced data set D through data enhancement technology. aug ; Taking open source projects as units, traverse the enhanced dataset D aug , extract grammatical features and structural features from several test case code files contained in each project, including variable declarations, control flow statements, and operators, thereby obtaining a text feature dataset D text_feature ; Through the TF-IDF vectorization method, the text feature dataset D text_featur The text features contained in are converted into digital features in turn, thereby obtaining the corresponding digital feature dataset D digital_feature ; For the digital feature dataset D digital_feature All the features contained in are normalized separately to ensure that the normalized value of each feature is in the interval [0,1], and finally the preprocessed data set D is obtained. process ; Step (2): Input the preprocessed data set D in step (1) process , divide it into training set D train and validation set D val ; Use XGBoost model in training set D train Train on the validation set D val The stability of the model is evaluated by cross-validation, and the parameter configuration of the model is optimized by combining hyperparameter tuning to ensure that it has good generalization ability and prediction accuracy on both the training set and the test set. Then, the gcov tool is used to extract code coverage and method coverage from the CAE open source projects collected in step (1). Finally, the fault probability prediction results and coverage information of all test cases are obtained. Step (3), taking the predicted failure probability prediction result and coverage information as input, and using the multi-armed bandit algorithm to sort the test cases; the reward value is constructed by combining the predicted probability and coverage dynamic calculation, and the weight of the predicted probability is set to ω p , the coverage weight is set to ω c , the final reward value is generated by weighted summation; the multi-armed bandit algorithm has four solver algorithms, including ThompsonSampling, UCB1, BayesianUCB and EpsilonGreedy; the solver algorithm uses the reward value to dynamically adjust the priority of test cases to ensure that test cases with high failure probability and high coverage can be executed first; by executing the sorted test cases, the average position fault detection rate APFD value is calculated to evaluate the sorting effect.
2. A test case sorting method based on a multi-armed bandit algorithm according to claim 1, characterized in that: Step (1) specifically includes the following steps: 1-1) Collect the test case code files of all CAE open source projects from Github to obtain the initial data set; In order to improve diversity and enhance the generalization ability of the model, the error injection method is used for data enhancement; Based on the code files in the initial data set, randomly insert code statements that may cause project errors to generate diverse code variants and simulate different programming scenarios; Specifically, the system first randomly selects a file from a set of predefined code files, and randomly selects a statement from a designed statement pool for insertion; The selected statements cover multiple code operations, including mesh refinement, definition of finite element space, operations related to mathematical models, and uninitialized variables and array out-of-bounds; Finally, the data enhanced data set D is obtained aug; 1-2) Traverse the enhanced dataset D based on open source projects aug For each project, an abstract syntax tree corresponding to each test case code file is generated, and key structural information contained in the code is extracted, including variable declarations, control flow statements, and operators, to obtain the text feature data of all code files in the project; finally, a text feature dataset D containing the text features corresponding to all project codes is obtained in units of projects. text_feature ; 1-3) Traverse the text feature dataset D text_feature , the text features contained in it are converted into digital features in turn through the TF-IDF vectorization method, thereby obtaining the corresponding digital feature dataset D digital_feature ; 1-4) Traverse the digital feature dataset D digital_feature For each eigenvalue, the Min-Max normalization formula is used to map it to the interval [0,1]; finally, the preprocessed data set D is obtained. process .
3. A test case sorting method based on a multi-armed bandit algorithm according to claim 1, characterized in that: Step (2) specifically includes the following steps: 2-1) Data set D process Divide into training set D train and the test set D val ; Then initialize the XGBoost model and set its basic parameters; the basic parameters include learning rate, maximum depth, and number of trees; during the training process, the model optimizes the parameters by minimizing the loss function, namely the cross entropy loss; in order to further verify the stability of the model, a 5-fold cross validation is used; 2-2) Use the initialized XGBoost model on the training set D train The model is trained on the validation set D to predict the failure probability of each test case. During the training process, the model optimizes the parameters by minimizing the loss function, namely the cross entropy loss, and gradually improves the prediction ability. val The stability of the model is evaluated through cross-validation, and the parameter configuration of the model is optimized in combination with hyperparameter tuning to ensure that it has good generalization ability and prediction accuracy on both the training set and the test set; finally, the fault prediction probability score of the test case is obtained; 2-3) Compile the CAE open source projects collected in step (1), enable the coverage instrumentation option of the GCC compiler during compilation, generate the instrumented executable file and the associated .gcno file; then, run the instrumented executable file and execute the test case, dynamically generate the .gcda file to record the code execution; then, use the gcov tool to parse the .gcda file and generate a detailed coverage report; finally, obtain the coverage information of the test case.
4. A test case sorting method based on a multi-armed bandit algorithm according to claim 1, characterized in that: Step (3) specifically includes the following steps: 3-1) Input the coverage information and fault probability prediction results of the test case as the initial data of the multi-armed bandit model; construct the reward value by combining the predicted probability and coverage dynamic calculation, and the weight of the predicted probability is set to ω p , the coverage weight is set to ω c , the final reward value is generated by weighted summation; 3-2) Four solver algorithms, Thompson Sampling, UCB1, Bayesian UCB, and Epsilon Greedy, are used to sort the test cases, and the optimal execution sequence is generated through dynamic trade-off exploration and utilization. Each algorithm adjusts the test order based on different strategies, and finally determines the solver with the best performance and optimizes the test efficiency by comparing their fault detection rate APFD indicators. The solver algorithms are as follows: a) The EpsilonGreedy solver is based on the ε-greedy strategy and makes a trade-off between exploration and exploitation. In each step, the EpsilonGreedy solver randomly selects a test case with probability ε and selects the test case with the highest current estimated reward with probability 1-ε. The estimated value of the reward is updated by formula (1): in: is the estimated reward value of test case i at time t; r i (t) is the actual reward value of test case i at time t; N i (t) is the number of times test case i was selected before time t; The EpsilonGreedy solver dynamically adjusts the value of ε as shown in formula (2); Among them: ε0 is the initial exploration probability; α is the decay coefficient, which controls the decay speed of ε; b) The UCB1 solver is based on the upper confidence bound strategy and uses confidence intervals to balance exploration and exploitation. The calculation formula of UCB1 is shown in formula (3): Where: UCB i (t) represents the upper confidence bound of test case i at time t; is the estimated reward value of test case i at time t; N i (t) is the number of times test case i was selected before time t; At each step, the UCB1 solver selects the test case i with the largest UCB value * , as shown in formula (4): i * =argmax i UCB i (t)(4) The estimated value of the reward is also updated using formula (1); c) The BayesianUCB solver is based on the Bayesian upper confidence bound strategy, assuming that the reward of each test case follows a Beta distribution. The distribution of the reward is estimated through Bayesian updating, and the upper confidence bound is calculated. The calculation of BayesianUCB is shown in formula (5): Among them: BayesianUCB i (t) is the Bayesian upper confidence bound of test case i at time t; α i (t) and β i (t) is the Beta distribution parameter of test case i at time t; c is the confidence parameter, which controls the intensity of exploration; At each step, the BayesianUCB solver selects the test case i with the largest BayesianUCB value * , as shown in formula (6): i * =argmax i BayesianUCB i (t) (6) According to the test case i * Rewards Update Beta distribution parameters and As shown in formulas (7) and (8) respectively: d) The ThompsonSampling solver selects test cases based on the Thompson sampling strategy by sampling from the Beta distribution of each test case; the specific steps are as follows: First, a value θ is sampled from the Beta distribution for each test case i i , as shown in formula (9): i i ~Beta(α i (t),β i (t)) (9) Then, select the test case i with the largest sampling value * , as shown in formula (10): i * =argmax i θ i (10) Finally, according to the test case i * Rewards Update Beta distribution parameters and The updating method is as shown in formulas (7) and (8); 3-3) Execute the test cases in the order of the test case sorting generated by the multi-armed bandit algorithm, and record the execution results of each test case; during the execution process, the pass or fail status of each test case will be recorded; these execution results are not only used to calculate the average position fault detection rate APFD value to evaluate the sorting effect, but also fed back to the multi-armed bandit algorithm to dynamically adjust the internal state of the solver, thereby further optimizing the subsequent test case sorting strategy; 3-4) In order to evaluate the sorting effect, the average position fault detection rate APFD is used as a quantitative indicator; the APFD value measures the efficiency of the sorted test cases in fault detection. The higher the value, the better the sorting strategy. The calculation formula of APFD is shown in formula (11): Where: m is the total number of faults detected in the test case; n is the total number of test cases; TF j is the location where the jth fault is first detected in the sorted test cases.
Citation Information
Cited By
Chip verification coverage rate improving method
CN120373222A