Software test case generation method and system, electronic equipment and storage medium
By combining ant colony algorithm and particle swarm optimization method, the software test cases are optimized to generate, and the problem of insufficient coverage and depth in complex path coverage scenarios is solved, and more efficient test case generation and more effective testing are achieved.
Patent Information
- Application Number
- CN202510476873.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-20
AI Technical Summary
When generating software test cases, the prior art faces the problem of insufficient path coverage, depth and complex security vulnerability disclosure capabilities in complex path coverage scenarios, and low search efficiency and slow convergence speed.
Using a method combining ant colony algorithm and particle swarm optimization, the user behavior model and input parameter space are built by initializing the test environment, and the ants are driven to explore paths in the path exploration space, and the path branch pheromone is updated to optimize the generation of test case sets.
It improves the generation efficiency, coverage and depth of test cases, enhances the effectiveness of tests, and can better reflect the real usage and complex interaction modes.
Smart Images

Figure CN120179560A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer software testing, and particularly to a method, a system, an electronic device and a storage medium for generating software test cases. Background Art
[0002] In the field of software engineering, software testing is a process for discovering and fixing defects in software, verifying whether the software meets the design requirements, and ensuring its quality and performance. It is a crucial link in the software development process. Test cases are the core part of software testing, directly related to the efficiency and coverage of testing, and affecting the quality and reliability of software products. A test case can be regarded as a specific test task, including elements such as input, operation, and expected output, and is a data set in various forms, such as execution paths, input data, and execution conditions.
[0003] The methods for generating test cases can generally be divided into the following several types:
[0004] First, test cases are manually written by testers. This means that testers need to have very rich test experience and relatively high professional levels. Moreover, there are certain drawbacks such as blindness, high costs, and difficulty in improving test coverage. Although the method of manually writing test cases is effective in some cases, with the continuous increase in the complexity of software systems, the manual writing of test cases has become increasingly difficult, time-consuming, and prone to missing errors, especially in cases where high coverage and in-depth detection are required. In addition, the method of manually writing test cases is difficult to adapt to the rapid iteration of software. Updating a function may require rewriting a large number of test cases, which significantly increases the costs of software development and maintenance.
[0005] Second, certain tools are used to automatically generate test cases based on certain methods, such as random testing, model-based testing (MBT), symbolic execution, search-based testing, etc. Although these tools have made progress in some fields, there are still many unsatisfactory aspects. For example, model-based testing requires detailed model design, which itself is a complex and time-consuming process; symbolic execution and search-based testing face the problem of state space explosion and are difficult to handle large software systems.
[0006] III. Combine with other algorithms. For example, adopt the genetic algorithm or the method of combining the genetic algorithm with other automated methods. For example, Chinese Patent Publication No. CN109344057B, with the invention name of "Combined Accelerated Test Case Generation Method Based on Genetic Method and Symbolic Execution", discloses a test case generation method combining the genetic algorithm, which generates test cases by using the search ability of the ant colony algorithm and the characteristic of being easy to integrate with other algorithms. However, for complex path coverage scenarios, the test cases generated by using the ant colony algorithm in the prior art are not ideal in terms of path coverage rate, depth, and the ability to reveal complex security vulnerabilities, and have low search efficiency and slow convergence speed. Since the diversity of user behavior patterns and software running environments is ignored when generating test case individuals, the generated test cases cannot comprehensively reflect the real usage situation, restricting the effectiveness of testing. Summary of the Invention
[0007] In view of the technical problems existing in the prior art, the present invention proposes a software test case generation method, system, electronic device, and storage medium to improve the generation efficiency of test cases and improve the test coverage rate, depth, and effectiveness of test cases.
[0008] To solve the above technical problems, according to one aspect of the present invention, the present invention provides a software test case generation method, including the following steps:
[0009] Initialize the test environment to obtain the path exploration space of the target software; wherein, the path exploration space includes nodes and directed edges between the nodes, and two nodes with a directed relationship form a path branch, and each path branch includes path branch conditions;
[0010] Construct a user behavior model and the input parameter space of the target software based on the collected user historical behavior data, wherein the user behavior model includes one or more user parameters simulating user behavior and their parameter values; the input parameter space includes multiple input variables, and the variable value of each input variable corresponds to one or more user parameter values;
[0011] Drive multiple ants to conduct path exploration in different regions of the path exploration space respectively, and update the path branch pheromone during the path exploration process; after each ant determines a path branch during the path exploration process, based on the path branch condition, select the corresponding input variable value from the input parameter space according to the user behavior model; after each ant finishes the path exploration, obtain the input variable values of all path branches corresponding to the explored path to form an initial test case individual; all initial test case individuals form an initial test case set; and
[0012] Map the initial test case set to an initial particle swarm, and iterate and optimize the initial particle swarm until the optimized test case set meets the optimization end condition.
[0013] Optionally, the steps for each ant to determine a path branch during path exploration include:
[0014] Based on the pointing relationship of the edge between two nodes, determine one or more candidate nodes that form a path branch with the current node;
[0015] Calculate the selection probability from the current node to each candidate node based on the path branch pheromone and the heuristic factor; and
[0016] Determine the candidate node with the highest selection probability as the next node that forms a path branch with the current node; wherein, the heuristic factor at least includes the user behavior matching degree weight.
[0017] Optionally, the heuristic factor further includes one or more of the following weights: detection execution time weight, running environment weight, path length weight, and security risk weight.
[0018] Optionally, the steps to update the path branch pheromone during path exploration include:
[0019] During path exploration, determine whether the currently explored path branch is marked with a security risk attribute mark. In response to the currently explored path branch being marked with a security risk attribute mark, increase the corresponding increment value based on the current pheromone value of the path branch marked with the security risk attribute mark.
[0020] Optionally, during path exploration, it further includes:
[0021] Determine the scenario corresponding to the current path exploration space area; and
[0022] Adjust the pheromone value of the path branch in a manner matching the scenario.
[0023] Optionally, the software test case generation method further includes: calculating the coverage rate of the explored paths for the path exploration space or a preset local path exploration space during the process of driving multiple ants to perform path exploration in different areas of the path exploration space, and adjusting the pheromone value of the path branches in the path exploration space based on the corresponding relationship between the coverage rate and the pheromone being negatively correlated.
[0024] Optionally, after each ant finishes path exploration, form candidate initial test case individuals with the input variable values of all path branches corresponding to the explored path. The method further includes:
[0025] Construct the maximum problem function of the initial test case set based on the environmental fitness and the user behavior model;
[0026] Configure the operating environment parameters for each candidate initial test case individual;
[0027] Calculate the environmental fitness score of each candidate initial test case individual based on the operating environment parameters of the candidate initial test case individual;
[0028] Calculate the user behavior simulation score of each candidate initial test case individual based on the user behavior model used when generating the candidate initial test case individual;
[0029] Calculate the maximum problem function value of the initial test case set based on the environmental fitness score and the user behavior simulation score of each candidate initial test case individual, and adjust the operating environment parameters of the candidate initial test case individual and / or the user behavior model used to generate the candidate initial test case individual to converge the maximum problem function to the maximum value; and
[0030] Determine the candidate initial test case individual when converging to the maximum value as the initial test case individual, and all the initial test case individuals constitute the initial test case set.
[0031] Optionally, one optimization step in the iterative optimization process of the initial particle swarm includes:
[0032] Construct the current test case set from the test case individuals corresponding to the current positions of each particle in the current particle swarm;
[0033] Select some test case individuals from the current test case set as the seed test cases and determine the corresponding seed particles;
[0034] Calculate the new moving speed of the seed particle based on the current position of the seed particle, the current speed, the individual best position of the seed particle, and the global pheromone highest position; and
[0035] Add the new moving speed to the current position of the seed particle to obtain the new position of the seed particle, and the new position of the seed particle corresponds to a new test case individual; wherein, the current seed test case and the new test case individual constitute the new test case set.
[0036] Optionally, the step of selecting some test case individuals from the current test case set as the seed test cases includes:
[0037] Obtain the fitness score of each test case individual in the current test case set; and
[0038] Traverse the fitness scores of each test case individual, and select the test case individuals with fitness scores greater than the threshold or select the preset number of test case individuals ranked at the front as seed test cases;
[0039] Alternatively, select some test case individuals from the current test case set based on the roulette wheel selection mechanism. When selecting some test case individuals from the current test case set based on the roulette wheel selection mechanism, calculate the selection probability of the test case individuals based on the pheromone value of the path covered by the test case individuals and the fitness scores of the test case individuals.
[0040] Optionally, after obtaining the new test case set, further include:
[0041] Obtain the fitness scores of each test case individual in the current test case set; and
[0042] Traverse the fitness scores of each test case individual, and eliminate the test case individuals with fitness scores less than the threshold.
[0043] Optionally, when the current test case set is the initial test case set, the steps of obtaining the fitness scores of each test case individual in the current test case set include:
[0044] Calculate the coverage rate of each initial test case individual for the target software metrics;
[0045] Evaluate each initial test case individual based on the threat model for security testing to obtain the defect revelation ability value; and
[0046] Calculate the weighted sum of the coverage rate of each initial test case individual for the software metrics and the defect revelation ability value as the initial fitness score of the initial test case individual;
[0047] When the current test case set is the test case set obtained after the last optimization or the new test case set obtained by the current optimization, the steps of obtaining the fitness scores of each test case individual in the current test case set include:
[0048] Run each test case individual in the current test case set and collect the corresponding running data;
[0049] Calculate the coverage rate of each test case individual for the target software metrics based on the running data of each test case individual;
[0050] Evaluate each test case individual based on the threat model for security testing to obtain the defect revelation ability value;
[0051] Query the user behavior model data and calculate the matching rate between the user behavior simulated by each test case individual and the user behavior model; and
[0052] Calculate the weighted sum of the coverage rate of each test case individual for the target software metrics, the defect revelation ability value, and the matching rate between the user behavior simulated by the test case individual and the user behavior model as the fitness score of each test case individual.
[0053] Optionally, when calculating the new moving speed of the particle, adjust the current moving speed of the particle based on one or more of safety risks, environmental impacts, user behavior matching degrees, business scenarios, coverage rate impacts, execution time impacts, and system load impacts.
[0054] Optionally, the optimization end condition is to reach the iteration number threshold or the fitness score of the current test case set reaches the threshold.
[0055] According to another aspect of the present invention, the present invention also provides a software test case generation system, including:
[0056] An initialization module configured to initialize the test environment to obtain the path exploration space of the target software; wherein, the path exploration space includes nodes and directed edges between the nodes, and two nodes with a directed relationship form a path branch, and each path branch corresponds to a path branch condition for implementing the path branch.
[0057] A user behavior model construction module configured to construct a user behavior model based on the collected user historical behavior data, wherein the user behavior model includes one or more user parameters simulating user behavior and their parameter values.
[0058] An input parameter space construction module configured to construct the input parameter space of the target software, wherein the input parameter space includes multiple input variables, and the variable value of each input variable corresponds to one or more user parameter values.
[0059] An initial set construction module configured to drive multiple ants to perform path exploration in different regions of the path exploration space respectively, and update the pheromone value of the path branch during the path exploration process; after each ant determines a path branch during the path exploration process, based on the path branch condition, select the corresponding input variable value from the input parameter space according to the user behavior model; after each ant finishes the path exploration, obtain the input variable values of all path branches corresponding to the explored path to form an initial test case individual; all initial test case individuals form an initial test case set; and
[0060] An optimization module configured to map the initial test case set to an initial particle swarm, and perform iterative optimization on the initial particle swarm until the optimized test case set meets the optimization end condition.
[0061] According to another aspect of the present invention, the present invention further provides an electronic device, which includes a processor and a memory storing computer program instructions; when the electronic device executes the computer program instructions, any of the foregoing software test case generation methods is implemented.
[0062] According to another aspect of the present invention, the present invention further provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, any of the foregoing software test case generation methods is implemented.
[0063] According to another aspect of the present invention, the present invention further provides a computer program product, including a set of computer program instructions, and when the set of computer program instructions is executed by a processor, any of the foregoing software test case generation methods is implemented.
[0064] The embodiments of the present invention improve the generation efficiency of test cases, as well as the test coverage rate, depth, and test effectiveness of test cases in complex path coverage scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Next, the preferred embodiments of the present invention will be further described in detail with reference to the drawings, where:
[0066] Figure 1 is a method flowchart for generating test cases by a method that combines the ant colony algorithm (ACO) and particle swarm optimization (PSO) according to an embodiment of the present invention;
[0067] Figure 2 is a schematic diagram of a path structure according to an embodiment of the present invention;
[0068] Figure 3 is a method flowchart for optimizing candidate test case individuals to obtain an initial test case set according to an embodiment of the present invention;
[0069] Figure 4 is a method flowchart for selecting seed particles for moving positions by using a roulette wheel selection mechanism according to an embodiment of the present invention;
[0070] Figure 5 is a schematic block diagram of the principle of a software test case generation system according to an embodiment of the present invention;
[0071] Figure 6 is a schematic diagram of the architecture of a software test system according to an embodiment of the present invention; and
[0072] Figure 7 is a schematic diagram of the hardware structure principle of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0074] In the following detailed description, reference may be made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration specific embodiments in which the invention may be practiced. In the drawings, like reference numerals describe substantially similar components in different views. The various specific embodiments of the present application have been described in sufficient detail below to enable those of ordinary skill in the art with relevant knowledge and technology to implement the technical solutions of the present application. It should be understood that other embodiments may be utilized or structural, logical, or electrical changes may be made to the embodiments of the present application.
[0075] See Figure 1 , Figure 1 is a flowchart of a method for generating test cases by using a method that combines the ant colony algorithm (ACO) and particle swarm optimization (PSO) according to an embodiment of the present invention. In this embodiment, the software test case generation method includes the following steps:
[0076] Step S101, initialize the test environment to obtain the path exploration space of the target software. Among them, the path exploration space includes nodes and directed edges between the nodes. Two nodes with a directed relationship form a path branch, and each path branch includes a path branch condition. The path branch condition includes the input operation for implementing the path branch and the corresponding input operation condition.
[0077] Step S102, construct a user behavior model and the input parameter space of the target software based on the collected user historical behavior data. Among them, the user behavior model includes one or more user parameters simulating user behavior and their parameter values; the input parameter space includes multiple input variables, and the variable value of each input variable corresponds to one or more user parameter values.
[0078] Step S103, drive multiple ants to perform path exploration in different regions of the path exploration space to obtain an initial test case set.
[0079] Step S104, map the initial test case set to an initial particle swarm. Among them, each initial test case individual corresponds to the initial position of a particle.
[0080] Step S105, move the position of the particle at a preset speed to obtain a new position of the particle, and the new position of the particle is a new test case individual.
[0081] In step S106, test case individuals with fitness scores not meeting the requirements are eliminated, and the test case individuals meeting the fitness score requirements form a new test case set.
[0082] In step S107, it is determined whether the new test case set meets the optimization end condition. If the optimization end condition is met, the current test case set is stored in step S108 and the optimization ends. If not, return to step S105.
[0083] In step S101, in one embodiment, when initializing the test environment, the dynamic call graph matrix of the target software is first obtained. It uses the matrix representation method M(t) in graph theory to describe the interaction and dependency relationships between software units. Specifically, the interaction and dependency relationships between software units are described by formula (1-1):
[0084] M(t) = M(t - 1) + ΔM(t) (1-1)
[0085] M(t) represents the call matrix of software units within time period t, M(t - 1) represents the call matrix of software units within time period t - 1, and ΔM(t) represents the change in call data within time period t. The dynamic change situation of the software during operation is represented by the matrix of the call data of software units.
[0086] Among them, software units, such as components, modules, and even more detailed functions, etc., serve as the rows and columns of the call matrix, and the matrix elements represent the call data between software units within a preset time period, such as the number of times or frequency. In one embodiment, the matrix element M(t) [i,j] represents the number of times or frequency that the i-th software unit calls (jumps / calls) the j-th software unit within time period t. For example, if "component A" and "component B" are mapped to the i-th row / column and the j-th column / row, then M(t) [i,j] is the number of times that component A calls component B within time period t. When further refining "component" to the "function" level, M(t) [f1,f2] represents the number of times that function f1 calls function f2 within time period t.
[0087] Over time accumulation, M(t), M(t + 1) …… can be merged or superimposed to obtain the overall call matrix M_total. Thus, it can be seen that the call graph matrix not only includes the structural composition of the software, but also can obtain whether the call relationships of certain components / modules / functions are covered, and can also obtain their call frequencies and control flows.
[0088] By parsing the dynamic call graph matrix to obtain two software units corresponding to each matrix element; using the software units as nodes and the call relationships between the software units as edges to construct a call flow graph, the call flow graph includes nodes and edges with directions between the nodes, where two nodes connected by an edge with a direction relationship form a path branch. According to the call relationships between the parsed software units, obtain the input operations and corresponding input parameter conditions for implementing the path branch, hereinafter referred to as branch conditions.
[0089] Specifically, query the elements greater than 0 in the current call graph matrix. If the count of a certain row and a certain column in M(t) is greater than 0, it means that the two software units represented by the corresponding row and column have indeed been called during the test. Therefore, it can be obtained that the two software units have "triggered" the corresponding function call actions during the test. When modeling the internal structure of the software, the two software units that have "triggered" the call actions are used as nodes, and the directed call relationship of "software unit i calls software unit j" is used as the "edge" connecting software unit i and software unit j, thus forming a path branch. According to this method, concatenate all the call relationships where the matrix elements (M(t) [i,j] ) are not 0, and all the actual paths taken during the test are obtained, which constitute the call flow graph, where multiple consecutive software units and their call relationships form a dynamic call chain. If it does not appear (or is 0) in the matrix, it means that the edge of the call relationship between the corresponding two software units has not been tested and executed. Therefore, based on the call flow graph, the path exploration space for the ants in the present invention to explore paths can be obtained.
[0090] In addition, through the aforementioned dynamic call graph matrix, it is also possible to obtain the coverage of modules / functions and the coverage of paths or path branches during the test process.
[0091] Since there is a specific mapping relationship between software units and code, the code file or code line corresponding to the software unit can be obtained. Based on the matrix element M(t)[X,Y]≠0, by combining a code-level coverage tool, it can be confirmed that "X calls a specific code branch or code line in Y".
[0092] Since after the test cases are generated, it is necessary to verify their coverage rate and effectiveness in various environments. Therefore, further, when initializing the test environment, various software running environment simulation configuration parameters are also obtained, which are used to simulate the running states of different environments and provide diverse test scenarios for the test. Among them, the running environment simulation configuration parameters include three types of parameters: the operating system version (abbreviated as O), the network condition (abbreviated as N), and the hardware configuration data (abbreviated as H). Each type of running environment simulation configuration parameter can include one or more running environment simulation configuration parameters. In one embodiment, each running environment simulation configuration parameter is set as an environmental factor ei Since the software running environment (or the system environment of the target software) has different impacts on different test cases, to measure this impact, an environment fitness function E(C, v) is defined. By measuring the fitness of a test case individual v to the environment under a specific system configuration, the impact of the software running environment on the test case individual v is quantified. Among them, C represents the set of software running environment simulation configuration parameters, C = {O, N, H}, O represents the operating system version, N represents the network condition, and H represents the hardware configuration data. The present invention adopts the set E of environmental factors e i to represent a specific system simulation configuration. Each environmental factor e n can correspond to a set of one or more running environment simulation configuration parameters. Each environmental factor e i is assigned a corresponding weight w i . Based on the comprehensive impact I of the fitness function E(C, v) on the environment configuration (abbreviated as the environmental impact value), it is calculated by the following formula (1-2): i where f(e
[0093]
[0094] ) is the impact value of the environmental factor e i . Among them, the calculation method of the impact value f(e i ) of the environmental factor e i is a conventional technique in the field of software testing and will not be elaborated here. i
[0095] In addition, security is usually also an important aspect in software testing. In one embodiment, when a software unit such as a software unit or a code segment is marked with a security risk attribute tag, the security risk attribute tag of the software unit is obtained during initialization, so that security testing can be performed. To evaluate the ability in terms of security testing during the test case generation process, the present invention further obtains a threat model for security testing during initialization. The threat model is a security module for analyzing and processing certain security vulnerabilities or security threats, such as the Common Vulnerability Scoring System (CVSS for short). The threat model can determine whether a vulnerability is detected based on the triggering conditions corresponding to the security vulnerability, and quantify the risk of the vulnerability or rate it when the vulnerability is detected. For example, the threat model determines whether a test case triggers a vulnerability through symbolic execution (such as SQL injection detection returns 1 / 0), and quantifies the severity s of the security vulnerability and the likelihood a of the vulnerability being exploited through the risk scoring function R(s, a) shown in formula (1-3).
[0096] R(s,a) = s × a (1-3)
[0097] Among them, the vulnerability severity s ∈ [0, 10], and the possibility a of the vulnerability being exploited (calculated by Monte Carlo simulation) a ∈ [0, 1]. The product of the two is the comprehensive quantitative risk value. For example, for a certain Cross-Site Scripting (XSS) vulnerability, R(s,a) = 0.7 × 8.5 + 0.3 × (probability 0.9) ≈ 7.2, which is determined to be a high risk according to the rating standard.
[0098] The types of security vulnerabilities in the present invention include SQL injection, cross-site scripting attacks, buffer overflows, etc.
[0099] In step S102, each piece of data in the user historical behavior data includes at least user identity information, time information, user operation information, and input information. These parameters representing various types of information are collectively referred to as user parameters. According to the category, user parameters include user operation parameters, input parameters corresponding to user operations, user identity parameters, and environment parameters. The parameter values of user operation parameters can be various user operations, such as clicking, selecting, inputting, etc.; the parameter values of input parameters are, for example, the specific content corresponding to the user operation, such as the confirmation key information corresponding to the click confirmation key operation, the product information corresponding to the selection operation, the input content corresponding to the input operation, etc.; the parameter values of user identity parameters are, for example, information such as user names and passwords, and the environment parameters are, for example, one or more types of information among the operating system version, network conditions, and hardware configuration.
[0100] For example, for an online shopping platform system, a piece of user historical behavior data collected is as follows (time information is omitted):
[0101] User operations and corresponding input information:
[0102] Search keyword - smart phone;
[0103] Select product category - electronics;
[0104] Select product price range - (1000 - 2000) yuan;
[0105] User account information:
[0106] User name - user123;
[0107] Password - pass456;
[0108] Environmental parameters:
[0109] User's device type - Android phone;
[0110] Operating System Version - Android 12;
[0111] Network Conditions - (Wi-Fi, bandwidth 10Mbps).
[0112] The parameter values of the user operation parameters extracted from the above data are "search keywords", "select product category", "select product price range", etc. The corresponding input parameters are, for example, specific keywords (such as smartphones), specific product categories (such as electronic products), etc. The parameter values of the environmental parameters include the user device type ("Android phone"), the user device operating system version ("Android 12"), the network conditions used by the user ("Wi-Fi, bandwidth 10Mbps"), etc. The parameter value of the user identity parameter is the specific account information, that is, the username "user123" and the password "pass456".
[0113] The user behavior model in the present invention is various data or combinations that can simulate user behavior. For example, the distribution probability and usage frequency of a specific type of user parameter value, a sequence composed of multiple user parameter values representing different user behaviors; the association relationship between different categories and / or the same category of user parameter values representing user behavior patterns.
[0114] Taking the user operation parameter as an example, the parameter value of the user operation parameter is a specific type of user operation, and the distribution probability of each user operation is calculated based on formula (1 - 4).
[0115]
[0116] Among them, x is a specific type of user operation, and n x is the number of user operations x extracted from the historical data of a preset time period (such as one week, three days, one month, etc.), and N is the total number of all user operations of this user within the same time period. User operation x is, for example, "search keywords", "select product category" or "select product price range" in the foregoing embodiments, etc. By calculating the distribution probability of each specific user operation, it is possible to distinguish whether a user operation conforms to the user's habitual operation. Similarly, the distribution probabilities of various other types of user parameters can also be obtained, such as the distribution probabilities of various specific input parameter values, the distribution probabilities of various environmental parameter values, etc.
[0117] The usage frequency of user parameters can also reflect the corresponding user behavior pattern. For example, by counting the button click frequency of the same user within their respective statistical time periods at different times, the user's habit of clicking the button can be obtained, such as clicking 5 times per minute. Another example is that according to the average number of product pages viewed by different users during a search process, the user behavior pattern of viewing 10 pages each time can be obtained.
[0118] A sequence composed of multiple user parameter values can also represent corresponding user behaviors. For example, for a shopping platform, one user behavior model obtained is: Login - Click on Product 1 - Click on Product 2 - Click on Product 3 - Click on Product 4 - Click on Product 5 - Click on the purchase button on the product page... Another user behavior model obtained is: Login - Select product category and search for product keywords - Click on Product 1 - Click on Product 2 - Click on Add to Cart - Click on Product 3 -... - Click on the purchase button.... The sequences composed of multiple categories of user parameter values representing the above two different user behaviors are respectively a kind of user behavior model, and these two user behavior models represent different user behavior patterns.
[0119] For another example, when the user behavior model is a sequence of specific user parameter values such as the operation of clicking on a product - specific product information - the operating system in the environmental parameters, the user behavior models composed of different specific parameter values represent different user behavior patterns.
[0120] The distribution probability of each of the above user parameter values, the usage frequency of each user parameter value, and multiple user parameter value sequences are collectively referred to as the user behavior model.
[0121] The input parameter space of the target software includes multiple input variables. The input variables correspond to user parameters, and the variable value of each input variable includes one or corresponding multiple user parameter values. In one embodiment, the input parameter space of the target software is represented by a set of variables: X = {x1, x2,..., x n}; X represents the input parameter space, and the variable x i represents the i-th input variable. The input variable can be an input operation or input parameter that the target software can accept. Here, the input operation corresponds to the user operation in the user behavior model, and the input parameter value is the specific information, data, etc. that the user can input, corresponding to the input parameter value corresponding to the user operation. For example, the input operation can be a login operation, inputting a search keyword, clicking on a relevant link, and the input parameter value can be the content that can be input for an input operation, such as a specific keyword, a specific link, etc. The input variable can also be a user identity parameter, such as a username and password. The input variable can also be an environmental parameter, such as user device parameters, the user device operating system, network configuration or conditions, etc. The specific variable values of these input variables form the complete space of the input parameters of the target software. For example, the input variable x1 represents the username, and its value can be "user123" or "admin", etc.; the input variable x2 represents the user operating system, and the parameter value can be "Windows10", "Android12", etc., the input variable x3 represents the login operation, the input variable x4 represents the operation of inputting a search keyword, the input variable x5 represents the keyword, and its value is one or more keywords in the keyword table, etc.
[0122] In the present invention, a test case individual refers to a specific instance for a complete test task, including multiple constituent parameters, and is the smallest unit in the test case set. In the following description, in order to highlight the relationship with the test case set, a specific instance for a complete test task is referred to as a test case individual. When not emphasizing the relationship with the test case set, a specific instance for a complete test task is referred to as a test case.
[0123] To achieve a complete test task, there are multiple categories of constituent parameters for a test case, and certain conditions need to be met. For example, the constituent parameters of a test case individual at least include specific input operations and corresponding input parameter values. The input operations can be one or multiple and consecutive, that is, an input operation sequence. Referring to the paths and path branches in the call flow graph, an input operation and its corresponding input parameter value in a test case can implement a transition from one node to another node, that is, implement a path branch. When a test case includes multiple consecutive input operations, multiple consecutive path branches can be implemented, that is, form a path.
[0124] In addition, the constituent parameters of a test case can also include environment parameters, and user identity parameters are required when necessary. The environment parameters are used to define the running scenario, and the user identity parameters are a special type of input parameter.
[0125] In step S103, before driving multiple ants to perform path exploration in different regions of the path exploration space, it is first necessary to set ant colony parameters, such as the initial pheromone value of each path branch, the pheromone evaporation rate ρ, the calculation method of the pheromone intensity constant Q, the number of ants, the number N of initial test case individuals to be obtained, the exploration termination condition, etc. Among them, according to the number of ants and the number N of test case individuals to be obtained, the exploration times of each ant are configured.
[0126] In one embodiment, the initial pheromone value is set for the paths in the path exploration space based on the user behavior model. For example, first, the corresponding user operation is determined based on the matching relationship between the paths in the path exploration space and the user operations. For example, each path branch formed by each node (such as a page or a state) among multiple nodes in the page browsing path corresponds to a user operation. Then, the user behavior model is queried based on the user operation, and the initial pheromone value of the path branch is set according to the distribution probability of the user behavior operations in the user behavior model. If the distribution probability is large, the initial pheromone value is large; if the distribution probability is small, the initial pheromone value is small. Thus, the ants can be enabled to select paths that are more matched with the user behavior during path exploration.
[0127] The process of path exploration for each ant specifically includes: First, determine the starting node for each ant, and then calculate the probabilities from the starting node to all possible next nodes according to the probability formula. Among them, the probability formula is shown as formula (1-5) below:
[0128]
[0129] τij: The pheromone concentration from node i to node j, which is calculated according to formula (1-6). is the probability from node i to node j, and l is one of multiple candidate nodes, that is, one of all allowed nodes.
[0130]
[0131] Among them, t is the time parameter in the ant colony algorithm. Corresponding to the present invention, it is the iteration number when the ant conducts path exploration; in this embodiment, when generating the initial test case set, it terminates when obtaining N test case individuals with preset conditions that are met. Before termination, the ant needs to conduct path exploration multiple times, that is, multiple iterations. τ ij (t + 1) is the pheromone value from node i to node j (referred to as path L i,j ) for the (t + 1)-th exploration after being updated through the t-th exploration, and τ ij (t) represents the pheromone concentration of path L i,j at the t-th exploration. ρ is the evaporation rate of the pheromone, is the pheromone value released (left) by the k-th ant when exploring path L i,j , and it is also the pheromone increment left after being explored by one ant on path L i,j . Q is the pheromone intensity constant, and L k is the path length of the k-th ant.
[0132] η ij is the heuristic factor from node i to node j. The heuristic factor includes at least the weight of user behavior simulation matching degree. For example, η ij = α · user operation probability. The user operation probability therein is taken as a specific value of the user behavior simulation matching degree. In addition, in order to give consideration to guiding the ant to choose a path with higher security risk, a security risk weight can also be added. For example, η ij = α · user operation probability + β · security risk value. Of course, other weights can also be set according to the needs of the test, such as test execution time weight, running environment weight, path length weight, business importance weight, high-frequency feature weight, etc. By introducing other concerned factors when selecting the next node, it can lead to the concerned path branches during the path exploration process. Specifically, it can be flexibly used according to the nature of the target software in actual applications.
[0133] α and β are the weights of the control pheromone and the heuristic factor respectively. Similarly, τ il is the pheromone from node i to node l, and η il is the heuristic factor from node i to node l.
[0134] After calculating the selection probabilities of all candidate nodes, the node with the maximum selection probability is selected as the next node, thereby obtaining a path branch. After each ant determines a path branch during the path exploration process, it selects the corresponding input variable from the input parameter space based on the path branch condition, and then selects the corresponding input operation and input parameter value from the input parameter space according to the user behavior model as the corresponding input variable value. And so on until the termination condition of the path exploration is reached, such as reaching the preset path length (such as at most 10 operation steps), or entering the termination state (such as triggering a crash, completing the core process), or covering the target branch or meeting the coverage threshold, etc.
[0135] The complete search trajectory of each ant is the explored path, and the corresponding input variable values selected from the input parameter space corresponding to all path branches constitute a test case individual.
[0136] The user behavior model provided by the present invention has various forms, such as the aforementioned distribution probability, usage frequency, specific operation sequence, etc. Therefore, after each ant determines a path branch during the path exploration process, it first determines the form of the user behavior model to be used based on all input variables corresponding to the path branch condition. When only one input variable is involved, it selects the input operation or input parameter value with the maximum distribution probability that can be used as the input variable value. For example, when the path is from the login page to the user's personal homepage, since a login variable is required, the corresponding input operation is the login operation, and the corresponding input parameters include two input parameters, namely the username and the password. There are multiple usernames and multiple passwords in the input space. When there are multiple input parameter values, the username and password with the maximum distribution probability are selected as the input parameter values, thereby obtaining a test case individual: [Click the login button; Input parameters: Username "user123" + Password "pass456"].
[0137] When there are multiple input operations corresponding to the input variable, select the input operation with the highest distribution probability or the highest usage frequency in the user behavior model. Figure 2 is a schematic diagram of the path structure according to an embodiment of the present invention. For Figure 2For node i, the corresponding input operations can be "click on product operation a" and "search operation b". Referring to the user behavior model, the distribution probability of "search operation b" is the highest, so "search operation b" is selected. Another example is when the determined path is from node i through node j1, node j3 and then to node j, including two consecutive input operations "search operation" and "click operation". According to the user behavior model: [search for "smartphone", select the "electronic product" category, browse products with a price range of "1000 - 2000 yuan"], the specific input operation values are determined, and the obtained test case individual is: [search for "smartphone"; select the "electronic product" category, price range "1000 - 2000 yuan"; click on "xxx phone"].
[0138] Since the user behavior model also provides the usage frequency of input parameters, further, the input parameters can be selected according to the usage frequency of input parameters. For example, according to the user behavior model, when the usage frequency of the input parameter that the user visits 10 pages on average each time when browsing the product page is obtained, add the browsing operation steps of 10 product pages in the generated test case. For example, the generated test case simulates that the user uses a "wired network" on a computer with the "Windows10" system to access the shopping platform and randomly clicks on 10 product pages.
[0139] Another example Figure 2 For the path shown, based on the user behavior model, when the frequency of clicking on products when the user browses products is 5 each time, 4 randomly selected products can be added on this basis, and the obtained test case individual is: [search for "smartphone"; select the "electronic product" category, price range "1000 - 2000 yuan"; click on "xxx phone 1"; click on "xxx phone 2"; click on "xxx phone 3"; click on "xxx phone 4"; click on "xxx phone 5"].
[0140] Finally, in order to specify the running scenario of the currently generated test case individual, usually the previously set running environment parameters can also be added, so as to obtain a test case with complete composition parameters, such as [username: user123; password: pass456; search for "smartphone"; select the "electronic product" category; select the price range "1000 - 2000 yuan"; click on "xxx phone 2"; click on "xxx phone 3"; click on "xxx phone 4"; click on "xxx phone 5"; operating system version: Windows10; network configuration: Wi-Fi, latency 50ms]. Of course, other more detailed running environment parameters such as hardware conditions can also be included.
[0141] The running environment parameters include three types of parameters: the operating system version (abbreviated as O), network conditions (abbreviated as N), and hardware configuration data (abbreviated as H). Each type of running environment parameter can include one or more running environment parameters. In order to enable test cases to test the target software in various environments, the present invention deploys and simulates the running states of different environments according to the software running environment parameters to provide diverse test scenarios for testing. In one embodiment, each running environment parameter is set as an environmental factor e i , since the software running environment (or the system environment of the target software) has different impacts on different test cases. To measure this impact, an environmental fitness function E(C, v) is defined. The environmental fitness of a test case individual v in a specific system configuration is measured through the environmental fitness function, so as to quantify the impact of the software running environment on the test case individual v. Among them, C represents the set of software running environment simulation configuration parameters, C = {O, N, H}, O represents the operating system version, N represents network conditions, and H represents hardware configuration data. The present invention adopts the set E of environmental factors e i = {e1, e2,..., e n} to represent a specific system simulation configuration. Each environmental factor e i can correspond to a set of one or more running environment simulation configuration parameters. Each environmental factor e i is assigned a corresponding weight w i . Based on the comprehensive impact I of the fitness function E(C, v) on the environmental configuration (abbreviated as the environmental impact value), it is calculated by the following formula (1-8):
[0142]
[0143] Among them, f(e i ) is the impact value of the environmental factor e i .
[0144] During the process of ants exploring paths, the pheromone values of path branches are updated in real time. Specifically, the following update methods are included:
[0145] (1) Adjusting pheromone update according to defect discovery priority: In actual application scenarios, some functional modules or code segments are more likely to have defects. Therefore, ants can be given priority to explore "high-risk areas" at this time. For example, before testing, the risk levels of functional modules or code segments are marked according to data analysis. During the path exploration of ants, it is judged whether the current explored path branch is marked with a security risk attribute mark. When the current explored path branch is marked with a security risk attribute mark, the pheromone increment of the path branch marked with the security risk attribute mark is increased. Moreover, the pheromone increment is adjusted according to the risk level. The higher the risk level, the greater the pheromone increment value. As a result, the probability of the path corresponding to the functional module or code segment being selected during path exploration is increased, and thus test cases that can reveal defects in potential high-risk areas are generated.
[0146] (2) Path optimization scenarios: Since different software systems have different test path requirements, according to the path optimization requirements of the target software and specific test scenarios, pheromone update rules or update formulas are defined for each specific scenario. The scenarios are, for example: user login scenario, data input verification scenario, API interface call scenario. During path exploration, it is judged which scenario the current path exploration space area corresponds to, and the pheromone value of the path branch is adjusted based on a manner matching the scenario.
[0147] (3) Path coverage scenarios: During the process of driving multiple ants to explore paths in different areas of the path exploration space, the coverage rate of the explored paths to the path exploration space or a preset local path exploration space is calculated. Based on the corresponding relationship where the coverage rate is negatively correlated with the pheromone, the pheromone value of the path branch in the path exploration space is adjusted, which can make the higher the path coverage rate, the lower the pheromone concentration, thus guiding ants to select paths with less coverage as test case individuals.
[0148] (4) Performance testing scenarios: When performing performance / stress testing, one often cares about the time that each test path may consume, the resource occupation during concurrent execution (CPU, memory, network bandwidth, etc.), the upper limit of test execution time or throughput metrics, etc. Therefore, in one embodiment, a calculation weight related to "execution time (or resource consumption) T" is added to the foregoing formulas (1-7). For example, the parameter Lk is replaced with α·path length + β·execution time + γ·resource overhead, or the execution time / resource is used as a penalty term Rk and introduced into the calculation formula of so that when ants explore paths, they can tend to select paths that are short, fast, and resource-saving, realizing the minimization of resource consumption.
[0149] (5) Adjustment based on environmental fitness: In actual test scenarios, the test performance of different system configurations is crucial. Therefore, in one embodiment, environmental factors are incorporated into the calculation process of pheromone concentration to generate test cases that are more in line with real usage scenarios. For example, the pheromone increment is set according to the magnitude of the influence value calculated based on the operating environment parameters of the current test environment.
[0150] (6) Optimization adjustment based on user behavior: According to the frequency and habits of actual user operations (such as login, query, submit form), the pheromone concentration is increased for common paths, and test cases for these common paths are preferentially generated to improve the practical value of the test cases. For example, after initializing the test environment to obtain the path exploration space of the target software and the user behavior model, the paths in the path exploration space of the target software are queried through the user behavior model, and the paths corresponding to the operations with high usage frequency or high distribution probability in the user behavior model are determined. A larger pheromone increment value is set for these paths, so that the pheromone value of these paths can be increased during information update, thereby achieving the purpose of preferentially generating test case individuals for these paths, and thus improving the practical value of the test case individuals.
[0151] To achieve the above adjustment of the pheromone value, the calculation weights of the foregoing various factors can be increased in the pheromone intensity constant Q or the path length L k so as to adjust the pheromone value left by the ants in the foregoing various situations to achieve the purpose of guiding other ants to explore towards a predetermined area.
[0152] Furthermore, in order to make the initialized test cases close to real user behavior and real operation scenarios, better face the final test goal, and make subsequent iterative optimization easier and convergence faster, in another embodiment of the present invention, when constructing the initial test case set, a batch of initially generated test case individuals are used as candidate initial test case individuals, and a maximum value problem function is constructed based on the user behavior model and environmental fitness. By adjusting the candidate initial test case individuals, or adjusting the user behavior model and / or operating environment parameters of the candidate initial test case individuals, the maximum value problem function is converged to the maximum value; the candidate initial test case individuals when converging to the maximum value are determined as the initial test case individuals.
[0153] In one embodiment, the maximum value problem function is as follows:
[0154]
[0155] where P(x i ) is a specific input variable value x in the test case iThe distribution probability, that is, the probability of being used, is calculated according to formula (1-4). n is the total number of various input variable values in all test cases. E(C,v) is the environmental impact value I of the operating environment parameters in each test case on the test case, which is calculated according to formula (1-8). m is the total number of operating environment parameters in the current test case set V. α and β are trade-off coefficients used to balance the influence of the user behavior model and environmental configuration on the initialization process.
[0156] The present invention measures the degree to which the overall test case set is close to the real user behavior and real operating scenario based on the user behavior model and the simulated configuration parameters of the operating environment, and optimizes the individual test case v according to the measurement result.
[0157] See Figure 3 , Figure 3 It is a flowchart of a method for optimizing an individual test case to obtain an initial test case set according to an embodiment of the present invention.
[0158] Step S1031, obtain a candidate test case individual from the current test case set V as the target test case individual.
[0159] Step S1032, calculate the user behavior simulation score of the target test case individual based on the user behavior model used when generating the target test case individual. In this embodiment, that is, calculate the distribution probability P(x i ) of each input variable value in the target test case individual. Specifically, for each input variable value, its distribution probability is calculated using formula (1-4).
[0160] Step S1033, calculate the environmental fitness score of the target test case individual based on the operating environment parameters of the target test case individual. In this embodiment, that is, calculate the influence value of the simulated configuration of the operating environment in the target test case on the target test case. Specifically, take each simulated configuration parameter of the operating environment in the target test case as an environmental factor e, and calculate the environmental impact value I of the current operating environment parameter on the target test case individual using formula (1-8).
[0161] Step S1034, determine whether there are still candidate test case individuals in the test case set V. If so, return to step S1031. If there are no candidate test case individuals in the test case set V, then execute step S1035.
[0162] Step S1035, calculate the weighted sum of the user behavior simulation scores and environmental fitness scores of all candidate test case individuals in the test case set V. In this embodiment, calculate the weighted sum of the probability sum of all input variable values of the current all test cases and the environmental impact value I according to formula (1-9).
[0163] Step S1036: Adjust the running environment parameters of the candidate test case individuals and / or the user behavior model used to generate the test case individuals to obtain new candidate test case individuals. For example, randomly change the input variable values of one or more candidate test cases in the test case set V in the direction of increasing the probability and / or randomly change the running environment parameters of one or more candidate test cases in the test case set V in the way of increasing the environmental impact value I, and use them as new candidate target test case individuals. For example, applying the genetic algorithm, taking each candidate test case individual as a chromosome in the genetic algorithm, and each input variable value and each running environment configuration parameter in the candidate test case individual as a gene in the genetic algorithm respectively, and changing the input variable values and / or running environment parameters in a candidate test case individual through gene crossover and mutation operations to obtain new candidate test case individuals.
[0164] Step S1037: Calculate the user behavior simulation score and the environmental fitness score of the new candidate test case individuals. For example, calculate the probability P(x i ) that each input variable value in the target test case is used and the environmental impact value I.
[0165] Step S1038: Calculate the weighted sum of the user behavior simulation scores and the environmental fitness scores of all candidate test case individuals in the current test case set V.
[0166] Step S1039: Determine whether the weighted sum of the current test case set V has converged to the maximum value. If it has converged to the maximum value, end the process; otherwise, return to Step S1036. Herein, "converging to the maximum value" means that the weighted sum of the current test case set V no longer increases, that is, by comparing the size of the current weighted sum with the weighted sum calculated last time, when the weighted sum no longer increases, it is confirmed that it has converged to the maximum value. The test case set V at this time is the initial test case set.
[0167] Furthermore, after obtaining the initial test case set, evaluate the test effectiveness of each test case individual. In one embodiment, calculate the fitness score of each test case individual through the following fitness function expression (1-10). In this embodiment, evaluate the test effectiveness through the software metric coverage ability and the ability to discover potential defects.
[0168] f(v) = γ·Coverage (v)+δ·Defects (v) (1-10)
[0169] Coverage(v) is the software metric coverage rate for an individual test case v, usually a percentage, ranging from (0 - 1). In one embodiment, the software metric is, for example, a path. When specifically calculating, the number of path branches included is determined through the path corresponding to the test case individual v, and the proportion of the number of path branches of the test case individual v to the total number is calculated. The path branch proportion is used as the coverage rate Coverage(v). The software metric can also be, for example, lines of code. Based on the test path corresponding to the test case individual v, the relevant software structure is determined, and based on the correspondence between the software structure and the code, the number of lines of code is determined. The proportion of the number of lines of code corresponding to the test case individual v to the total number of lines of code is calculated, and the line of code proportion is used as the coverage rate Coverage(v). Or, the software metric can also be, for example, a functional module of the software. The ratio of the number of functional modules covered by the test case individual v to the total number of functional modules is calculated, and the functional module ratio is used as the coverage rate Coverage(v). For example, if the target software system has 100 functional modules and the test case individual v covers 20 of them, then Coverage(v) = 20%.
[0170] Defects(v) is the number of potential defects that a single test case individual v can reveal, representing the defect detection ability of a single test case individual v. The defect detection ability Defects(v) is a positive integer, and its specific range depends on the test scenario and system complexity. Its calculation method is, for example: using the aforementioned security threat model to analyze the potential defect triggering ability of the test case individual v for the target software through symbolic execution, such as predicting the types and quantities of vulnerabilities that can be exposed.
[0171] γ and δ are weight coefficients used to balance the contributions of the coverage rate and defect detection ability to fitness.
[0172] In this embodiment, the weight coefficients can be set according to the needs of testers. If the coverage rate is given priority, γ > δ can be set. If defect discovery is emphasized, δ > γ is set.
[0173] For example, in a medium - sized or small software system, for a test case, its Coverage(v) may be between 0.1 and 0.5 (i.e., 10% - 50%), and its Defects(v) may be between 0 and 5 (the number of detected defects is limited).
[0174] For example, set γ = 0.6 and δ = 0.4. If the coverage rate of the test case individual v is 0.3 (30%) and 2 defects are detected, then the fitness score is: f(v) = 0.6·0.3 + 0.4·2 = 0.18 + 0.8 = 0.98. This score indicates that the test effectiveness performance of this individual is very good and it is suitable to be retained in the current set for further optimization.
[0175] In a large and complex software system, for a test case individual v, Coverage(v) may be between 0.01 and 0.3. Defects(v) may be between 0 and 10 (it is easier to detect defects).
[0176] For example, set γ = 0.7 and δ = 0.3. If the coverage rate of the test case individual v is 0.1 and 5 defects are detected, then the fitness score is: f(v) = 0.7·0.1 + 0.3·5 = 0.07 + 1.5 = 1.57. This score indicates that the test case individual v is relatively effective.
[0177] The score calculated for each test case individual through expression (1-6) is used as the initial fitness score of the test case individual. When the score calculated based on the fitness function expression (1-6) is higher, it means that the test case individual v performs better in both aspects of the software function coverage degree and the ability to discover potential defects, that is, the test effectiveness is better, has higher potential value, and can more effectively discover software defects or verify software functions.
[0178] In one embodiment, after the weighted sum of the current test case set V converges, the candidate initial test case individuals whose initial fitness scores meet the requirements are determined as the initial test case individuals. For example, sort the test case individuals in the current test case set V according to the initial fitness scores, and use a certain number of test case individuals ranked at the front as the initial test case individuals, so as to obtain the final initial test case set. Or, use the test case individuals whose initial fitness scores are greater than the threshold as the initial test case individuals, so as to obtain the final initial test case set.
[0179] In addition, referring to the evaluation of the test effectiveness of each test case, in another embodiment, the test effectiveness of the test case set V can also be evaluated as a whole. Specifically, see the fitness function expression (1-11):
[0180] f(V) = γ·Coverage(V) + δ·Defects(V) (1-11)
[0181] Among them, Coverage(V) represents the software metric coverage rate of the test case set V, reflecting the coverage of the test case set V for the target software. Defects(V) represents the number of potential defects that can be revealed through the test case set V, reflecting the effectiveness of the defect detection ability of the test case set. γ and δ are trade-off coefficients used to balance the contributions of the coverage rate and the defect detection ability to the fitness.
[0182] The initial test case set of the present invention has a certain degree of diversity, which can initially cover key or common use case scenarios, and to a certain extent, targets or takes into account goals such as function coverage, security testing, and user behavior coverage. Therefore, in subsequent iterative optimizations, combined with fitness evaluations (such as coverage rate, security vulnerability discovery ability, user behavior matching degree, etc.), these initial test case individuals are continuously optimized, gradually approaching or continuously improving the degree of satisfaction with the test goals, avoiding a sudden expansion of the search space during optimization, thus reducing the search difficulty and also improving the optimization efficiency.
[0183] In step S104, the initial test case set is mapped to an initial particle swarm. Among them, each initial test case individual corresponds to the initial position of a particle. When mapping, first map the test case to a multi-dimensional position vector of the particle, and each dimension corresponds to an input variable of the test case individual. Among them, the position of the particle corresponds to a specific test case individual. For example, the dimension parameters of a particle x1 are configured as [input field length, number of concurrent requests, timeout threshold], and the dimension parameters of a particle x2 are configured as [username length, password complexity, network latency, operating system].
[0184] Then, each particle is encoded, that is, the specific dimension parameter values are encoded into numbers. If the dimension parameter value is a continuous numerical type, such as network latency, input value, etc., the continuous parameter value is directly used. If the dimension parameter value is a discrete numerical type, such as operating system type, operation sequence, it is encoded into discrete values. When encoding into discrete values, integer encoding can be performed. For example, for the operating system type, different numbers are used to represent different operating systems, such as: 0 = Windows, 1 = macOS, 2 = Linux, 3 = Android. Multi-segment encoding can also be performed. For example, the operation sequence dimension is divided into multiple sub-dimensions, and each sub-dimension represents the selection of an operation step.
[0185] For example, for a particle x2 with a specific dimension parameter configuration of: [username: User123; password: 456789; network latency: 200ms; operating system: Android], according to its position encoding rule, the username length is 6 characters, the password complexity is a weak password, and the corresponding encoding is 0. Among them, the encoding of a medium-strength password pair is 1, and the encoding of a strong password pair is 2. The network latency is the corresponding number 200, and the encoding of the operating system is 3. Therefore, the particle x2 [username: User123; password: 456789; network latency: 200ms; operating system: Android] is encoded as [6; 0; 200; 3], which is the current position of the particle x2.
[0186] Then, determine its position range, speed range, and initial speed according to the dimension parameter type. The position range is, for example, the value range of each dimension parameter. For example, for the first dimension parameter "username length" of particle x2, its value range is 1 - 20. The speed of each particle is a vector, and its dimension corresponds one-to-one with the dimension of the particle's position vector (i.e., the input variable values in the test case individual). Each component of the speed represents the change amount of the input variable value corresponding to the dimension. When the data type of the dimension parameter is continuous numerical, such as input value, network delay, etc., the speed component is a real number, representing the increase or decrease amplitude of the dimension parameter value. When the range of the input value is [0, 100], the speed range can be set to ±10%, so the maximum adjustment amplitude of the speed component in each iteration is ±10. When the data type of the dimension parameter is discrete data, such as operating system type, operation steps, etc., the speed component needs to be converted into the probability of discrete selection or adjustment strategy. For example, for the operating system type [0 = Windows, 1 = macOS, 2 = Linux, 3 = Android], the Softmax function or roulette wheel selection method is used to reduce the probability of selecting the current operating system (such as macOS) and increase the probability of adjacent options (such as Windows). Another example is to use the integer speed conversion method to round the speed represented by a decimal to an integer, so as to select the parameter value corresponding to the integer. For example, when v = 0.7, it is converted to v = 1, so as to select the corresponding macOS.
[0187] In step S105, in one embodiment, select some particles from the current particle swarm as seed particles, and move the seed particles with their respective speeds to obtain the new positions of the seed particles. However, it can be known that each particle in the current particle swarm can also be moved to obtain the new positions of the particles. The particle swarm for the first optimization is the initial particle swarm, and in each subsequent optimization, the current particle swarm is the particle swarm obtained after optimization. In each optimization process, a certain number of particles can be randomly selected, or the roulette wheel selection mechanism can be used to select particles.
[0188] See Figure 4 , Figure 4 is a flowchart of a method for selecting particles for moving positions using the roulette wheel selection mechanism according to an embodiment of the present invention. In the current particle swarm, the current position of each particle corresponds to a specific test case individual, and the test case individuals corresponding to the current positions of all particles in the particle swarm constitute the current test case set. Selecting particles is also selecting the corresponding test case individuals. Specifically, the method for selecting particles for moving positions includes the following steps:
[0189] Step S201, obtain the fitness scores of the test case individuals corresponding to the current positions of each particle in the current particle swarm.
[0190] Step S202, calculate the selection probability p select (v) of the test case individual corresponding to the current position of each particle.
[0191] Step S203, calculate the cumulative probability P select (v) of the test case individual corresponding to the current position of each particle.
[0192] Step S204, generate a random number, and select the corresponding test case individual based on the position where the random number is located.
[0193] Step S205, determine whether the number of selected test case individuals has reached the preset quantity. If it has reached the preset quantity, end. If it has not reached the preset quantity, return to Step S204.
[0194] Among them, in Step S201, in the first optimization, the fitness score of the test case individual corresponding to the current position of each particle is the initial fitness score, which is calculated based on formula (1-10). In the subsequent optimization process, it is the fitness score calculated by applying the following formula (1-12), which reflects the quality of the test case individual v from three aspects: coverage rate, security testing ability, and user behavior matching degree.
[0195]
[0196] Among them, covered(i,v), detect(d,v), and match(u,v) are respectively indicator functions. The three terms in (1-8) respectively correspond to the coverage criterion, security criterion, and user behavior matching criterion during evaluation, and are respectively used to judge the software metrics covered by the test case individual v (such as lines of code, functions, modules, and path coverage and conditions), the ability to detect security threats or vulnerabilities, and the matching degree with user behavior. λ1, λ2, and λ3 are respectively the weights for the three criteria.
[0197] Specifically, c is a coverage criterion, C is the set of coverage criteria, and the set of coverage criteria C includes various coverage criteria for lines of code, functions, modules, and path coverage and condition coverage. i is a coverage unit, such as a line of code or a code item, function, branch, module, path, or condition, etc. I c is the set of coverage units for the coverage criterion c, w c is the weight for the coverage criterion c, |I c | is the total number of coverage units i, and covered(i,v) is an indicator function. If the test case individual v covers the coverage unit i, it is 1, otherwise it is 0.
[0198] s is the type of security test, S is the set of security test types, d is a vulnerability or security threat (hereinafter referred to as a security test item), D s is the set of security test items for the security test type s, w s is the weight of the security test type s, detect(d, v) is an indicator function, which is 1 if the test case individual v can detect the security test item v, otherwise 0, T s is the detection threshold of the security test type s, and e is the base of the natural logarithm.
[0199] b is the user behavior pattern, B is the set of user behavior patterns, U b is the set of user operation sequences for the user behavior pattern b, w b is the weight for the user behavior pattern b, p(u) is the probability that the user operation sequence u appears, and match(u, v) is a matching function that reflects the matching degree between the test case individual v and the user operation sequence u.
[0200] In step S202, in one embodiment, the selection probability p of the particle can be calculated by the following formula (1-13) select (v):
[0201]
[0202] where p select (v) represents the probability that the test case individual v corresponding to the current position of the particle is selected; p select (v) ∈ [0, 1], and the sum of the selection probabilities of all test case individuals v is 1; F(v) is the fitness score of the test case individual v corresponding to the particle. The higher F(v) is, the better the quality of the test case individual v is, and the greater the probability of being selected; Population represents the test case set corresponding to the current position of the particle.
[0203] ∑ v′∈Population F(v′) represents the sum of the fitness scores of the test case individuals corresponding to the current positions of all particles in the current particle swarm.
[0204] In another embodiment, in order to enhance the influence of the fitness difference of the test case individuals and make the influence of the fitness difference of the test case individual v on the selection probability more significant. Calculate the selection probability p of the test case individual based on formula (1-14) select (v):
[0205]
[0206] T is a regulatory parameter used to adjust the sensitivity of the selection mechanism to fitness differences. When the value of T is large, the probability distribution is smoother, and test case individuals with low fitness have a certain probability of being selected, which helps to maintain the population diversity. When T is small, the selection is more concentrated on test case individuals with high fitness, accelerating convergence but possibly leading to premature convergence to a local optimum. By adjusting the parameter T, the smoothness of the selection probability distribution is controlled, allowing individuals with lower fitness scores to also have a certain probability of being selected. Among them, in the selection process of the same round of optimization, the same regulatory parameter T is used to calculate for all test case individuals, so as to ensure the implementation of a consistent selection strategy for the entire test case set, facilitating the comparison of the fitness of different test case individuals and the assignment of selection probabilities. For different iteration rounds, the value of T can be dynamically adjusted to enhance or weaken the "amplification" or "smoothing" effect on fitness differences.
[0207] In this embodiment, through the operation of e F(V) / T the fitness score of the test case individual v is exponentiated to enhance the influence of the fitness difference of the test case individual, making the influence of the fitness difference of the test case individual v on the selection probability more significant.
[0208] In step S203, in order to calculate the cumulative probability P select (v) of the test case individual, all test case individuals are queued and numbered n, where n = 1... N, and N is the total number of test case individuals in the seed set. For any test case individual v in the queue i Take the test case individual v i The selection probability P select (v i ) and the sum of the selection probabilities p select (v) of all test case individuals sorted in front of it are used as its cumulative probability P select (v i ), that is Thus, N cumulative probability values are obtained, forming a cumulative probability value queue. Then in step S204, the range of the random number is 1 to N, and the random number is used as the sorting of the cumulative probability value queue, so as to obtain the test case individual corresponding to the position indicated by the random number. Since each test case individual corresponds to a particle, after selecting a test case individual, the corresponding particle is determined.
[0209] In step S105, when moving the position of a particle at a preset speed, first calculate the current moving speed of the particle according to the following formula (1-15) based on the current position of the particle, the individual best position of the particle, and the global pheromone highest position.
[0210] v i (t + 1) = w·v i(t)+c1·rand1·(pbest i -x i (t))+c2·rand2·(τ gbest -x i (t)) (1-15)
[0211] Among them, v i (t) and x i (t) represent the velocity and position of particle i at the tth optimization, respectively. When t = 1, v i (t) is the preset initial speed value, and x i (t) is the initial position of the particle, i.e., the initial test case individual.
[0212] w is the inertia weight, c1 and c2 are learning factors, rand1 and rand2 generate random numbers in the interval [0,1], pbest i is the individual best position of particle i. In one embodiment, to obtain the individual best position pbest of particle i i During the movement of particle i, the fitness score of the test case individual corresponding to the new position obtained after each movement is calculated, and the test case individual with the highest fitness score is taken as the individual best position pbest i τ gbest It is the global best position selected according to the pheromone strength. In one embodiment, based on the global best reinforcement strategy, after each pheromone update, the global best position with the largest pheromone strength in the current optimization process can be obtained and recorded. Then the global best position in all optimization processes is queried, that is, the historical best position can be obtained, that is, the global best position with the largest pheromone strength recorded in all iterations is used.
[0213] Then, based on formula (1-16), the new position x(t+1) of the particle is calculated.
[0214] x(t+1)=v(t+1)+x(t) (1-16)
[0215] That is, the current moving speed v(t+1) is added to the current position x(t) of the particle to obtain the new position x(t+1) of the particle, and the new position x(t+1) corresponds to a new test case individual.
[0216] Then the new test case individuals are merged with the original test case individuals into a new test case set.
[0217] As can be seen from Formulas 1-15, when calculating the new velocity of a particle, a part of the current velocity is retained through the inertial term of the first item to maintain the historical inertia of the search direction. Through the second item, the particle is adjusted towards its own historical optimal position. Through the third item, the particle is adjusted towards the global historical optimal position. The global historical optimal position in this embodiment is the position with the strongest pheromone concentration in the ant colony algorithm.
[0218] From the foregoing description, it can be seen that the position of a particle is closely related to the particle velocity. In order to obtain test case individuals that meet different emphasis aspects, in one embodiment, the learning factors c1, c2, and / or the inertia weight are as shown in the following Formulas (1-17), (1-18), and (1-19), including one or more influencing factors and weights.
[0219]
[0220] where g k and are the kth influencing factor and its weight respectively, and m is the total number of influencing factors. The influencing factors are, for example, one or more of a security risk factor, an environmental impact factor, a user behavior matching rate factor, a business scenario factor, a coverage impact factor, an execution time impact factor, and a system load impact factor. Therefore, when calculating the new velocity of a particle, according to requirements, the learning factors c1, c2, and / or the inertia weight required in Formula (1-15) are first calculated. For example, when the learning factor includes the security risk factor weight, it is queried whether the software unit identifier corresponding to the current position in the software structure diagram includes a security risk identifier. If so, the value of the security risk factor is determined according to the preset defect priority assignment weight calculation method based on the security risk identifier type, quantity, etc. In another embodiment, the fitness calculation formula of the test case corresponding to the current position is queried, and the score calculated from the security test standard therein is used as the value of the security risk factor. For example, the value calculated according to in the second item of Formula (1-12) is used as the value of the security risk factor. By incorporating the security test standard into the global search strategy of the particle swarm algorithm, the purpose of preferentially searching for modules with higher risks can be achieved.
[0221] Similarly, when the environmental impact factor is added to the learning factors c1, c2, and / or the inertia weight, when calculating the particle velocity, the environmental impact value calculated according to Formula (1-8) based on the foregoing fitness function E(C, V) can be used as the environmental impact factor. When the user behavior matching degree is added to the learning factors c1, c2, and / or the inertia weight, the value calculated according to The calculated value is used as the user behavior matching degree factor, and is incorporated into the calculation process of the particle position through environmental factors and the user behavior model to generate test cases that are more in line with the real usage scenario. Additionally, in the ant colony algorithm, pheromones are increased for certain scenarios corresponding to some test cases. For example, when the software is mainly used in scenarios with frequent user interactions (such as e-commerce platforms), for some paths in these scenarios, such as paths related to user input and order processing, the pheromone values are higher. Therefore, in order to preferentially search for these critical paths when passing through the particle swarm optimization algorithm, a pheromone factor can be added to the learning factors c1, c2, and / or the inertia weight, that is, the pheromone of the test case individual corresponding to this position is used as one of the influencing factors. When the pheromone value is high, it can preferentially move towards this path.
[0222] In another embodiment, when evaluating pbest and gbest, one or more of the security risk factor, environmental impact factor, user behavior matching degree factor, business scenario factor, coverage impact factor, execution time impact factor, and system load impact factor can also be added to the evaluation formula, and a test case set for multiple scenarios can also be generated accurately and efficiently.
[0223] After obtaining the new test case set through step S105, in step S106, first calculate the fitness score of each test case individual according to formula (1-12). Then, eliminate the test case individuals that do not meet the requirements through the fitness score. The conditions to be met are, for example, that the fitness score is greater than the threshold, or when sorted according to the fitness score, it needs to be ranked among the top N. The test case individuals that meet the conditions are used as the test case individuals for the next optimization, and those that do not meet the conditions are eliminated, which not only avoids the large number of test cases but also retains high-quality test cases, accelerating the convergence speed of the optimization.
[0224] The optimization end condition in step S107 is, for example, that the current test case set meets the system test objectives, and the system test objectives include, for example, software system test objectives, security test objectives, and user behavior test objectives. For the software system test objectives, its corresponding coverage criterion is, for example, specifically to judge whether the current test case set covers the specified software internal structure and reaches the specified coverage rate, whether it reaches the code coverage rate, function call coverage rate, and condition judgment coverage rate of the module or component, and whether it reaches the specified environmental fitness (E(C,V)); for the security test objectives, it corresponds to the security criterion, for example, whether the specified vulnerabilities or threats are detected, or whether a specified number or proportion of vulnerabilities or threats are detected. For the user behavior test objectives, it corresponds to the user behavior matching criterion, for example, whether the specified user behavior coverage and matching degree are achieved. In one embodiment, the fitness score of the test case set is calculated according to the following formula (1-20):
[0225]
[0226] Among them, C, S, and B represent the coverage criteria of the software system, the types of security tests, and the set of user behavior patterns respectively, and w c 、w s and w b are the weight coefficients corresponding to each type; c is a coverage criterion, C is the set of coverage criteria, and the set of coverage criteria C includes various coverage criteria for the number of code lines, functions, modules, and path coverage and condition coverage. Coverage c (V) is the score of the software unit (such as various coverage criteria such as the number of code lines, functions, modules, and path coverage and conditions) covered by the test case set V calculated according to formula (1-21).
[0227]
[0228] Among them, I c is the set of covered units for the coverage criterion c, |I c | is the total number of covered units i, and covered(i, V) is an indicator function that is 1 if the test case set covers the covered unit i, and 0 otherwise.
[0229] s is the type of security test, S is the set of types of security tests; Satety s (V) is the score calculated based on the following formula (1-22).
[0230]
[0231] Among them, d is a vulnerability or security threat (hereinafter referred to as a security test item, such as SQL injection, cross-site scripting attack, and buffer overflow, etc.), D s is the set of security test items for the type of security test s, detect(d, v) is an indicator function that is 1 if the test case individual v can detect the security test item v, and 0 otherwise, T s is the detection threshold for the type of security test s, and e is the base of the natural logarithm.
[0232] b is the user behavior pattern, B is the set of user behavior patterns; S45, the user behavior coverage degree Behavior b (V) is the score calculated based on formula (1-23), which is used to evaluate the degree of coincidence between the test case set and the user behavior model, and reflects the effect of the test case set in simulating real user behavior.
[0233]
[0234] Among them, Ub is a set of user operation sequences for user behavior pattern b. p(u) is the probability of the occurrence of user operation sequence u, and match(u, V) is a matching function that reflects the matching degree between the test case set V and the user operation sequence u. The user behavior coverage Behavior b (V) reflects the user behavior coverage and matching degree, that is, the user behavior covered by the user behavior in the test case and the corresponding matching degree.
[0235] When the fitness score of the test case set reaches the threshold, it is considered that the optimization end condition is met. At this time, the optimization is stopped and the current test case set V is stored final , V final ={V|V∈Populationattermination}. At this time, the finally obtained test case set V final contains the test case individuals with the highest fitness after multiple rounds of iterative optimization. In addition, the optimization end condition also includes the number of iterations. When the number of iterations is reached, the optimization is also stopped. When the optimization is terminated only by the number of iterations, the finally generated test case set V final can also be evaluated. For example, the fitness score of the test case set V is calculated according to the foregoing fitness calculation formula final to evaluate whether the predetermined test coverage criteria are met, including code coverage rate, the ability to discover potential security vulnerabilities, and simulating user behavior.
[0236] As can be seen from the foregoing method, this embodiment utilizes the path optimization ability of the ant colony algorithm (Ant Colony Optimization, abbreviated as ACO) and the global search ability of the particle swarm optimization algorithm (Particle Swarm Optimization, abbreviated as PSO), and can realize the rapid exploration of software test paths and the optimized generation of test cases in complex path coverage scenarios. At the same time, according to the update rule of path pheromone and the generation rule of particle velocity in the ant colony algorithm, during multiple particle swarm iterative optimization processes, it can optimize in the direction of low path coverage rate, high defect risk level, many security vulnerabilities, more matching with the user's real scenario, and more conforming to the user's real behavior, that is, it improves the depth and breadth of path exploration and can also converge quickly. By adding calculation weights in aspects such as user behavior matching degree, environmental impact, and security risk to the information update and particle velocity update rules, the generated test cases are more in line with user behavior, more matching with the real application scenario, and also improve the probability of discovering potential defects and the diversity of the operating environment.
[0237] In addition, in a further embodiment, corresponding to Figure 1, before performing step S103, it also includes a step of scenario classification. For example, calculate the following scenario classification features and conditions: the McCabe complexity (McCabe's Cyclomatic Complexity) is greater than or equal to a threshold, or the average branch depth of a software module is greater than or equal to a threshold, and the number of input variable value combinations corresponding to a software unit in the input parameter space is less than a threshold (for example, the input variable dimension is lower than the dimension number threshold, or the value range of the parameter value is small, less than the range threshold); the number of path branches in the target area of the call flow graph is greater than a threshold; the area with unsatisfactory coverage of existing test cases; the area with coverage less than a threshold during the historical test process, or the known high-risk paths that have not been fully tested. Among them, the McCabe complexity is also known as the Cyclomatic Complexity, and its measurement method is well-known to those of ordinary skill in the art and will not be elaborated here. When any two or more of the above classification conditions are met, it can be considered that the current test scenario is a complex path coverage scenario. At this time, step S103 is executed, and then the foregoing method is used to generate test cases. If none of the above classification conditions are met, or less than two, it is determined whether other test scenarios are met, such as whether the input combination explosion scenario is met, and whether the high-risk security test scenario is met. If the input combination explosion scenario is met, a method that combines the Genetic Algorithm (GA) and Particle Swarm Optimization (PSO) is used to generate test cases. If the high-risk security test scenario is met, a method that combines the Ant Colony Optimization (ACO) and Genetic Algorithm (GA) is used to generate test cases. If none of the obvious test scenarios are met, any of the foregoing methods or a method applicable to the general scenario can be used to generate test cases.
[0238] The present invention captures low-frequency behaviors and complex operation patterns by analyzing historical data, and saves the data representing real user behaviors in the form of a user behavior model, and uses the user behavior model to guide the generation of test cases, so that the test cases can simulate the actual interaction scenarios. The present invention utilizes the excellent performance of the ant colony algorithm in path exploration and the particle swarm optimization algorithm in global optimization. By dynamically adjusting the path selection probability and using the global search ability, it not only enhances the coverage ability of difficult-to-reach paths, but also ensures that the test cases can handle complex interaction patterns. The present invention evaluates the fitness of test case individuals and the overall test case set in three dimensions: coverage rate, security test ability, and user behavior simulation, strengthens the contribution of complex path and interaction coverage and defect discovery ability in fitness evaluation, ensures that the iterative optimization converges rapidly towards the test target, improves the generation efficiency of test cases, and moreover enhances the test coverage rate and security verification ability of the obtained test case set for the system.
[0239] SeeFigure 5 , Figure 5 is a schematic block diagram of a software test case generation system according to an embodiment of the present invention. The system includes an initialization module 11, a user behavior model construction module 12, an input parameter space construction module 13, an initial set construction module 14, and an optimization module 15. Among them, the initialization module 11 is used to initialize the test environment to obtain the path exploration space of the target software. Among them, the path exploration space includes nodes and directed edges between the nodes. Two nodes with a directed relationship form a path branch, and each path branch corresponds to a path branch condition for implementing the path branch. The user behavior model construction module 12 is used to construct a user behavior model based on the collected user historical behavior data. Among them, the user behavior model includes one or more user parameters simulating user behavior and their parameter values. The input parameter space construction module 13 is used to construct the input parameter space of the target software. Among them, the input parameter space includes multiple input variables, and the variable value of each input variable corresponds to one or more user parameter values. The initial set construction module is used to drive multiple ants to perform path exploration in different regions of the path exploration space respectively, and update the path branch pheromone during the path exploration process. After each ant determines a path branch during the path exploration process, based on the path branch condition, it selects the corresponding input variable value from the input parameter space according to the user behavior model. After each ant finishes the path exploration, the input variable values of all path branches corresponding to the explored path are obtained to form an initial test case individual. All initial test case individuals form an initial test case set. The optimization module 15 is used to map the initial test case set to an initial particle swarm, and iteratively optimize the initial particle swarm based on the particle swarm optimization algorithm until the obtained test case set meets the end condition. For specific details, please refer to the foregoing description of the method, which will not be elaborated here.
[0240] Figure 6 is a schematic architecture diagram of a software test system according to an embodiment of the present invention. The software test system includes an operation end and a service end. Among them, the operation end is located at the terminal device 102, and the service end is located at the server 104 or a server cluster. The terminal device 102 communicates with the server 104 through the network. The terminal device 102 includes a desktop computer, a notebook computer, or a mobile intelligent terminal device, such as a mobile phone, a tablet computer, etc. The service end includes the foregoing Figure 5The software test case generation system shown has an operation terminal including an interactive interface. When test cases need to be generated, testers configure corresponding parameters through the interactive interface and send them to the server via the network. The server generates a test case set based on the received instructions and corresponding parameters according to the method described above in the present invention. The parameters configured by the testers through the interactive interface include, for example, the name of the target software, version, storage address of user historical data, number of iterations, storage address of the test case set, and so on. After the test case set is generated, the testers can send a test instruction to the server through the interactive interface. After receiving the test instruction, the server executes software testing using the test case set and records the test results.
[0241] Figure 7 It is a schematic diagram of the hardware structure principle of an electronic device according to an embodiment of the present invention. The electronic device can be implemented as a server or various other terminal devices, such as a desktop personal computer, a tablet computer, a laptop computer, a mobile phone, etc. It includes a processor 601 and a memory 602. A program instruction set is stored on the memory 602, and when the processor 601 executes the program instruction set on the memory 602, the aforementioned software test case generation method is implemented.
[0242] Specifically, the aforementioned processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0243] The memory 602 may include a mass storage for data or instructions. By way of example and not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 602 may include removable or non-removable (or fixed) media. In a suitable case, the memory 602 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 602 is a non-volatile solid state memory.
[0244] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the software test case generation method provided by the present invention.
[0245] In one example, the electronic device may further include a communication interface 603 and a bus 604. The processor 601, the memory 602, and the communication interface 603 are connected through the bus 604 to complete communication with each other.
[0246] The communication interface 603 is mainly used to implement communication between each module, device, unit, and / or device in the embodiments of the present invention.
[0247] The bus 604 includes hardware, software, or both, and couples the components of the online data flow charging device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 604 may include one or more buses. Although the embodiments of the present invention describe and illustrate specific buses, the present invention contemplates any suitable bus or interconnect.
[0248] The present invention also provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, they can implement any one of the software test case generation methods in the foregoing embodiments. The computer-readable storage medium may be any medium that is tangible and contains or stores computer-executable instructions for use by or in connection with an instruction execution system, apparatus, and device. The storage medium may be a transitory computer-readable storage medium or a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium may include, but is not limited to, a magnetic storage device, an optical storage device, and / or a semiconductor storage device. Corresponding embodiments of such storage devices include, for example, magnetic disks, optical discs based on CD, DVD, or Blu-ray technology, and persistent solid-state memories such as flash memories, solid-state drives, etc.
[0249] The present invention also provides a computer program product, which includes a set of computer program instructions. When the set of computer program instructions is executed by a processor, it implements any one of the software test case generation methods in the foregoing embodiments. The computer program product includes, but is not limited to, application installation packages, application plugins, applets that can run in certain applications, etc. published on websites and application stores.
[0250] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0251] The above embodiments are only for illustrating the present invention and are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can also make various changes and variations without departing from the scope of the present invention. Therefore, all equivalent technical solutions should also fall within the scope of the disclosure of the present invention.
Claims
1. A method for generating software test cases, characterized in that: include: Initialize the test environment to obtain a path exploration space of the target software; wherein the path exploration space includes nodes and edges with pointing between the nodes, two nodes with pointing relationships constitute a path branch, and each path branch includes a path branch condition; Based on the collected user historical behavior data, a user behavior model and an input parameter space of the target software are constructed, wherein the user behavior model includes one or more user parameters and parameter values simulating user behavior; the input parameter space includes a plurality of input variables, and the variable value of each input variable corresponds to one or more user parameter values; Drive multiple ants to perform path exploration in different areas of the path exploration space, and update the path branch pheromone value during the path exploration process; after each ant determines the path branch during the path exploration process, based on the path branch condition, select the corresponding input variable value from the input parameter space according to the user behavior model; after each ant completes the path exploration, the input variable values of all path branches of the explored path constitute an initial test case individual; all initial test case individuals constitute an initial test case set; and The initial test case set is mapped to the initial particle swarm, and the initial particle swarm is iteratively optimized until the optimized test case set meets the optimization end condition.
2. The method for generating software test cases according to claim 1, characterized in that: The steps for each ant to determine the path branch during the path exploration process include: Based on the directional relationship of the edge between the two nodes, determine one or more candidate nodes that form a path branch with the current node; Calculate the selection probability from the current node to each candidate node based on the path branch pheromone value and the heuristic factor; and The candidate node with the highest selection probability is determined as the next node that forms a path branch with the current node; wherein the heuristic factor at least includes a user behavior matching weight.
3. The software test case generation method according to claim 2, characterized in that: The heuristic factor also includes one or more of the following weights: detection execution time weight, operating environment weight, path length weight and security risk weight.
4. The software test case generation method according to claim 1, characterized in that: The steps for updating the path branch pheromone value during the path exploration process include: During the path exploration process, it is determined whether the currently explored path branch is marked with a security risk attribute tag. In response to the currently explored path branch being marked with a security risk attribute tag, a corresponding incremental value is increased based on the current pheromone value of the path branch marked with the security risk attribute tag.
5. The method for generating software test cases according to claim 1, characterized in that: The path exploration process further includes: Determine the scene corresponding to the current path exploration space area; and The pheromone value of the current path branch is adjusted based on a method that matches the scenario.
6. The software test case generation method according to claim 1, characterized in that: Further including: In the process of driving multiple ants to explore paths in different areas of the path exploration space, the coverage rate of the explored paths to the path exploration space or the preset local path exploration space is calculated, and the pheromone values of the path branches in the path exploration space are adjusted based on the corresponding relationship between the coverage rate and the negative correlation between pheromones.
7. The software test case generation method according to claim 1, characterized in that: After each ant finishes the path exploration, the input variable values of all path branches of the explored path constitute candidate initial test case individuals. The method further includes: Construct the maximum problem function of the initial test case set based on the environmental fitness and user behavior model; Configure the operating environment parameters for each candidate initial test case individual; Calculate the environmental fitness score of each candidate initial test case individual based on the operating environment parameters of the candidate initial test case individual; Calculate the user behavior simulation score of each candidate initial test case individual based on the user behavior model used when generating the candidate initial test case individual; Calculating the maximum problem function value of the initial test case set based on the environmental fitness score and the user behavior simulation score of each candidate initial test case individual, and adjusting the operating environment parameters of the candidate initial test case individual and / or the user behavior model used to generate the candidate initial test case individual so that the maximum problem function converges to a maximum value; and The candidate initial test case individuals that converge to the maximum value are determined as the initial test case individuals, and all initial test case individuals constitute the initial test case set.
8. The software test case generation method according to claim 1 or 7, characterized in that: An optimization step in the iterative optimization process of the initial particle swarm includes: The current test case set is constructed by the test case individuals corresponding to the current position of each particle in the current particle swarm; Selecting some test case individuals from the current test case set as seed test case individuals, and determining seed particles corresponding to the seed test case individuals; Calculating a new moving speed of the seed particle based on the current position and current speed of the seed particle, the individual best position of the seed particle, and the highest position of the global pheromone; and The new moving speed is added to the current position of the seed particle to obtain a new position of the seed particle, and the new position of the seed particle corresponds to a new test case individual; wherein the current seed test case individual and the new test case individual constitute a new test case set.
9. The software test case generation method according to claim 8, characterized in that: The steps of selecting some test case individuals from the current test case set as seed test cases include: Get the fitness score of each test case in the current test case set; and Traverse the fitness scores of each test case individual, select the test case individuals whose fitness scores are greater than the threshold or a preset number of test case individuals that are ranked first when the fitness scores are sorted from large to small as seed test cases; Alternatively, some test case individuals are selected from the current test case set based on a roulette wheel selection mechanism, wherein, when some test case individuals are selected from the current test case set based on the roulette wheel selection mechanism, the selection probability of the test case individual is calculated based on the pheromone value of the test case individual coverage path and the fitness score of the test case individual.
10. The method for generating software test cases according to claim 8, characterized in that: After getting the new test case set, further include: Get the fitness score of each test case in the current test case set; and Traverse the fitness scores of each test case individual and eliminate the test case individuals whose fitness scores are less than the threshold.
11. The method for generating software test cases according to claim 9 or 10, characterized in that: When the current test case set is the initial test case set, the steps of obtaining the fitness score of each individual test case in the current test case set include: Calculate the coverage of each initial test case individual on the target software indicator; Evaluate each individual initial test case based on the threat model for security testing to obtain a defect disclosure capability value; and Calculate the weighted sum of the coverage rate and defect revealing capability value of each initial test case individual for the software indicator as the initial fitness score of the initial test case individual; When the current test case set is a test case set obtained through the previous optimization or a new test case set obtained through the current optimization, the step of obtaining the fitness score of each individual test case in the current test case set includes: Run each test case in the current test case set and collect the corresponding running data; Calculate the coverage of the target software indicator by the individual test case based on the running data of each individual test case; Evaluate each test case based on the threat model for security testing to obtain the defect disclosure capability value; Querying the user behavior model data and calculating the matching rate between the user behavior simulated by each test case individual and the user behavior model; and The fitness score of each test case is calculated as the weighted sum of the coverage of the target software indicator, the defect revealing capability value, and the matching rate between the user behavior simulated by the test case and the user behavior model.
12. The software test case generation method according to claim 8, characterized in that: When calculating the new moving speed of the seed particle, the current moving speed of the seed particle is adjusted based on one or more of security risks, environmental impact, user behavior matching, business scenarios, coverage impact, execution time impact, and system load impact.
13. The method for generating software test cases according to claim 11, characterized in that: The optimization end condition is that the number of iterations reaches a threshold or the fitness score of the current test case set reaches a threshold.
14. A software test case generation system, characterized in that: include: An initialization module, configured to initialize the test environment to obtain a path exploration space of the target software; wherein the path exploration space includes nodes and edges with pointing between the nodes, two nodes with pointing relationships constitute a path branch, and each path branch includes a path branch condition; A user behavior model building module, configured to build a user behavior model based on the collected user historical behavior data, wherein the user behavior model includes one or more user parameters and parameter values simulating user behavior; An input parameter space construction module configured to construct an input parameter space of the target software, wherein the input parameter space includes a plurality of input variables, and a variable value of each input variable corresponds to one or more user parameter values; An initial set construction module is configured to drive multiple ants to perform path exploration in different areas of the path exploration space, and update path branch pheromones during the path exploration process; after each ant determines a path branch during the path exploration process, based on the path branch condition, selects a corresponding input variable value from the input parameter space according to the user behavior model; after each ant completes the path exploration, the input variable values of all path branches of the explored path constitute an initial test case individual; all initial test case individuals constitute an initial test case set; and The optimization module is configured to map the initial test case set to an initial particle group, and iteratively optimize the initial particle group until the optimized test case set meets the optimization end condition.
15. An electronic device, characterized in that: The electronic device comprises a processor and a memory storing computer program instructions; when the electronic device executes the computer program instructions, the method for generating software test cases as described in any one of claims 1 to 13 is implemented.
16. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the software test case generation method as described in any one of claims 1 to 13 is implemented.
17. A computer program product, characterized in that It includes a computer program instruction set, which, when executed by a processor, implements the software test case generation method described in any one of claims 1 to 13.
Citation Information
Patent Citations
A combined approach to accelerate test case generation using genetic methods and symbolic execution.
CN109344057B