Software test case generation method and system, electronic equipment and storage medium
By building a user behavior model and input parameter space, combining genetic algorithms and particle swarm optimization technology, test cases reflecting the real use of the target software are generated, and the problems of low efficiency and insufficient coverage in the existing technology are solved, and efficient and comprehensive coverage test case generation is achieved.
Patent Information
- Application Number
- CN202510476875.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to generate test cases that reflect the real use of target software, and the test case generation efficiency is inefficient.
By initializing the test environment, building a user behavior model and input parameter space, generating an initial test case set, and iteratively optimized through a combination of genetic algorithm and particle swarm optimization to generate a test case set that meets the optimization end conditions.
It improves the efficiency of the generation of test cases, enables the generated test cases to better reflect the actual use of the software, and enhances the test coverage and depth detection capabilities.
Smart Images

Figure CN120216384A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer software testing, and particularly to a method, system, electronic device and storage medium for generating software test cases. Background Art
[0002] In the field of software engineering, software testing is a process for discovering and fixing defects in software, verifying whether the software meets the design requirements, and ensuring its quality and performance. It is a crucial link in the software development process. Test cases are the core part of software testing, directly related to the efficiency and coverage of testing, and affecting the quality and reliability of software products. A test case can be regarded as a specific test task, containing elements such as input, operation, and expected output, and is a data set in various forms, such as execution paths, input data, and execution conditions.
[0003] The methods for generating test cases can generally be divided into the following several types:
[0004] First, test cases are manually written by testers. This means that testers need to have very rich test experience and relatively high professional levels, and there are certain drawbacks such as blindness, high cost, and difficulty in improving test coverage. Although the method of manually writing test cases is effective in some cases, with the continuous increase in the complexity of software systems, it has become increasingly difficult to manually write test cases, which is time-consuming and prone to missing errors, especially in cases where high coverage and in-depth detection are required. In addition, the manual method is difficult to adapt to the rapid iteration of software. Updating a function may require rewriting a large number of test cases, which significantly increases the cost of software development and maintenance.
[0005] Second, some tools are used to automatically generate test cases based on certain methods, such as random testing, model-based testing (MBT), symbolic execution, search-based testing, etc. Although these tools have made progress in some fields, there are still many unsatisfactory aspects. For example, model-based testing requires detailed model design, which is itself a complex and time-consuming process; symbolic execution and search-based testing face the problem of state space explosion and are difficult to handle large software systems.
[0006] III. Combine with other algorithms. For example, adopt the genetic algorithm (Genetic Algorithm, abbreviated as GA) or the method of combining the genetic algorithm with other automated methods. For example, the Chinese patent with the publication number CN109344057B and the invention title "Combined accelerated test case generation method based on genetic method and symbolic execution" discloses a test case generation method combining the genetic algorithm. However, the prior art often ignores the diversity of user behavior patterns and software operating environments, which makes the generated test cases unable to comprehensively reflect the real usage situation. Moreover, in the case of a huge input parameter scale, the search efficiency is low and the convergence speed is slow. Summary of the Invention
[0007] In view of the technical problems existing in the prior art, the present invention proposes a software test case generation method, system, electronic device and storage medium to generate test cases that reflect the real usage situation of the target software and improve the generation efficiency of test cases.
[0008] To solve the above technical problems, according to one aspect of the present invention, the present invention provides a software test case generation method, including the following steps:
[0009] Initialize the test environment to obtain the operation environment simulation configuration parameters and the structure parameters of the target software;
[0010] Construct a user behavior model and the input parameter space of the target software based on the collected user historical behavior data, wherein the user behavior model includes one or more user parameters simulating user behavior and their parameter values; the input parameter space includes multiple input variables, and the variable value of each input variable corresponds to one or more user parameter values used to form the component parameters of the test case individual;
[0011] Construct an initial test case set based on the input parameter space of the target software, the user behavior model and the test environment, wherein the initial test case set includes a preset number of initial test case individuals, and each initial test case individual includes multiple types of component parameters, and the categories of the component parameters at least include input operations and corresponding input parameters and operation environment parameters; and
[0012] Perform the following iterative optimization steps on the test case individuals based on the initial test case set:
[0013] Regard the test case individual as a chromosome and the component parameters of the test case as genes, and perform a reproduction operation on the test case individuals in the current test case set to obtain new test case individuals, wherein, in the first optimization, the current test case set is the initial test case set;
[0014] Map the test case individuals to particles, use the new test case individuals obtained by reproduction as the current positions of the corresponding particles, and calculate the new moving speeds of the particles based on the current positions, current speeds, individual best positions, and global best positions of the particles;
[0015] Superimpose the new moving speed on the current position of the particle to obtain the new position of the particle, the new position of the particle corresponds to a new test case individual, and the new test case individuals and the test case individuals without reproduction operations are combined together to form a new test case set; and
[0016] Evaluate whether the new test case set meets the optimization end condition. When the new test case set meets the optimization end condition, stop the iterative optimization and store the new test case set; when the new test case set does not meet the optimization end condition, repeat the foregoing iterative optimization steps for the new test case set.
[0017] Optionally, the steps of constructing the initial test case set based on the input parameter space, user behavior model, and test environment of the target software include:
[0018] Configure the running environment parameters of the initial test case individuals based on the running environment simulation configuration parameters obtained from the initialized test environment;
[0019] Determine the initial path based on the test path space constructed from the target software structure parameters and the user behavior model;
[0020] Determine the input variables that can implement the initial path from the input parameter space;
[0021] Select the input variable values that can simulate user behavior for each input variable from the input parameter space according to the user behavior model; and
[0022] Use the running environment parameters and one or more input variable values that implement the initial path as the component parameters of the test case respectively to form an initial test case individual.
[0023] Optionally, use the running environment parameters and one or more input variable values that implement the initial path as the component parameters of the test case respectively to form a candidate initial test case individual. After obtaining multiple candidate initial test case individuals, it further includes:
[0024] Construct the maximum problem function of the initial test case set based on the environmental fitness and the user behavior model;
[0025] Calculate the environmental fitness scores of each candidate initial test case individual based on the running environment parameters of the candidate initial test case individuals;
[0026] Calculate the user behavior simulation score for each candidate initial test case individual based on the user behavior model when determining the input variable values;
[0027] Calculate the maximum problem function value of the initial test case set based on the environmental fitness score and the user behavior simulation score of each candidate initial test case individual, and adjust the candidate initial test case individual running environment parameters and / or the user behavior model to make the maximum problem function converge to the maximum value; and
[0028] Determine the candidate test case individual when converging to the maximum value as the initial test case individual, and all initial test case individuals constitute the initial test case set.
[0029] Optionally, taking the test case individual as a chromosome and the composition parameters of the test case as genes, the steps of performing a reproduction operation on the test case individuals in the current test case set to obtain new test case individuals include:
[0030] Select some test case individuals from the current test case set as the parent test case individuals; and
[0031] Perform a crossover operation and / or a mutation operation on the genes in the parent test case individuals to obtain new test case individuals.
[0032] Optionally, the steps of selecting some test case individuals from the current test case set as the parent test case individuals include:
[0033] Obtain the fitness score of each test case individual in the current test case set; and
[0034] Traverse the fitness scores of each test case individual, and use the test case individuals with fitness scores greater than the threshold as the parent test case individuals;
[0035] Or, sort the test case individuals in the current test case set in descending order according to the fitness score; use the preset number of test case individuals ranked at the front as the parent test case individuals;
[0036] Or, select the parent test case individuals from the current test case set based on the roulette wheel selection mechanism, where when selecting the parent test case individuals from the current test case set based on the roulette wheel selection mechanism, calculate the selection probability of the test case individuals to be tested based on the fitness scores of the test case individuals.
[0037] Optionally, when the current test case set is the initial test case set, the steps of obtaining the fitness score of each test case individual in the current test case set include:
[0038] Calculate the coverage rate of each initial test case individual for the target software metrics;
[0039] Evaluate each initial test case individual based on a threat model for security testing to obtain a defect revelation ability value; and
[0040] Calculate the weighted sum of the coverage rate of each initial test case individual for the target software metric and the defect revelation ability value as the initial fitness score of the initial test case individual;
[0041] When the current test case set is a test case set that has been optimized, the steps of obtaining the fitness score of each test case individual in the current test case set include:
[0042] Run each test case individual in the current test case set and collect the corresponding running data;
[0043] Calculate the coverage rate of the test case individual for the target software metric based on the running data of each test case individual;
[0044] Evaluate each test case individual based on a threat model for security testing to obtain a defect revelation ability value;
[0045] Query the user behavior model data and calculate the matching rate between the user behavior simulated by each test case individual and the user behavior model; and
[0046] Calculate the weighted sum of the coverage rate of each test case individual for the target software metric, the defect revelation ability value, and the matching rate between the user behavior simulated by the test case individual and the user behavior model as the fitness score of each test case individual.
[0047] Optionally, when performing a mutation operation on the genes in the parent test case individual, obtain the input variable value corresponding to the mutation gene category from the input parameter space in a manner that simulates new user behaviors and / or triggers security risks.
[0048] Optionally, when calculating the new moving speed of the particle, adjust the current moving speed of the particle based on one or more of security risks, environmental impacts, user behavior matching degrees, business scenarios, coverage rate impacts, execution time impacts, and system load impacts.
[0049] Optionally, after obtaining the new test case individual, further include:
[0050] Calculate the fitness score of the new test case individual; and
[0051] Traverse the fitness scores of each new test case individual, retain the new test case individuals with fitness scores greater than or equal to the threshold, and eliminate the new test case individuals with fitness scores less than the threshold;
[0052] Alternatively, sort the new test case individuals in descending order of fitness scores; retain the preset number of new test case individuals with the highest rankings.
[0053] Alternatively, construct a problem of maximizing the fitness score, calculate the maximum fitness score based on the current test case set, and retain the test case individuals when the maximum fitness score is obtained.
[0054] Optionally, the end condition is reaching the iteration count threshold or the fitness score of the new test case set reaching the threshold.
[0055] According to another aspect of the present invention, the present invention also provides a software test case generation system, including:
[0056] An initialization module configured to initialize the test environment to obtain operation environment simulation configuration parameters and the structural parameters of the target software;
[0057] A user behavior model construction module configured to construct a user behavior model based on the collected user historical behavior data, where the user behavior model includes one or more user parameters simulating user behavior and their parameter values;
[0058] An input parameter space construction module configured to construct a test parameter space for the target software, where the input parameter space includes multiple input variables, and the variable value of each input variable corresponds to one or more user parameter values for constituting the component parameters of the test case individuals;
[0059] An initial set construction module configured to construct an initial test case set based on the input parameter space of the target software, the user behavior model, and the test environment, where the initial test case set includes a preset number of initial test case individuals, and each initial test case individual includes multiple types of component parameters, and the types of the component parameters at least include input operations and corresponding input parameters and operation environment parameters; and
[0060] An optimization module configured to iteratively optimize the test case individuals based on the initial test case set; the optimization module includes:
[0061] A reproduction unit configured to use the test case individuals as chromosomes and the component parameters of the test cases as genes to perform reproduction operations on the test case individuals in the current test case set to obtain new test case individuals, where, in the first optimization, the current test case set is the initial test case set;
[0062] The particle swarm optimization unit is configured to map test case individuals into particles, use the newly obtained test case individuals through reproduction as the current positions of the corresponding particles, and calculate the new moving speeds of the particles based on the current positions, current speeds, the individual best positions, and the global best positions of the particles; superimpose the new moving speeds on the current positions of the particles to obtain the new positions of the particles, where the new positions of the particles correspond to new test case individuals, and the new test case individuals and the test case individuals without reproduction operations are combined to form a new set of test case individuals; and
[0063] The evaluation unit is configured to evaluate whether the new set of test case individuals meets the optimization end condition, stop iterative optimization and store the new set of test case individuals when the new set of test case individuals meets the optimization end condition; when the new set of test case individuals does not meet the optimization end condition, trigger the reproduction unit to start a new round of optimization.
[0064] According to another aspect of the present invention, the present invention also provides an electronic device, which includes a processor and a memory storing computer program instructions; when the electronic device executes the computer program instructions, the foregoing software test case generation method is implemented.
[0065] According to another aspect of the present invention, the present invention also provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing software test case generation method is implemented.
[0066] According to another aspect of the present invention, the present invention also provides a computer program product, including a set of computer program instructions, and when the set of computer program instructions is executed by a processor, the foregoing software test case generation method is implemented.
[0067] The embodiments of the present invention improve the generation efficiency of test cases and can reflect the real usage situation of software during testing. Description of the Drawings
[0068] Next, the preferred embodiments of the present invention will be further described in detail with reference to the drawings, where:
[0069] Figure 1 is a flowchart of a software test case generation method according to an embodiment of the present invention;
[0070] Figure 2 is a flowchart of a method for constructing an initial set of test case individuals according to an embodiment of the present invention;
[0071] Figure 3 is a schematic diagram of a path structure according to an embodiment of the present invention;
[0072] Figure 4 is a flowchart of a method for optimizing test cases according to an embodiment of the present invention to obtain an initial test case set;
[0073] Figure 5 is a flowchart of a method for selecting parent test case individuals for reproduction using a roulette wheel selection mechanism according to an embodiment of the present invention;
[0074] Figure 6 is a schematic block diagram of the principle of a software test case system according to an embodiment of the present invention;
[0075] Figure 7 is a schematic diagram of the architecture of a software test system according to an embodiment of the present invention; and
[0076] Figure 8 is a schematic diagram of the hardware structure principle of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0077] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0078] In the following detailed description, reference may be made to the accompanying drawings that form a part hereof and in which are shown, by way of illustration, specific embodiments in which the application may be practiced. In the drawings, like reference numerals describe substantially similar components in different views. The various specific embodiments of the present application are described in sufficient detail below to enable those of ordinary skill in the art with relevant knowledge and technology to implement the technical solutions of the present application. It should be understood that other embodiments may be utilized or structural, logical, or electrical changes may be made to the embodiments of the present application.
[0079] Figure 1 is a flowchart of a software test case generation method according to an embodiment of the present invention. In this embodiment, the software test case generation method includes the following steps:
[0080] Step S101, initialize the test environment to obtain the running environment simulation configuration parameters and the structure parameters of the target software.
[0081] Step S102: Construct a user behavior model and an input parameter space of the target software based on the collected historical user behavior data. The user behavior model includes one or more user parameters simulating user behavior and their parameter values. The input parameter space includes multiple input variables, and the variable value of each input variable corresponds to one or more user parameter values used to form the component parameters of a test case individual.
[0082] Step S103: Construct an initial test case set based on the input parameter space of the target software, the user behavior model, and the test environment. The initial test case set includes a preset number of initial test case individuals, and each initial test case individual includes multiple types of component parameters. The types of component parameters at least include input operations and corresponding input parameters and running environment parameters.
[0083] Step S104: Perform a reproduction operation on the test case individuals in the current test case set based on the genetic algorithm to obtain new test case individuals. The new test case individuals form a test case subset.
[0084] Step S105: Map the test case subset to an initial particle swarm.
[0085] Step S106: Move the particles to obtain new positions of the particles, and the new positions of the particles correspond to new test case individuals.
[0086] Step S107: Update the test case set.
[0087] Step S108: Evaluate whether the updated test case set meets the optimization end condition. When the updated test case set meets the optimization end condition, store the updated test case set and stop the iterative optimization in Step S109. When the updated test case set does not meet the optimization end condition, in Step S110, use the updated test case set as the current test case set for the new round of optimization and return to Step S104.
[0088] Among them, the running environment simulation configuration parameters in Step S101 include three types of parameters: the operating system version (abbreviated as O), network conditions (abbreviated as N), and hardware configuration data (abbreviated as H). Each type of running environment simulation configuration parameter can include one or more running environment simulation configuration parameters. In order to enable test cases to test the target software in various environments and verify its coverage rate and effectiveness, the present invention simulates the running states of different environments through software running environment simulation configuration parameters to provide diverse test scenarios for testing. In one embodiment, each running environment simulation configuration parameter is set as an environmental factor e i, since the software operating environment (or the system environment of the target software) has different impacts on different test cases, in order to measure this impact, an environment fitness function E(C, v) is defined. By measuring the fitness of a test case individual v to the environment under a specific system configuration, the impact of the software operating environment on the test case individual v is quantified. Among them, C represents the set of software operating environment simulation configuration parameters, C = {O, N, H}, where O represents the operating system version, N represents the network condition, and H represents the hardware configuration data. The present invention uses the set of environmental factors e i E = {e1, e2,..., e n} to represent a specific system simulation configuration. Each environmental factor e i can correspond to a set of one or more operating environment simulation configuration parameters. Each environmental factor e i is assigned a corresponding weight w i . Based on the comprehensive impact I of the fitness function E(C, v) on the environment configuration (abbreviated as the environmental impact value), it is calculated by the following formula (1-1):
[0089]
[0090] where f(e i ) is the impact value of the environmental factor e i . Among them, the calculation method of the impact value f(e i ) of the environmental factor e i is a conventional technique in the field of software testing and will not be elaborated here.
[0091] In one embodiment, the target software structure parameters in step S101 include a dynamic call graph matrix. When initializing in step S101 of the present invention, the dynamic call graph of the target software system is initialized, and the matrix representation method M(t) in graph theory is used to describe the interaction and dependency relationships between software units. Specifically, the interaction and dependency relationships between software units are described by formula (1-2):
[0092] M(t) = M(t - 1) + ΔM(t) (1-2)
[0093] M(t) represents the call matrix of software units within time period t, M(t - 1) represents the call matrix of software units within time period t - 1, and ΔM(t) represents the change in call data within time period t. The dynamic change situation of the software during operation is represented by the matrix of the call data of software units.
[0094] Among them, software units, such as components, modules, and even more detailed functions, etc., serve as the rows and columns of the call matrix, and the matrix elements represent the call data between software units within a preset time period, such as the number of times or frequency. In one embodiment, the matrix element M(t)[i,j] Denotes the number or frequency of the $i$-th software unit calling (jumps / calls) the $j$-th software unit within the time period $t$. For example, if "Component A" and "Component B" are mapped to the $i$-th row / column and the $j$-th column / row, then $M(t)$ [i,j] is the number of times Component A calls Component B during the time period $t$. When further refining "component" to the "function" level, $M(t)$ [f1,f2] denotes the number of times function $f1$ calls function $f2$ during the time period $t$.
[0095] As time accumulates, $M(t)$, $M(t + 1)$,... can be merged or superimposed to obtain the overall call matrix $M_{total}$. Thus, it can be seen that the call graph matrix not only includes the structural composition of the software, but also can obtain whether the call relationships of certain components / modules / functions are covered, as well as their call frequencies and control flows.
[0096] By parsing the dynamic call graph matrix to obtain two software units corresponding to each matrix element; using the software units as nodes and the call relationships between the software units as edges to construct a call flow graph, the call flow graph includes nodes and directed edges between the nodes, where two nodes connected by a directed edge form a path branch. According to the call relationships between the parsed software units, obtain the input operations and corresponding input parameter conditions for implementing the path branch, hereinafter referred to as branch conditions.
[0097] Specifically, query the elements greater than 0 in the current call graph matrix. If the count of a certain row and column in $M(t)$ is greater than 0, it means that the two software units represented by the corresponding row and column have indeed made calls during the test. Therefore, it can be obtained that these two software units have "triggered" the corresponding function call actions during the test. When modeling the internal structure of the software, use the two software units that have "triggered" the call actions as nodes, and use the directed call relationship of "software unit $i$ calls software unit $j$" as the "edge" connecting software unit $i$ and software unit $j$, thus forming a path branch. According to this method, connect all the call relationships where the matrix elements ($M(t)$ [i,j] ) are not 0 in series, and all the actual paths taken during the test are obtained, which constitute the call flow graph, where multiple consecutive software units and their call relationships constitute a dynamic call chain. If it does not appear (or is 0) in the matrix, it means that the call relationship between the corresponding two software units has not been tested and executed.
[0098] Therefore, based on the dynamic call graph matrix and the call flow graph, a test path space can be constructed, and the coverage of modules / functions and paths or path branches can be obtained through the elements of the dynamic call graph matrix obtained during the test.
[0099] Since there is a specific mapping relationship between software units and code, the code file or code line corresponding to the software unit can be obtained. Based on the matrix element M(t) [X,Y] ≠0, by combining with the coverage tool at the code level, it can be confirmed that "X calls a specific code branch or code line in Y".
[0100] In addition, security is usually also an important aspect in software testing. In one embodiment, when software units such as software units or code segments are marked with security risk attribute marks, the security risk attribute marks of the software units are obtained during initialization, so as to facilitate determining test case individuals that can perform security tests at this location when generating test cases.
[0101] In step S102, each piece of data in the user historical behavior data includes at least user identity information, time information, user operation information, and corresponding input information. These parameters representing various types of information are collectively referred to as user parameters. According to the category, user parameters include user operation parameters, input parameters corresponding to user operations, user identity parameters, and environment parameters. The parameter values of user operation parameters can be various user operations, such as click, select, input, etc.; the parameter values of input parameters are, for example, the specific content corresponding to the user operation, such as the confirmation key information corresponding to the click confirmation key operation, the product information corresponding to the select operation, the input content corresponding to the input operation, etc.; the parameter values of user identity parameters are, for example, information such as username and password, and the environment parameters are, for example, one or more types of information among the operating system version, network condition, and hardware configuration.
[0102] For example, for an online shopping platform system, a piece of user historical behavior data collected is as follows (time information is omitted):
[0103] User operations and corresponding input information:
[0104] Search keyword - smartphone;
[0105] Select product category - electronics;
[0106] Select product price range - (1000 - 2000) yuan;
[0107] User account information:
[0108] Username - user123;
[0109] Password - pass456;
[0110] Environment parameters:
[0111] User's device type - Android phone;
[0112] Operating system version - Android 12;
[0113] Network conditions - (Wi-Fi, bandwidth 10Mbps).
[0114] The parameter values of the user operation parameters extracted from the above data are "search keywords", "select product categories", "select product price ranges", etc. The corresponding input parameters are, for example, specific keywords (such as smartphones), specific product categories (such as electronic products), etc. The parameter values of the environmental parameters include the user device type ("Android phone"), the user device operating system version ("Android 12"), the network conditions used by the user ("Wi-Fi, bandwidth 10Mbps"), etc. The parameter value of the user identity parameter is specific account information, that is, the username "user123" and the password "pass456".
[0115] The user behavior model in the present invention includes various data capable of simulating user behavior. For example, the distribution probability or / and usage frequency of specific types of user parameter values, the sequence composed of multiple user parameter values representing different user behaviors; the association relationship between different categories / same category user parameter values representing user behavior patterns.
[0116] Taking the user operation parameters as an example, the parameter values of the user operation parameters are specific categories of user operations, and the distribution probability of each user operation is calculated based on formula (1-3).
[0117]
[0118] where x is a specific category of user operation, n x is the number of user operations x extracted from the historical data in a preset time period (such as one week, three days, one month, etc.), and N is the total number of all user operations of the user in the same time period. User operation x is, for example, "search keywords", "select product categories" or "select product price ranges" in the foregoing embodiments, etc. By calculating the distribution probability of each specific user operation, it is possible to distinguish whether a user operation conforms to the user's habitual operation. Similarly, the distribution probabilities of various other categories of user parameters can also be obtained, such as the distribution probabilities of various specific input parameter values, the distribution probabilities of various environmental parameter values, etc.
[0119] The usage frequency of user parameters can also reflect the corresponding user behavior pattern. For example, by counting the button click frequency of the same user in their respective statistical time periods at different times, the user's habit of clicking the button can be obtained, such as clicking 5 times per minute. Another example is that according to the average number of product pages browsed by different users during a search process, a user behavior pattern of accessing 10 pages each time can be obtained.
[0120] A sequence composed of multiple user parameter values can also represent corresponding user behaviors. For example, for a shopping platform, one user behavior model obtained is: Login - Click on Product 1 - Click on Product 2 - Click on Product 3 - Click on Product 4 - Click on Product 5 - Click on the purchase button on the product page... Another user behavior model obtained is: Login - Select product category, search for product keywords - Click on Product 1 - Click on Product 2 - Click on Add to Cart - Click on Product 3 -... - Click on the purchase button.... The sequences composed of multiple categories of user parameter values representing the above two different user behaviors are each a user behavior model, and these two user behavior models represent different user behavior patterns.
[0121] For another example, when the user behavior model is a sequence of specific user parameter values such as the operation of clicking on a product - specific product information - the operating system in the environmental parameters, the user behavior models composed of different specific parameter values represent different user behavior patterns.
[0122] The distribution probability of each of the above user parameter values, the usage frequency of each user parameter value, and sequences of multiple user parameter values are collectively referred to as the user behavior model.
[0123] The input parameter space of the target software includes multiple input variables, where the input variables correspond to user parameters, and the variable value of each input variable includes one or multiple corresponding user parameter values. In one embodiment, the input parameter space of the target software is represented by a set of variables: X = {x1, x2,..., x n}; X represents the input parameter space, and the variable x i represents the i-th input variable. The input variable can be an input operation or input parameter that the target software can accept. Here, the input operation corresponds to the user operation in the user behavior model, and the input parameter value is the specific information, data, etc. that the user can input, corresponding to the input parameter value corresponding to the user operation. For example, the input operation can be a login operation, inputting a search keyword, clicking on a relevant link, and the input parameter value can be the content that can be input for an input operation, such as a specific keyword, a specific link, etc. The input variable can also be a user identity parameter, such as a username and password. The input variable can also be an environmental parameter, such as user device parameters, the user device operating system, network configuration or conditions, etc. The specific variable values of these input variables form the complete space of the input parameters of the target software. For example, the input variable x1 represents the username, and its value can be "user123" or "admin", etc.; the input variable x2 represents the user operating system, and the parameter value can be "Windows10", "Android12", etc., the input variable x3 represents the login operation, the input variable x4 represents the operation of inputting a search keyword, the input variable x5 represents the keyword, and its value is one or more keywords in the keyword table, etc.
[0124] In the present invention, a test case individual refers to a specific instance for a complete test task, including multiple component parameters, and is the smallest unit in the test case set. In the following description, to highlight the relationship with the test case set, a specific instance for a complete test task is referred to as a test case individual. When the relationship with the test case set is not emphasized, a specific instance for a complete test task is referred to as a test case.
[0125] To achieve a complete test task, there are multiple categories of component parameters that make up a test case, and certain conditions need to be met. For example, the component parameter body for testing should at least include specific input operations and corresponding input parameter values. The input operations can be one or multiple and consecutive, that is, an input operation sequence. Referring to the paths and path branches in the call flow graph, an input operation and its corresponding input parameter value in a test case can implement a transition from one node to another, that is, implement a path branch. When a test case includes multiple consecutive input operations, multiple consecutive path branches can be implemented, that is, a path is formed.
[0126] In addition, the component parameters of a test case can also include environmental parameters, and when necessary, user identity parameters, which can be considered a special type of input parameter.
[0127] In step S103, when generating initial test case individuals to construct an initial test case set, according to the categories and specific variable values of input variables that can be provided in the input parameter space, various user behavior models, and the current test environment, a preset number of initial test case individuals are generated according to the conditions that the composition of the test case needs to meet.
[0128] See Figure 2 , Figure 2 is a flowchart of a method for constructing an initial test case set according to an embodiment of the present invention. In this embodiment, the method for constructing an initial test case set based on the input parameter space, user behavior model, and test environment of the target software includes the following steps:
[0129] Step S1031: Configure the operating environment parameters based on the operating environment simulation configuration parameters obtained from the initialized test environment. Define the operating scenarios of the target software when running test case individuals by configuring the operating environment parameters for the initial test case set. Among them, the operating environment parameters include one or more of the operating system O, network conditions N, and hardware configuration H. For example, a specific operating scenario is obtained by setting the network condition to "10 Mbps with a delay of 200 ms". Specific operating scenarios can be configured for each initial test case individual, or multiple operating scenarios can be configured for the initial test case individuals in the initial test case set, and multiple initial test case individuals are run in each operating scenario. When the operating environment parameters are configured, the test environment is deployed based on the operating environment parameters in the subsequent optimization process.
[0130] Step S1032: Determine the initial path based on the target software structure parameters and user behavior model obtained from the initialized test environment. The user behavior model provided by the present invention has various forms, such as the distribution probability of the aforementioned user operations or input operations, the usage frequency of user operations, specific operation sequences, etc. And since each user operation can usually implement a path branch, the path or path branch is determined with reference to the user behavior model. For example, there are multiple paths corresponding to going from the product home page (node i) to a specific product page (node j). Refer to Figure 3 , Figure 3 which is a schematic diagram of the path structure according to an embodiment of the present invention. Figure 3 Among the paths shown in
[0131] it is possible to implement the path branch from node i to node j through the operation of clicking on a specific product (input operation a), or indirectly implement the path from node i to node j through the operation of searching for a product (input operation b) by combining different search conditions (conditions c, d, etc.) through other nodes (nodes j1, j2, or j3). In one embodiment, first, different paths are formed according to the starting node and the ending node, and then the paths are decomposed into multiple consecutive path branches. The corresponding input operations and corresponding input parameters are determined according to the corresponding path branch conditions, so as to determine the input operation sequence, and then compared with the operation sequence provided by the user behavior model, and a path that best matches the user behavior model is selected as the initial path.
[0132] Step S1033: Determine the input variables in the input parameter space that can implement the initial path. Each path branch has a corresponding branch condition. Through the branch condition, the required input operations and input parameter conditions can be determined, and the input parameter space is queried based on the input operations to obtain the corresponding input variables.
[0132] Step S1034: Select input variable values that can simulate user behavior for each input variable from the input parameter space according to the user behavior model.
[0133] Step S1035: Use the running environment parameters and the values of one or more input variables for the initial path as the component parameters of the test case to form an initial test case individual.
[0134] The user behavior models provided by the present invention have various forms, such as the distribution probability of the aforementioned user operations or input operations, the usage frequency of user operations, specific operation sequences, etc. Therefore, first, determine the form of the user behavior model to be used based on all input variables corresponding to the initial test target. When only one input variable is involved, select the input operation or input parameter value with the highest distribution probability that can be used as the input variable value. When the path is from the login page to the user's personal homepage, since a login variable is required, the corresponding input operation is the login operation, and the corresponding input parameters include two input parameters: username and password. There are multiple usernames and multiple passwords in the input space. At this time, when there are multiple input parameter values, select the username and password with the highest distribution probability as the input parameter values. Thus, the obtained test case individual is: [Click the login button; Input parameters: username "user123" + password "pass456"].
[0135] When there are multiple input operations corresponding to an input variable, select the input operation with the highest distribution probability or the highest usage frequency in the user behavior model. For example, for the aforementioned Figure 3 node i, the corresponding input operations can be "click product operation a" and "search operation b". Referring to the user behavior model, the distribution probability of "search operation b" is the highest, so "search operation b" is selected. Another example is when the determined path is from node i through node j1 to node j, including two consecutive input operations: "search operation" and "click operation". According to the user behavior model: [Search for "smartphone", select the "electronic product" category, and browse products with a price range of "1000 - 2000 yuan"], determine the specific input operation values. Thus, the obtained test case individual is: [Search for "smartphone"; Select the "electronic product" category, price range "1000 - 2000 yuan"; Click "xxx phone"].
[0136] Since the user behavior model also provides the usage frequency of input parameters, further, the input parameters can be selected according to the usage frequency of input parameters. For example, according to the user behavior model, when obtaining the usage frequency of the input parameter that the user views 10 pages on average each time when browsing the product page, add the operation step of browsing 10 product pages to the generated test case. For example, the generated test case simulates that the user uses a "wired network" on a computer with the "Windows 10" system to access the shopping platform and randomly clicks 10 product pages.
[0137] Based on the user behavior model, when the frequency of a user clicking on a product during product browsing is 5 products per such behavior, 4 randomly selected products can be added on this basis, that is, the test case individual is obtained: [Search for "smartphone"; Select the "electronic product" category, and the price range is "1000 - 2000 yuan"; Click on "xxx Phone 1"; Click on "xxx Phone 2"; Click on "xxx Phone 3"; Click on "xxx Phone 4"; Click on "xxx Phone 5"].
[0138] Finally, in order to specify the running scenario of the currently generated test case individual, the above-mentioned set running environment parameters can usually be added, so as to obtain a test case with complete composition parameters, such as [Username: user123; Password: pass456; Search for "smartphone"; Select the "electronic product" category; Select the price range of "1000 - 2000 yuan"; Click on "xxx Phone 2"; Click on "xxx Phone 3"; Click on "xxx Phone 4"; Click on "xxx Phone 5"; Operating system version: Windows10; Network configuration: Wi-Fi, latency 50ms]. Of course, other more detailed running environment parameters such as hardware conditions can also be included.
[0139] The present invention generates a preset number of initial test case individuals according to the foregoing steps to form an initial test case set.
[0140] From the above process of generating the initial test case individual, it can be seen that by guiding the input operations, operation sequences, and selection of input parameters for the test case individual through the user behavior model, the generated test case individual can simulate the real operations of the user, so as to reflect the real usage situation of the target software during testing. In addition, in a further embodiment, when guiding the input operations, operation sequences, and selection of input parameters for the test case individual through the user behavior model, high-frequency or high-distribution probability user input operations and input parameter values are preferentially used as the input operations and input parameter values of the test case, or high-frequency or high-distribution probability input operations and low-frequency or low-distribution probability but high-value user operations or input parameter values are respectively used in a certain proportion, so as to achieve diversification and comprehensiveness of testing.
[0141] Furthermore, when generating a preset number of initial test case individuals, paths in different regions of the test path space are selected as the initial paths, so that the initial test case set can evenly cover different regions for subsequent optimization processing.
[0142] Further, when the security risk attribute flag in the path branch or software unit is obtained during initialization, when determining the initial path in step S1032, select the path branch or software unit with the security risk attribute flag to generate a part of the initial test case individuals, so that the generated test cases can cover some path areas with security risks. Among them, in order to generate test cases that can cover the paths with security risks and test whether they contain security vulnerabilities or security threats, when selecting input parameters in step S1033, select the input parameters based on the trigger conditions of security testing (such as a specific vulnerability). For example, for the login operation, in order to obtain security test cases for SQL injection, select the user name in the corresponding input parameters as "admin' OR '1' = '1". If the output when performing the login operation with this user name is login failure, it means that the system of the target software can prevent SQL injection.
[0143] The present invention further obtains a threat model for security testing during initialization. The threat model is a security module used to analyze and process one or some security vulnerabilities or security threats. The threat model can judge whether a vulnerability is detected based on the trigger condition corresponding to the security vulnerability, and quantify and score or rate the risk of the vulnerability when the vulnerability is detected. For example, the threat model determines whether a test case triggers a vulnerability (such as SQL injection detection returns 1 / 0) through symbolic execution, and combines the CVSS framework to quantify the severity s of the security vulnerability and the possibility a of the vulnerability being exploited through the risk scoring function R(s,a) shown in formula (1-4).
[0144] R(s,a) = s × a (1-4)
[0145] Among them, s represents the vulnerability severity, s ∈ [0, 10], a represents the possibility of the vulnerability being exploited (calculated through Monte Carlo simulation), a ∈ [0, 1], and the product of the two is used as the comprehensive quantified risk value. For example, for a certain XSS vulnerability, R(s,a) = 0.7 × 8.5 + 0.3 × (probability 0.9) ≈ 7.2, and it is determined to be a high-risk according to the corresponding rating standard.
[0146] The types of security vulnerabilities in the present invention include SQL injection, cross-site scripting attack, buffer overflow, etc.
[0147] Further, in order to make the initialized test cases closer to real user behaviors and real operating scenarios, better target the ultimate test goal, and make subsequent iterative optimization easier and converge faster, in another embodiment of the present invention, when constructing the initial test case set, a batch of initially generated test case individuals are used as candidate initial test case individuals, and a maximum value problem function is constructed based on the user behavior model and environmental fitness. By adjusting the candidate initial test case individuals, or adjusting the user behavior model and / or operating environment parameters of the candidate initial test case individuals, the maximum value problem function is made to converge to the maximum value; the candidate initial test case individuals when converging to the maximum value are determined as the initial test case individuals.
[0148] In one embodiment, the maximum value problem function is shown as the following formula (1-5):
[0149]
[0150] Wherein, P(x i ) is the distribution probability of a specific input variable value x i in the test case, that is, the probability of being used, which is calculated according to formula (1-3). n is the total number of various input variable values in all test cases. E(C, v) is the environmental impact value I of the operating environment parameters in each test case on this test case, which is calculated according to formula (1-1). m is the total number of operating environment parameters in the current test case set V. α and β are trade-off coefficients used to balance the influence of the user behavior model and environmental configuration on the initialization process.
[0151] The present invention measures the degree to which the test case set as a whole is close to real user behaviors and real operating scenarios based on the user behavior model and the simulated configuration parameters of the operating environment, and optimizes the test case individual v according to the measurement result.
[0152] See Figure 4 , Figure 4 is a flowchart of a method for optimizing test cases to obtain an initial test case set according to an embodiment of the present invention. In this embodiment, the previously obtained test case individuals are used as candidate initial test case individuals, and the candidate initial test case individuals constitute the candidate initial test case set V.
[0153] Step S1031a, obtain a candidate test case individual from the candidate initial test case set V as the target test case individual.
[0154] Step S1032a, calculate the distribution probability P(x i ) of each input variable value in the target test case individual. Specifically, for each input variable value, its probability is calculated using formula (1-3).
[0155] Step S1033a: Calculate the influence value of the running environment simulation configuration in the target test case individual on this target test case individual. Specifically, take each running environment parameter in the target test case as an environmental factor e, and use formula (1-1) to calculate the environmental influence value I of the current running environment parameter on the target test case individual.
[0156] Step S1034a: Determine whether there are still candidate test case individuals in the candidate initial test case set V. If there are, return to Step S1031a. If there are no candidate test case individuals in the candidate initial test case set V, then execute Step S1035a.
[0157] Step S1035a: Calculate the weighted sum of the distribution probabilities of all input variable values and the environmental influence value I of all current candidate test case individuals.
[0158] Step S1036a: Optimize the candidate test case individuals to obtain new test case individuals. For example, randomly change the input variable values of one or more test cases in the candidate initial test case set V in the direction of increasing the probability and / or randomly change the running environment parameters of one or more test case individuals in the candidate initial test case set V in the way of increasing the environmental influence value I, and use them as the new target test case individuals. For example, apply the genetic algorithm. Take each test case individual as a chromosome in the genetic algorithm, and each component parameter value (such as each input variable value and each running environment configuration parameter) in the test case individual as a gene respectively. During the process of optimizing the test cases, change the input variable values and / or running environment parameters in a test case individual through gene crossover and mutation operations to obtain new test case individuals.
[0159] Step S1037a: Calculate the distribution probability p(x) of each input variable value and the environmental influence value I in the new target test case individual.
[0160] Step S1038a: Calculate the weighted sum of the probabilities P(x i ) of all input parameters and the environmental influence value I of all test case individuals in the current candidate initial test case set V.
[0161] Step S1039a: Determine whether the weighted sum of the current candidate initial test case set V has converged to the maximum value. If it has converged to the maximum value, end. If not, return to Step S1036a. Here, "converging to the maximum value" means that the weighted sum of the current candidate initial test case set V no longer increases, that is, by comparing the size of the current weighted sum with the weighted sum calculated last time, when the weighted sum no longer increases, it is confirmed that it has converged to the maximum value.
[0162] Further, after obtaining the initial test case set, evaluate the test effectiveness of each test case individual. In one embodiment, calculate the fitness score of each test case individual through the following fitness function expression (1-6). In this embodiment, evaluate the test effectiveness through the software metric coverage ability and the ability to discover potential defects.
[0163] f(v) = γ·Coverage(v) + δ·Defects(v) (1-6)
[0164] Coverage(v) is the software metric coverage rate of a single test case individual v, usually a percentage, ranging from (0-1). In one embodiment, the software metric is, for example, a path. When specifically calculating, determine the number of path branches it includes through the path corresponding to the test case individual v, calculate the proportion of the number of path branches of the test case individual v to the total number, and take the path branch proportion as the coverage rate Coverage(v). The software metric can also be, for example, lines of code. Based on the test path corresponding to the test case individual v, determine the relevant software structure, and based on the correspondence between the software structure and the code, determine the number of lines of code. Calculate the proportion of the number of lines of code corresponding to the test case individual v to the total number of lines of code, and take the line of code proportion as the coverage rate Coverage(v). Or, the software metric can also be, for example, a functional module of the software. Calculate the ratio of the number of functional modules covered by the test case individual v to the total number of functional modules, and take the functional module ratio as the coverage rate Coverage(v). For example, if the target software system has 100 functional modules and the test case individual v covers 20 of them, then Coverage(v) = 20%.
[0165] Defects(v) is the number of potential defects that a single test case individual v can reveal, representing the defect detection ability of a single test case individual v. The defect detection ability Defects(v) is a positive integer, and its specific range depends on the test scenario and system complexity. Its calculation method is, for example: through the aforementioned security threat module, use symbolic execution to analyze the potential defect triggering ability of the test case individual v for the target software, such as predicting the type and number of vulnerabilities that can be exposed.
[0166] γ and δ are weight coefficients, used to balance the contributions of the software metric coverage rate and the defect detection ability to the fitness.
[0167] In this embodiment, the weight coefficients can be set as needed. If the coverage rate is given priority, γ > δ can be set. If defect discovery is emphasized, δ > γ can be set.
[0168] For example, in a medium-sized or small software system, for a test case, its Coverage(v) can be set between 0.1 and 0.5 (i.e., 10% - 50%), and its Defects(v) can be set between 0 and 5 (the number of detected defects is limited).
[0169] For example, set γ = 0.6 and δ = 0.4. If the coverage rate of the test case individual v is 0.3 (30%) and 2 defects are detected, then the fitness score is: f(v) = 0.6·0.3 + 0.4·2 = 0.18 + 0.8 = 0.98. This score indicates that the test effectiveness performance of this individual is very good and it is suitable to be retained in the current set for further optimization.
[0170] In a large and complex software system, for a test case individual v, Coverage(v) may be between 0.01 and 0.3. Defects(v) may be between 0 and 10 (it is easier to discover defects).
[0171] For example, set γ = 0.7 and δ = 0.3. If the coverage rate of the test case individual v is 0.1 and 5 defects are detected, then the fitness score is: f(v) = 0.7·0.1 + 0.3·5 = 0.07 + 1.5 = 1.57. This score indicates that the test case individual v is relatively effective.
[0172] The score calculated for each test case individual through expression (1 - 6) is used as the initial fitness score of the test case individual. The higher the score calculated based on the fitness function expression (1 - 6), the better the performance of the test case individual v in terms of both the software function coverage degree and the ability to discover potential defects, that is, the better the test effectiveness, the higher the potential value, and the more effectively it can discover software defects or verify software functions.
[0173] In one embodiment, after the weighted sum of the current candidate initial test case set V converges to the maximum value, the candidate initial test case individuals with initial fitness scores meeting the requirements are determined as the initial test case individuals. For example, sort the test case individuals in the current candidate initial test case set V according to the initial fitness scores, and use a certain number of test case individuals ranked at the front as the initial test case individuals to obtain the final initial test case set. Or, use the test case individuals with initial fitness scores greater than the threshold as the initial test case individuals to obtain the final initial test case set.
[0174] In addition, referring to the evaluation of the test effectiveness of each test case individual, in another embodiment, the test effectiveness of the current test case set can also be evaluated as a whole. Specifically, see the fitness function expression (1 - 7):
[0175] f(V) = γ·Coverage(V) + δ·Defects(V) (1-7)
[0176] Among them, Coverage(V) represents the software metric coverage rate of the test case set V, reflecting the coverage of the test case set V for the target software. Defects(V) represents the number of potential defects that can be revealed by the test case set V, reflecting the effectiveness of the defect detection ability of the test case set. γ and δ are trade-off coefficients used to balance the contributions of the coverage rate and the defect detection ability to the fitness.
[0177] The initial test case set of the present invention has a certain degree of diversity and can initially cover key or common use case scenarios. To a certain extent, it faces or takes into account goals such as function coverage, security testing, and user behavior coverage. Therefore, in subsequent iterative optimization (from step S104 to step S108), combined with fitness evaluation (such as coverage rate, security vulnerability discovery ability, user behavior matching degree, etc.), a combination of genetic algorithm and particle swarm optimization is continuously adopted for these initial test case individuals. In the case of a huge scale of the input parameter space, it can quickly generate test case individuals that meet the test goals, avoiding a sudden expansion of the search space during optimization, thus reducing the search difficulty and also improving the optimization efficiency.
[0178] In step S104, when performing a reproduction operation on the test case individuals in the current test case set based on the genetic algorithm, all the test case individuals in the current test case set can be used as the parent individuals for the reproduction operation, or some test case individuals can be selected from the current test case set as the parent test case individuals. When selecting some parent test case individuals, some test case individuals with high fitness can be selected as the parent test case individuals, or the parent test case individuals can be selected from the current test case set based on the roulette wheel selection mechanism. For example, first calculate the fitness score of each test case individual in the current test case set; then traverse the fitness scores of each test case individual, and use the test case individuals with fitness scores greater than the threshold as the parent test case individuals; or sort the test case individuals in the current test case set in descending order of fitness scores, and use the preset number of test case individuals ranked at the front as the parent test case individuals.
[0179] When the current test case set is the initial test case set, obtain the fitness score calculated according to the foregoing formula (1-6). When the current test case set includes the optimized test case set, first collect the corresponding running data when running each test case individual in the current test case set; then calculate the coverage rate of each test case individual for the target software metrics based on the running data of each test case individual; query the user behavior model data and calculate the matching rate between the user behavior simulated by each test case individual and the user behavior model; evaluate each test case individual based on the threat model for security testing to obtain the defect revelation ability value; finally, calculate the weighted sum of the coverage rate of each test case individual for the target software metrics, the defect revelation ability value, and the matching rate between the user behavior simulated by the test case individual and the user behavior model as the fitness score of each test case individual.
[0180] In one embodiment, calculate the fitness score of each test case individual based on the following formula (1-8):
[0181]
[0182] Among them, covered(i,v), detect(d,v), and match(u,v) are indicator functions respectively. The three terms in formula (1-8) respectively correspond to the coverage criterion, security criterion, and user behavior matching criterion during evaluation, and are respectively used to judge the software metrics covered by the test case individual v (such as lines of code, functions, modules, and path coverage and conditions, etc.), the ability to detect security threats or vulnerabilities, and the matching degree with user behavior. λ1, λ2, and λ3 are the weights for the three criteria respectively.
[0183] Specifically, c is a coverage criterion, C is the set of coverage criteria, and the set of coverage criteria C includes various coverage criteria for lines of code, functions, modules, and path coverage and condition coverage. i is a coverage unit, such as a line of code or a code item, a function, a branch, a module, a path, or a condition, etc., and I c is the set of coverage units for the coverage criterion c, and w c is the weight for the coverage criterion c, |I c | is the total number of coverage units i, and covered(i,v) is an indicator function. If the test case individual v covers the coverage unit i, it is 1, otherwise it is 0.
[0184] s is the type of security testing, S is the set of security testing types, d is a vulnerability or security threat (hereinafter referred to as a security testing item), and D s is the set of security testing items for the security testing type s, and w sis the weight of the security test type s. detect(d, v) is an indicator function that is 1 if the test case individual v can detect the security test item v, and 0 otherwise. T s is the detection threshold of the security test type s, and e is the base of the natural logarithm.
[0185] b is the user behavior pattern, B is the set of user behavior patterns, U b is the set of user operation sequences for the user behavior pattern b, w b is the weight for the user behavior pattern b. p(u) is the probability of the user operation sequence u occurring, and match(u, v) is a matching function that reflects the matching degree between the test case individual v and the user operation sequence u.
[0186] See Figure 5 , Figure 5 is a flowchart of a method for selecting parental test case individuals for reproduction using a roulette wheel selection mechanism according to an embodiment of the present invention. It includes the following steps:
[0187] Step S10421, obtain the fitness score of each test case individual in the current test case set.
[0188] Step S10422, calculate the selection probability p select (v) of each test case individual in the current test case set.
[0189] Step S10423, calculate the cumulative probability s elect (v) of each test case individual.
[0190] Step S10424, generate a random number, and select the corresponding test case individual as the parental test case individual based on the position where the random number is located.
[0191] Step S10425, determine whether the number of selected parental test case individuals has reached the preset number. If it has reached, end. If it has not reached the preset number, return to step S10424.
[0192] Among them, in step S10421, the fitness score of the test case individual can be the initial fitness score calculated based on formula (1-6) as described above or the fitness score calculated based on formula (1-8).
[0193] In step S10422, in one embodiment, the selection probability p select (v) of the test case individual can be calculated by the following formula (1-9):
[0194]
[0195] Among them, pselect (v) represents the probability that the test case individual v is selected for reproduction; p select (v) ∈ [0, 1], and the sum of the selection probabilities of all test case individuals v is 1; F(v) is the fitness score of the test case individual v. In the first optimization, it is the initial fitness score, and in subsequent optimization processes, it is calculated using formula (1 - 8), reflecting the quality of the test case individual v from three aspects: coverage rate, security testing ability, and user behavior matching degree. The higher F(v) is, the better the quality of the test case, and the greater the probability of being selected; Population represents the current test case set, that is, the seed set obtained after the selection in step S1043.
[0196] ∑ v′∈Population F(v) represents the sum of the fitness scores of all test case individuals in the seed set.
[0197] In another embodiment, to enhance the influence of the fitness difference of test case individuals and make the influence of the fitness difference of test case individual v on the selection probability more significant. Calculate the selection probability p of the test case individual based on formula (1 - 10) select (v):
[0198]
[0199] T is a regulation parameter used to adjust the sensitivity of the selection mechanism to fitness differences. When the value of T is larger, the probability distribution is smoother, and individuals with low fitness have a certain probability of being selected, which helps to maintain the species characteristics. When T is smaller, the selection is more concentrated on individuals with high fitness, accelerating convergence, but may lead to premature convergence to a local optimum. By adjusting the parameter T, the smoothness of the selection probability distribution is controlled, allowing individuals with lower fitness scores to also have a certain probability of being selected. Among them, in the selection process of the same round of optimization, the same regulation parameter T is used for all test cases to calculate, so as to ensure the implementation of a consistent selection strategy for the entire population, facilitating the comparison of the fitness of different individuals and the allocation of selection probabilities. For different iteration rounds, the value of T can be dynamically adjusted to enhance or weaken the "amplification" or "smoothing" effect on fitness differences.
[0200] This embodiment passes e F(V) / T operation to exponentiate the fitness score of the test case individual v, enhancing the influence of the fitness difference of test case individuals and making the influence of the fitness difference of test case individual v on the selection probability more significant.
[0201] In step S10423, to calculate the cumulative probability P of the test case individual select(v), queue all test case individuals and label them with serial numbers n, where n = 1... N, and N is the total number of test case individuals in the seed set. For any test case individual v in the queue i , take the test case individual v i 's selection probability p select (v i ) and the sum of the selection probabilities p select (v) of all the test case individuals sorted in front of it as its cumulative probability P select (v i ), that is Thus, N cumulative probability values are obtained, forming a cumulative probability value queue. Then in step S10424, the range of the random number is 1 to N. Use the random number as the sorting of the cumulative probability value queue to obtain the test case individual corresponding to the position indicated by the random number.
[0202] After obtaining the parent test case individuals, the breeding operations performed on the parent test case individuals are, for example, crossover operations and / or mutation operations. Take each test case individual as a chromosome, and take each input variable value and each running environment configuration parameter in the test case as chromosome genes. Change the input variable values and / or running environment parameters in a test case individual through gene crossover and mutation operations to obtain new test case individuals.
[0203] For example, perform crossover operations through the following formula (1-11) and perform mutation operations through the following formula (1-12):
[0204] v(t + 1)' i = crossover(v(t) i , v(t) r1 , cr) (1-11)
[0205] where crossover represents the crossover operation, cr is the crossover rate, v(t + 1)' i represents the test case individual after crossover, v(t) i is the test case individual selected for the crossover operation, and v(t) r1 is another randomly selected test case individual for the crossover operation.
[0206] v(t + 1)″ i = mutate(v(t) i , mr) (1-12)
[0207] where mutate represents the mutation operation, mr is the mutation rate, and v(t + 1)″ i represents the test case individual after mutation, and v(t) iIs the test case individual selected for mutation operation.
[0208] In this embodiment, the number, ratio, position, etc. of genes for crossover or mutation can be specified, so that the specific input variable values (such as user identity parameters, input operations, input parameters, etc.) and / or environmental parameters for crossover or mutation can be determined.
[0209] For input parameters with a parameter value range, the mutation operation of the input parameters can be performed using formula (1-13):
[0210] R mutated = R + μ · (rand([-1,1]) · (max(R) - min(R))) (1-13)
[0211] Wherein, R mutated Represents the mutated input parameter value; R represents the original input parameter value; rand([-1,1]) represents generating a random number within the range of [-1,1]. By introducing randomness in the present invention, the mutation direction and amplitude of each test parameter are different; max(R) and min(R) respectively represent the upper and lower limits of the input parameter value, which are used to control the range of the mutation operation to ensure that the input parameter value does not exceed the acceptable limit. By changing the mutation rate μ, the purpose of diversity control can be achieved. For example: setting a higher μ: making the input parameter span larger. For example, for the username and password in the input parameters, completely different username formats can be tried, while for some numerical parameters, a lower mutation rate μ can be set: fine-tuning near the original parameter. For example, changing the time interval of the login request, fine-tuning the price range of the goods clicked by the user, etc.
[0212] By performing crossover and mutation on input operations, input parameters, user identity parameters, running environment parameters, etc., the present invention can obtain new test case individuals that simulate different user behaviors, meet the trigger conditions for security testing, increase the running scenarios, increase the software functions, code, functions, branches, conditions, paths, etc. coverage.
[0213] For example, an original test case individual is as follows:
[0214] Input parameters: {"user123","pass456"};
[0215] Input operation: Click the navigation bar 3 times after logging in;
[0216] Running environment parameters: The operating system is "Windows10" and the bandwidth is 10Mbps.
[0217] During mutation, the input parameters can be mutated. For example, the username "user123" in the original input data can be mutated to user'OR'1'='1 for testing SQL injection.
[0218] When mutating the input operations, the user behavior patterns can be adjusted. For example, the original operation of clicking the navigation bar 3 times can be mutated to clicking the navigation bar 5 times, and a search operation is added.
[0219] The operating environment parameters are mutated to adjust the environment configuration. For example, the original configuration is a bandwidth of 10 Mbps; after mutation, the bandwidth becomes 5 Mbps, thus simulating a network fluctuation scenario.
[0220] The final mutated test cases are as follows:
[0221] Input data: {"user'OR'1'='1","pass456"};
[0222] Input operations: Click the navigation bar 5 times and perform a search;
[0223] Operating environment parameters: Operating system "Windows10", bandwidth 5 Mbps.
[0224] Therefore, it can be seen that during the mutation operation, by mutating the input parameters (such as username, password), various possible vulnerability trigger conditions can be generated, generating test case individuals that are more likely to trigger security vulnerabilities, thereby expanding the security test coverage. In addition, test case individuals that simulate other user behaviors can also be generated through the mutation operation, making the test process closer to the real user operation habits and improving the user behavior coverage. For example, by changing the input operations in the test cases, new behavior patterns can be generated. For example, when mutating the click count and path, new operation habits can be obtained, increasing the diversity of test cases. Through the mutation operation of the operating environment, the running states of different environments can be simulated, providing diverse test scenarios for the execution of the test.
[0225] The new test case individuals obtained after the genetic algorithm in step S104 form a test case subset. Then in step S105, the test case subset is mapped to an initial particle swarm. During mapping, first, the test cases are mapped to the multi-dimensional position vectors of the particles, and each dimension corresponds to an input variable of the test case. Among them, the position of the particle corresponds to a specific test case individual. For example, the dimension parameters of a particle x1 are configured as [input field length, concurrent request count, timeout threshold], and the dimension parameters of a particle x2 are configured as [username length, password complexity, network latency, operating system].
[0226] Then, position encoding is performed on each particle, that is, the specific dimensional parameter values are encoded as numbers. If the dimensional parameter value is a continuous numerical type, such as network latency, input value, etc., the continuous parameter value is directly used. If the dimensional parameter value is discrete, such as operating system type, operation sequence, it is encoded as a discrete value. When encoding as a discrete value, integer encoding can be performed. For example, for the operating system type, different numbers are used to represent different operating systems, such as: 0 = Windows, 1 = macOS, 2 = Linux, 3 = Android. Multi-segment encoding can also be performed. For example, the operation sequence dimension is divided into multiple sub-dimensions, and each sub-dimension represents the selection of an operation step.
[0227] For example, for a particle x2 with a specific dimensional parameter configuration of [Username: User123; Password: 456789; Network latency: 200ms; Operating system: Android], according to its position encoding rule, the username length is 6 characters and the password complexity is a weak password, and the corresponding encoding is 0. Among them, the encoding of medium-strength password pairs is 1, and the encoding of strong password pairs is 2. The network latency is the corresponding number 200, and the encoding of the operating system is 3. Therefore, the particle x2 [Username: User123; Password: 456789; Network latency: 200ms; Operating system: Android] is encoded as [6; 0; 200; 3], which is the current position of the particle x2.
[0228] Then, determine its position range, speed range, and initial speed according to the dimension parameter type. The position range is, for example, the value range of each dimension parameter. For example, for the first dimension parameter "username length" of particle x2, its value range is 1 - 20. The speed of each particle is a vector, and its dimension corresponds one-to-one with the dimension of the particle's position vector (i.e., the input variable values in the test case individual). Each component of the speed represents the change amount of the input variable value corresponding to the dimension. When the data type of the dimension parameter is continuous numerical type, such as input value, network delay, etc., the speed component is a real number, representing the increase or decrease amplitude of the dimension parameter value. When the range of the input value is [0, 100], the speed range can be set to ±10%, then the maximum adjustment amplitude of the speed component in each iteration is ±10. When the data type of the dimension parameter is discrete numerical type, such as operating system type, operation steps, etc., the speed component needs to be converted into the probability of discrete selection or adjustment strategy. For example, for the operating system type [0 = Windows, 1 = macOS, 2 = Linux, 3 = Android], the probability of selecting the current operating system (such as macOS) is reduced by using the Softmax function or roulette wheel selection method, and the probability of adjacent options (such as Windows) is increased. Another example is to use the method of integer speed conversion to round the speed represented by a decimal to an integer, so as to select the parameter value corresponding to the integer. For example, when v = 0.7, it is converted to v = 1, so as to select the corresponding macOS.
[0229] In step S106, when moving the particle, in one embodiment, first calculate the current moving speed of the particle based on formula (1 - 14).
[0230] v i (t + 1) = w·v i (t) + c1·rand1·(pbest i -x i (t)) + c2·rand2·(τ gbest -x i (t)) (1 - 14)
[0231] Where, v i (t) and x i (t) respectively represent the speed and position of particle i at the t-th optimization, w is the inertia weight, c1 and c2 are learning factors, rand1 and rand2 generate random numbers in the range of [0, 1], pbest i is the individual best position of particle i. τ gbestis the global historical best position of the particle swarm. In one embodiment, after each movement of the particle swarm is completed, the fitness score of the test case individual corresponding to the new position is calculated and recorded. By querying the historical fitness score record of a particle, the position with the highest historical fitness score can be obtained and used as pbest i , query the historical fitness score records of all particles, and use the position with the highest historical fitness score among all particles as the global best position τ of the particle swarm gbest .
[0232] Then, based on formula (1-15), the new position x(t + 1) of the particle is calculated
[0233] x(t + 1) = v(t + 1) + x(t) (1-15)
[0234] That is, the new position x(t + 1) of the particle is obtained by adding the current movement speed v(t + 1) to the current position x(t) of the particle, and the new position x(t + 1) corresponds to a new test case individual
[0235] It can be seen from formula (1-14) that when calculating the new speed of the particle, a part of the current speed is retained through the inertial term of the first item to maintain the historical inertia of the search direction. Through the second item, it is adjusted towards the historical optimal position of the particle itself. Through the third item, it is adjusted towards the global historical optimal position of the group
[0236] As can be seen from the above, the position of the particle is closely related to the particle speed. In order to meet the test cases with different focuses, in one embodiment, the learning factors c1, c2 and / or the inertia weight are as shown in the following formulas (1-16), (1-17) and (1-18), including one or more influencing factors and weights
[0237]
[0238] where g k and They are the k-th impact factor and its weight respectively, and m is the total number of impact factors. The impact factors mentioned above are, for example, one or more of the safety risk factor, environmental impact factor, user behavior matching rate factor, business scenario factor, coverage impact factor, execution time impact factor, and system load impact factor. Therefore, when calculating the new velocity of a particle, if necessary, first calculate the learning factors c1, c2, and / or the inertia weight required in formula (1-14). For example, when the safety risk factor weight is included in the learning factors, query whether the software unit identifier corresponding to the current position in the software structure diagram includes a safety risk identifier. If so, determine the value of the safety risk factor according to the preset weight calculation method for defect priorities based on the type, quantity, etc. of the safety risk identifier. In another embodiment, query the fitness calculation formula of the test case corresponding to the current position, and use the score calculated from the security test standard therein as the value of the safety risk factor. For example, use the value calculated according to the second term in formula (1-8) as the value of the safety risk factor. By incorporating the security test standard into the global search strategy of the Particle Swarm Optimization (PSO) algorithm, the purpose of preferentially searching for modules with higher risks can be achieved. The value calculated as the value of the safety risk factor.
[0239] Similarly, when the environmental impact factor is added to the learning factors c1, c2, and / or the inertia weight, when calculating the particle velocity, the environmental impact value calculated according to formula (1-1) based on the aforementioned fitness function E(C, V) can be used as the environmental impact factor. When the user behavior matching degree is added to the learning factors c1, c2, and / or the inertia weight, the value calculated according to the third term in formula (1-8) can be used as the user behavior matching degree factor. By adding environmental factors and the user behavior model to the calculation process of the particle position, test cases that are more in line with the actual usage scenario can be generated.
[0240] In another embodiment, when evaluating pbest and gbest, one or more of the safety risk factor, environmental impact factor, user behavior matching degree factor, business scenario factor, coverage impact factor, execution time impact factor, and system load impact factor can also be added to the evaluation formula, and test cases for multiple scenarios can be generated accurately and efficiently as well.
[0241] After obtaining new test case individuals through the aforementioned genetic algorithm reproduction and particle movement, they are merged with the test case individuals not selected for reproduction to form a new test case set.
[0242] Then, in step S107, the fitness score F(v) of each test case individual in the new test case set is calculated according to formulas (1-8) respectively to measure its performance in terms of coverage, security testing ability, and user behavior simulation. Then, they are sorted according to the fitness score F(v), and the top N test case individuals with higher fitness scores are retained for the next round of optimization, while the test case individuals with lower fitness scores are eliminated. Alternatively, in another embodiment, the maximum value of the new test case set is obtained through the following formula (1-19), and the test case individual corresponding to the obtained maximum value is used as the updated test case set.
[0243]
[0244] Among them, Population combined is a combination of the original test case set and the newly generated test case individuals after reproduction, and N is the size of the test case set. Population new is the updated test case set.
[0245] The optimization end condition in step S108 is, for example, that the current test case set meets the system test objectives, and the test objectives include, for example, software system test objectives, security test objectives, and user behavior test objectives, and are calculated through the fitness calculation formulas of the corresponding test case sets. For the software system test objectives, for example, specifically determine whether the current test case set covers the specified software internal structure and reaches the specified coverage rate, whether it reaches the code coverage rate, function call coverage rate, and conditional judgment coverage rate of modules or components, and whether it reaches the specified environmental fitness (E(C,V)); for the security test objectives, for example, whether the specified vulnerabilities or threats are detected, or whether a specified number or proportion of vulnerabilities or threats are detected. For the user behavior test objectives, for example, whether the specified user behavior coverage and matching degree are achieved. In one embodiment, the fitness score of the test case set is calculated according to the following formula (1-20):
[0246]
[0247] Among them, C, S, and B respectively represent the set of coverage criteria, security test types, and user behavior patterns of the software system, and w c , w s and w b are the weight coefficients corresponding to each type; c is a coverage criterion, C is the set of coverage criteria, and the set of coverage criteria C includes various coverage criteria for the number of code lines, functions, modules, path coverage, and conditional coverage. Coverage c(V) is the score of the software units (such as the number of code lines, functions, modules, and various coverage criteria such as path coverage and conditions) covered by the test case set V calculated according to formula (1-21).
[0248]
[0249] Among them, I c is the set of covering units for the covering criterion c, |I c | is the total number of coverage unit i, and covered(i,V) is an indicator function that is 1 if the test case set covers coverage unit i, otherwise it is 0.
[0250] s is the safety test type, S is the set of safety test types; Safety s (V) is the score calculated based on the following formula (1-22).
[0251]
[0252] Where d is a vulnerability or security threat (hereinafter referred to as security test item, such as SQL injection, cross-site scripting attack and buffer overflow, etc.), D s is a set of security test items for security test type s. detect(d,v) is an indicator function. If the test case individual v can detect the security test item v, it is 1, otherwise it is 0. T s is the detection threshold of safety test type s, and e is the base of the natural logarithm.
[0253] b is the user behavior pattern, B is the user behavior pattern set; S45, user behavior coverage Behavior b (V) is a score calculated based on formula (1-23), which is used to evaluate the degree of consistency between the test case set and the user behavior model, reflecting the effectiveness of the test case set in simulating real user behavior.
[0254]
[0255] Among them, U b It is a set of user operation sequences for user behavior pattern b, p(u) is the probability of user operation sequence u appearing, and match(u,V) is a matching function that reflects the degree of match between the test case set V and the user operation sequence u. b (V) reflects the user behavior coverage and matching degree, that is, the user behaviors covered by the user behaviors in the test case and the corresponding matching degree.
[0256] In step S108, calculate the fitness score of the updated test case set. When the fitness score of the updated test case set reaches the threshold, it is considered that the optimization end condition is satisfied. At this time, stop the optimization and store the current test case set V final , V final ={V|V∈Population at temination}, and the finally obtained test case set V final contains the test case individuals with the highest fitness after multiple rounds of iterative optimization. In addition, the optimization end condition also includes the number of iterations. When the number of iterations is reached, stop the optimization. When terminating the optimization only by the number of iterations, the finally generated test case set V final can also be evaluated. Calculate the fitness score of the test case set V according to the aforementioned fitness calculation formula final to evaluate whether the predetermined test coverage criteria are met, including code coverage rate, the ability to discover potential security vulnerabilities, and simulating user behavior.
[0257] As can be seen from the foregoing method, the present invention further optimizes on the basis of the new test case individuals obtained by the genetic algorithm, combines the social learning behavior of the particle swarm optimization algorithm with the genetic evolution mechanism of the genetic algorithm, and uses the velocity and position update mechanism of the particles to guide the crossover and mutation operations of the genetic algorithm. It can maintain the population diversity and accelerate the convergence speed under the condition of a huge input parameter space scale, thereby improving the generation efficiency and quality of test cases.
[0258] In addition, in a further embodiment, corresponding to Figure 1, before performing step S104, it also includes a step of scenario classification. For example, the dimension or the number of combinations of input parameters is greater than a threshold, or the number of Cartesian products of the values of all variable parameters is greater than a threshold; the path complexity is not very high, but the requirement for "breadth" coverage of the input parameter space is relatively high: such as the McCabe complexity is less than a threshold and the conventional process / function coverage needs to be satisfied; the time / resources are relatively loose, allowing multiple generations of iteration to explore the huge input space: such as the executable test time is greater than or equal to a threshold. When one or more of the above conditions are met, it can be determined as an input combination explosion scenario, and step S104 is executed. Otherwise, it is judged whether other preset scenarios are satisfied. For example, it is judged whether the complex path coverage scenario is satisfied. If it is satisfied, a method that combines the Ant Colony Optimization (ACO) and the Particle Swarm Optimization (PSO) is used to generate test cases; for another example, it is judged whether the high-risk security test scenario is satisfied. If the high-risk security test scenario is satisfied, a method that combines the Ant Colony Algorithm (ACO) and the Genetic Algorithm (GA) is used to generate test cases. If none of the obvious test scenarios are satisfied, any of the foregoing methods or a method applicable to the general scenario can be used to generate test cases.
[0259] This embodiment adopts a combination of the Genetic Algorithm (GA) and the Particle Swarm Optimization Algorithm (PSO). By combining the social learning behavior of the particle swarm algorithm with the crossover and mutation mechanisms of the genetic algorithm, and using the velocity and position update mechanism of the particles to guide the genetic algorithm operation, it can quickly explore the high-dimensional input parameter space, which can not only improve the population diversity but also accelerate the optimization convergence, effectively solving the problems of difficulty in generating appropriate test cases caused by input combination explosion and slow optimization convergence.
[0260] See Figure 6 , Figure 6It is a principle block diagram of a software test case system based on an intelligent swarm fusion algorithm according to an embodiment of the present invention. The system includes an initialization module 11, a user behavior model construction module 12, an input parameter space construction module 13, an initial set construction module 14, and an optimization module 15. Among them, the initialization module 11 is used to initialize the test environment to obtain the running environment simulation configuration parameters and the structure parameters of the target software. The user behavior model construction module 12 is used to construct a user behavior model based on the collected user historical behavior data. Among them, the user behavior model includes one or more user parameters simulating user behavior and their parameter values. The input parameter space construction module 13 is used to construct the test parameter space of the target software. Among them, the input parameter space includes multiple input variables, and the variable value of each input variable corresponds to one or more user parameter values used to form the component parameters of the test case individual. The initial set construction module 14 constructs an initial test case set based on the input parameter space, user behavior model, and test environment of the target software. Among them, the initial test case set includes a preset number of initial test case individuals, and each initial test case individual includes multiple types of component parameters. The categories of the component parameters at least include input operations and corresponding input parameters and running environment parameters. The optimization module 15 is used to iteratively optimize the test case individuals based on the initial test case set. Specifically, the optimization module includes a reproduction unit 151, a particle swarm optimization unit 152, and an evaluation unit 153. Among them, the reproduction unit 151 is used to take the test case individual as a chromosome and the component parameters of the test case as genes, and perform a reproduction operation on the test case individuals in the current test case set to obtain new test case individuals. Among them, in the first optimization, the current test case set is the initial test case set; the particle swarm optimization unit 152 can map the test case individual to a particle, take the new test case individual obtained by reproduction as the current position of the corresponding particle, and calculate the new moving speed of the particle based on the current position of the particle, the current speed, the individual best position of the particle, and the global best position; superimpose the new moving speed on the current position of the particle to obtain the new position of the particle. The new position of the particle corresponds to a new test case individual. The new test case individuals and the test case individuals that have not undergone the reproduction operation are combined together to form a new test case individual set. The evaluation unit 153 is used to evaluate whether the new test case set meets the optimization end condition. When the new test case set meets the optimization end condition, stop the iterative optimization and store the new test case set; when the new test case set does not meet the optimization end condition, trigger the reproduction unit 151 to start a new round of optimization.
[0261] Figure 7It is a schematic structural diagram of a software testing system according to an embodiment of the present invention. The software testing system includes an operation terminal and a server terminal. Among them, the operation terminal is located on the terminal device 102, and the server terminal is located on the server 104 or a server cluster. The terminal device 102 communicates with the server 104 through a network. The terminal device 102 includes a desktop computer, a laptop computer or a mobile intelligent terminal device, such as a mobile phone, a tablet computer, etc. The server terminal includes the Figure 6 software test case generation system shown above. The operation terminal includes an interactive interface. When it is necessary to generate test cases, the tester configures corresponding parameters through the interactive interface and sends them to the server terminal through the network. The server terminal generates a test case set according to the received instruction and the corresponding parameters according to the method described above in the present invention. The parameters configured by the tester through the interactive interface are, for example, the target software name, version, user historical data storage address, number of iterations, storage address of the test case set, etc. After the test case set is generated, the tester can send a test instruction to the server terminal through the interactive interface. After receiving the test instruction, the server terminal executes the software test using the test case set and records the test results.
[0262] Figure 8 It is a schematic diagram of the hardware structure principle of an electronic device according to an embodiment of the present invention. The electronic device can be implemented as a server or various other terminal devices, such as a desktop personal computer, a tablet computer, a laptop computer, a mobile phone, etc. It includes a processor 601 and a memory 602. A program instruction set is stored on the memory 602. When the processor 601 executes the program instruction set on the memory 602, the above-mentioned software test case generation method is implemented.
[0263] Specifically, the above-mentioned processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0264] The memory 602 may include a mass storage for data or instructions. By way of example and not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 602 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 602 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 602 is a non-volatile solid state memory.
[0265] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the software test case generation method provided by the present invention.
[0266] In one example, the electronic device may further include a communication interface 603 and a bus 604. The processor 601, the memory 602, and the communication interface 603 are connected through the bus 604 to complete communication with each other.
[0267] The communication interface 603 is mainly used to implement communication between each module, device, unit, and / or device in the embodiments of the present invention.
[0268] The bus 604 includes hardware, software, or both, and couples the components of the online data flow charging device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 604 may include one or more buses. Although the embodiments of the present invention describe and illustrate specific buses, the present invention contemplates any suitable bus or interconnect.
[0269] The present invention also provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, they can implement any one of the software test case generation methods in the foregoing embodiments. The computer-readable storage medium may be any medium that is tangible and contains or stores computer-executable instructions for use by or in connection with an instruction execution system, apparatus, and device. The storage medium may be a transitory computer-readable storage medium or a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media may include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Corresponding embodiments of such storage devices include, for example, magnetic disks, optical discs based on CD, DVD, or Blu-ray technology, and persistent solid-state memories such as flash memories, solid-state drives, and the like.
[0270] The present invention also provides a computer program product, which includes a set of computer program instructions. When the set of computer program instructions is executed by a processor, it implements any one of the software test case generation methods in the foregoing embodiments. The computer program product includes, but is not limited to, application installation packages, application plugins, applets that can run in certain applications, etc. published on websites and application stores.
[0271] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0272] The above embodiments are only for illustrating the present invention, rather than limiting the present invention. Those of ordinary skill in the relevant technical field can make various changes and variations without departing from the scope of the present invention. Therefore, all equivalent technical solutions should also fall within the scope of the disclosure of the present invention.
Claims
1. A method for generating software test cases, characterized in that: include: Initialize the test environment to obtain the simulation configuration parameters of the operating environment and the structural parameters of the target software; A user behavior model and an input parameter space of the target software are constructed based on the collected user historical behavior data, wherein the user behavior model includes one or more user parameters and parameter values simulating user behavior; the input parameter space includes a plurality of input variables, and the variable value of each input variable corresponds to one or more user parameter values used to constitute individual component parameters of the test case; Constructing an initial test case set based on the input parameter space of the target software, the user behavior model and the test environment, wherein the initial test case set includes a preset number of initial test case individuals, each of which includes a plurality of component parameters, and the categories of the component parameters include at least input operations and corresponding input parameter values and operating environment parameters; and Based on the initial test case set, the following iterative optimization steps are performed on the individual test cases: Taking the test case individual as a chromosome and the composition parameters of the test case as genes, breeding the test case individuals in the current test case set to obtain new test case individuals, wherein, during the first optimization, the current test case set is the initial test case set; Mapping individual test cases to particles, taking new individual test cases obtained by breeding as current positions of corresponding particles, and calculating new moving speeds of the particles based on current positions, current speeds, individual optimal positions of the particles, and global optimal positions of the particles; superimposing the new moving speed on the current position of the particle to obtain a new position of the particle, the new position of the particle corresponds to a new test case individual, and the new test case individual is merged with the test case individual that has not been propagated to form a new test case set; and Evaluate whether the new test case set meets the optimization end condition, stop iterative optimization and store the new test case set when the new test case set meets the optimization end condition; repeat the aforementioned iterative optimization steps for the new test case set when the new test case set does not meet the optimization end condition.
2. The software test case generation method according to claim 1, characterized in that: The steps to construct an initial set of test cases based on the input parameter space of the target software, the user behavior model, and the test environment include: Configure the operating environment parameters of the initial test case individuals based on the operating environment simulation configuration parameters obtained by initializing the test environment; Determine the initial path based on the test path space constructed based on the target software structure parameters and the user behavior model; Determine the input variables that can realize the initial path from the input parameter space; Selecting, for each input variable, an input variable value that can simulate user behavior from the input parameter space according to the user behavior model; and The operating environment parameters and one or more input variable values for implementing the initial path are respectively used as component parameters of the test case to form an initial test case individual.
3. The software test case generation method according to claim 2, characterized in that: The operating environment parameters and one or more input variable values for implementing the initial path are respectively used as component parameters of the test case to form a candidate initial test case individual. After obtaining multiple candidate initial test case individuals, the method further includes: Construct the maximum problem function of the initial test case set based on the environmental fitness and user behavior model; Calculate the environmental fitness score of each candidate initial test case individual based on the operating environment parameters of the candidate initial test case individual; Calculate the user behavior simulation score of each candidate initial test case individual based on the user behavior model when determining the input variable value; Calculating the maximum problem function value of the initial test case set based on the environmental fitness score and the user behavior simulation score of each candidate initial test case individual, and adjusting the candidate initial test case individual operating environment parameters and / or the user behavior model to make the maximum problem function converge to a maximum value; and The candidate test case individuals that converge to the maximum value are determined as the initial test case individuals, and all initial test case individuals constitute the initial test case set.
4. The software test case generation method according to claim 1, characterized in that: Taking the test case individual as a chromosome and the composition parameters of the test case as genes, the steps of performing a breeding operation on the test case individuals in the current test case set to obtain a new test case individual include: Select some test case individuals from the current test case set as parent test case individuals; and A crossover operation and / or a mutation operation is performed on the genes in the parent test case individual to obtain a new test case individual.
5. The method for generating software test cases according to claim 4, characterized in that: The steps of selecting some test case individuals from the current test case set as parent test case individuals include: Get the fitness score of each test case in the current test case set; and Traverse the fitness scores of each test case individual, and take the test case individual whose fitness score is greater than the threshold as the parent test case individual; Alternatively, the test case individuals in the current test case set are sorted in descending order of fitness scores; a preset number of test case individuals in the front order are used as parent test case individuals; Alternatively, a parent test case individual is selected from the current test case set based on a roulette wheel selection mechanism, wherein when selecting a parent test case individual from the current test case set based on the roulette wheel selection mechanism, the selection probability of the tested case individual is calculated based on the fitness score of the test case individual.
6. The method for generating software test cases according to claim 5, characterized in that: When the current test case set is the initial test case set, the steps of obtaining the fitness score of each individual test case in the current test case set include: Calculate the coverage of each initial test case individual on the target software indicator; Evaluate each individual initial test case based on the threat model for security testing to obtain a defect disclosure capability value; and Calculate the weighted sum of the coverage rate and defect revealing capability value of each initial test case individual for the target software indicator as the initial fitness score of the initial test case individual; When the current test case set includes an optimized test case set, the step of obtaining the fitness score of each individual test case in the current test case set includes: Run each test case in the current test case set and collect the corresponding running data; Calculate the coverage of the target software indicator by the individual test case based on the running data of each individual test case; Evaluate each test case based on the threat model for security testing to obtain the defect disclosure capability value; Querying the user behavior model data and calculating the matching rate between the user behavior simulated by each test case individual and the user behavior model; and The fitness score of each test case is calculated as the weighted sum of the coverage of the target software indicator, the defect revealing capability value, and the matching rate between the user behavior simulated by the test case and the user behavior model.
7. The method for generating software test cases according to claim 4, characterized in that: When performing mutation operations on genes in the parent test case individuals, input variable values corresponding to the mutated gene categories are obtained from the input parameter space in a manner that simulates new user behaviors and / or triggers security risks.
8. The method for generating software test cases according to claim 1, characterized in that: When calculating the new moving speed of the particle, the current moving speed of the particle is adjusted based on one or more of security risks, environmental impact, user behavior matching, business scenarios, coverage impact, execution time impact, and system load impact.
9. The software test case generation method according to claim 1, characterized in that: After obtaining a new test case individual, further including: Calculate the fitness score of the new test case individual; and Traverse the fitness scores of each new test case individual, retain the new test case individuals whose fitness scores are greater than or equal to the threshold, and eliminate the new test case individuals whose fitness scores are less than the threshold; Alternatively, the new test case individuals are sorted in descending order of fitness scores; a preset number of new test case individuals that are sorted first are retained; Alternatively, a maximum fitness score problem is constructed, the maximum fitness score is obtained based on the current test case set, and the test case individuals with the maximum fitness score are retained.
10. The software test case generation method according to claim 4, characterized in that: The optimization termination condition is that a threshold number of iterations is reached or a fitness score of a new test case set reaches a threshold value.
11. A software test case generation system, characterized in that: include: an initialization module configured to initialize the test environment to obtain operating environment simulation configuration parameters and structural parameters of the target software; A user behavior model building module, configured to build a user behavior model based on the collected user historical behavior data, wherein the user behavior model includes one or more user parameters and parameter values simulating user behavior; An input parameter space construction module configured to construct a test parameter space of the target software, wherein the input parameter space includes a plurality of input variables, and a variable value of each input variable corresponds to one or more user parameter values for constituting individual component parameters of the test case; an initial set construction module, configured to construct an initial test case set based on the input parameter space of the target software, the user behavior model and the test environment, wherein the initial test case set includes a preset number of initial test case individuals, each of which includes multiple types of component parameters, and the types of component parameters at least include input operations and corresponding input parameters and operating environment parameters; and An optimization module, configured to iteratively optimize individual test cases based on an initial test case set; the optimization module comprises: a reproduction unit configured to use the test case individual as a chromosome and the composition parameters of the test case as genes, and to perform a reproduction operation on the test case individual in the current test case set to obtain a new test case individual, wherein, during the first optimization, the current test case set is an initial test case set; A particle swarm optimization unit is configured to map the test case individuals to particles, use the new test case individuals obtained by breeding as the current position of the corresponding particle, calculate the new moving speed of the particle based on the current position, current speed, individual optimal position of the particle and global optimal position of the particle; superimpose the new moving speed on the current position of the particle to obtain a new position of the particle, the new position of the particle corresponds to a new test case individual, and the new test case individual is merged with the test case individual that has not been bred to form a new test case individual set; and The evaluation unit is configured to evaluate whether the new test case set meets the optimization end condition, stop iterative optimization and store the new test case set when the new test case set meets the optimization end condition; when the new test case set does not meet the optimization end condition, trigger the breeding unit to start a new round of optimization.
12. An electronic device, characterized in that: The electronic device comprises a processor and a memory storing computer program instructions; when the electronic device executes the computer program instructions, the method for generating software test cases as described in any one of claims 1 to 10 is implemented.
13. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the software test case generation method as described in any one of claims 1 to 10 is implemented.
14. A computer program product, characterized in that It includes a computer program instruction set, which, when executed by a processor, implements the software test case generation method according to any one of claims 1 to 10.
Citation Information
Patent Citations
A combined approach to accelerate test case generation using genetic methods and symbolic execution.
CN109344057B