Software test case generation method and system, electronic equipment and storage medium
By using the Ant algorithm and user behavior model in software testing, a test case can be generated that can truly reflect the software usage situation, which solves the problem of difficulty in revealing security vulnerabilities in the existing technology and improves the coverage and depth of test cases.
Patent Information
- Application Number
- CN202510476879.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-20
AI Technical Summary
It is difficult for the prior art to generate test cases that can truly reflect the use of software, especially in high-risk security scenarios, which are difficult to reveal security vulnerabilities.
By initializing the test environment, building a user behavior model and input parameter space, ants are driven to explore paths in the path exploration space, update path branch pheromones, and iteratively optimized based on pheromone values and fitness scores to generate an initial test case set.
The generated test cases can truly reflect the software usage, improve the ability to reveal security vulnerabilities, and ensure the coverage and depth of test cases.
Smart Images

Figure CN120179561A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer software testing, and particularly to a method, a system, an electronic device and a storage medium for generating software test cases. Background Art
[0002] In the field of software engineering, software testing is a process for discovering and fixing defects in software, verifying whether the software meets the design requirements, and ensuring its quality and performance. It is a crucial link in the software development process. Test cases are the core part of software testing, directly related to the efficiency and coverage of testing, and affecting the quality and reliability of software products. A test case can be regarded as a specific test task, including elements such as input, operation, and expected output, and is a data set in various forms, such as execution paths, input data, and execution conditions.
[0003] The methods for generating test cases can generally be divided into the following several types:
[0004] First, test cases are manually written by testers. This means that testers need to have very rich test experience and relatively high professional levels, and there are certain drawbacks such as blindness, high cost, and difficulty in improving test coverage. Although the method of manually writing test cases is effective in some cases, with the continuous increase in the complexity of software systems, it has become increasingly difficult to write test cases manually, time-consuming and prone to missing errors, especially in the case of high coverage and in-depth detection. In addition, the manual method is difficult to adapt to the rapid iteration of software. Updating a function may require rewriting a large number of test cases, which significantly increases the cost of software development and maintenance.
[0005] Second, certain tools are used to automatically generate test cases based on certain methods, such as random testing, model-based testing (MBT), symbolic execution, search-based testing, etc. Although these tools have made progress in some fields, there are still many unsatisfactory aspects. For example, model-based testing requires detailed model design, which is itself a complex and time-consuming process; symbolic execution and search-based testing face the problem of state space explosion and are difficult to handle large software systems.
[0006] III. Combine with other algorithms. For example, adopt a genetic algorithm or a method that combines a genetic algorithm with other automated methods. For example, Chinese Patent with publication number CN109344057B and invention title "Combined Accelerated Test Case Generation Method Based on Genetic Method and Symbolic Execution" discloses a test case generation method that combines a genetic algorithm. However, the prior art often ignores the diversity of user behavior patterns and software operating environments, which makes the generated test cases unable to comprehensively reflect the real usage scenarios. Especially in high-risk security scenarios, it is difficult to generate test cases that can effectively reveal security vulnerabilities. Summary of the Invention
[0007] In view of the technical problems existing in the prior art, the present invention proposes a software test case generation method, system, electronic device and storage medium, which can generate test cases that can truly reflect the software usage scenarios and improve the ability to reveal security vulnerabilities.
[0008] To solve the above technical problems, according to one aspect of the present invention, the present invention provides a software test case generation method, including:
[0009] Initialize the test environment to obtain the path exploration space of the target software; wherein, the path exploration space includes nodes and directed edges between the nodes, and two nodes with a directed relationship form a path branch, and each path branch includes a corresponding path branch condition;
[0010] Construct a user behavior model and the input parameter space of the target software based on the collected user historical behavior data, wherein the user behavior model includes one or more user parameters simulating user behavior and their parameter values; the input parameter space includes multiple input variables, and the variable value of each input variable corresponds to one or more user parameter values;
[0011] Drive multiple ants to perform path exploration in different regions of the path exploration space respectively, and update the path branch pheromone during the path exploration process; after each ant determines a path branch during the path exploration process, based on the path branch condition, select the corresponding input variable value from the input parameter space according to the user behavior model; after each ant finishes the path exploration, obtain the input variable values of all path branches corresponding to the explored path to form an initial test case individual, and all initial test case individuals form an initial test case set;
[0012] Perform the following iterative optimization steps based on the initial test case set:
[0013] Screen the parent test case individuals from the current test case set based on the pheromone value and the fitness score of the test case individuals; the current test case set at the first optimization is the initial test case set;
[0014] Performing a reproduction operation based on a parent test case individual to obtain a new test case individual, wherein the new test case individual is merged into the current test case set to obtain a new test case set;
[0015] Updating the pheromone value of the path branch covered by the parent test case individual, and updating the pheromone value of the path branch covered by the reproduced test case individual;
[0016] Evaluating whether the new test case set meets the optimization end condition, stopping the iterative optimization and storing the new test case set when the new test case set meets the optimization end condition; repeating the foregoing iterative optimization steps for the new test case set when the new test case set does not meet the optimization end condition.
[0017] Optionally, the steps for each ant to determine a path branch during path exploration include:
[0018] Based on the pointing relationship of the edge between two nodes, determining one or more candidate nodes that form a path branch with the current node;
[0019] Calculating the selection probability from the current node to each candidate node based on the path branch pheromone and the heuristic factor; and
[0020] Determining the candidate node with the highest selection probability as the next node that forms a path branch with the current node; wherein the heuristic factor at least includes the user behavior weight and the security risk weight determined based on the user behavior model.
[0021] Optionally, the heuristic factor further includes one or more of the following weights: detection execution time weight, running environment weight, and path length weight.
[0022] Optionally, the steps for updating the pheromone value of the path branch during path exploration include:
[0023] During path exploration, determining whether the currently explored path branch is marked with a security risk attribute mark, and in response to the currently explored path branch being marked with a security risk attribute mark, increasing the corresponding increment based on the current pheromone value of the path branch marked with the security risk attribute mark.
[0024] Optionally, each ant further includes during path exploration:
[0025] Determining the scenario corresponding to the current path exploration space region; and
[0026] Adjusting the pheromone value of the path branch in a manner matching the scenario.
[0027] Optionally, the software test case generation method further includes: calculating the coverage rate of the explored paths for the path exploration space or a preset local path exploration space during the process of driving multiple ants to perform path exploration in different regions of the path exploration space, and adjusting the pheromone values of the path branches in the path exploration space based on the corresponding relationship between the coverage rate and the pheromone being negatively correlated.
[0028] Optionally, after each ant finishes path exploration, the input variable values of all path branches corresponding to the explored path form a candidate initial test case individual, and the method further includes:
[0029] Constructing a maximum problem function for the initial test case set based on the environmental fitness and the user behavior model;
[0030] Configuring operating environment parameters for each candidate initial test case individual;
[0031] Calculating the environmental fitness score of each candidate initial test case individual based on the operating environment parameters of the candidate initial test case individual;
[0032] Calculating the user behavior simulation score based on the user behavior model used when generating the candidate initial test case individual;
[0033] Calculating the function value of the maximum problem function for the initial test case set based on the environmental fitness score and the user behavior simulation score of each candidate initial test case individual, and adjusting the operating environment parameters of the candidate initial test case individual and / or the user behavior model used to generate the candidate initial test case individual to make the maximum problem function converge to the maximum value; and
[0034] Determining the candidate initial test case individual when converging to the maximum value as the initial test case individual, and all the initial test case individuals form the initial test case set.
[0035] Optionally, the step of screening parent test case individuals from the current test case set based on the pheromone value and the fitness of the test case individual includes:
[0036] Obtaining the fitness score of each test case individual in the current test case set;
[0037] Obtaining the total pheromone value of the path branches covered by each test case individual in the current test case set to obtain the pheromone value of each test case individual;
[0038] Calculating the selection probability of each test case individual in the current test case set based on the pheromone value and the fitness score;
[0039] Use the test case individuals with a selection probability greater than the threshold as the parent test case individuals selected from the current test case set; or sort the test case individuals in descending order of the selection probability, and use the preset number of test case individuals ranked at the front as the parent test case individuals selected from the current test case set.
[0040] Optionally, when the current test case set is the initial test case set, the steps of obtaining the fitness score of each test case individual in the current test case set include:
[0041] Calculate the coverage rate of each initial test case individual for the target software metric;
[0042] Evaluate each initial test case individual based on the threat model for security testing to obtain the defect revelation ability value; and
[0043] Calculate the weighted sum of the coverage rate of each initial test case individual for the target software metric and the defect revelation ability value as the initial fitness score of the initial test case individual;
[0044] When the current test case set is a test case set that has been optimized, the steps of obtaining the fitness score of each test case individual in the current test case set include:
[0045] Run each test case individual in the current test case set and collect the corresponding running data;
[0046] Calculate the coverage rate of the test case individual for the target software metric based on the running data of each test case individual;
[0047] Evaluate each test case individual based on the threat model for security testing to obtain the defect revelation ability value;
[0048] Query the user behavior model data and calculate the matching rate between the user behavior simulated by each test case individual and the user behavior model; and
[0049] Calculate the weighted sum of the coverage rate of each test case individual for the target software metric, the defect revelation ability value, and the matching rate between the user behavior simulated by the test case individual and the user behavior model as the fitness score of each test case individual.
[0050] Optionally, the steps of performing a reproduction operation based on the parent test case individuals further include:
[0051] Calculate the selection probability of each parent test case individual in the parent test case set;
[0052] Calculate the cumulative probability of each parent test case individual in the parent test case set;
[0053] Generate random numbers; and
[0054] Select the parent test case individuals corresponding to the random number indication positions for reproduction operations.
[0055] The method according to claim 1 or 10, characterized in that the reproduction operation includes a crossover operation and / or a mutation operation.
[0056] Optionally, when performing a mutation operation based on the parent test case individuals, obtain the input variable values corresponding to the mutation gene categories from the input parameter space in a manner that simulates new user behaviors and / or triggers security risks.
[0057] Optionally, after performing a reproduction operation based on the parent test case individuals to obtain new test case individuals, it further includes:
[0058] Calculate the fitness scores of the new test case individuals;
[0059] Traverse the fitness scores of each new test case individual, retain the new test case individuals whose fitness scores are greater than or equal to the threshold, and eliminate the new test case individuals whose fitness scores are less than the threshold;
[0060] Alternatively, sort the new test case individuals in descending order of fitness scores; retain the preset number of new test case individuals ranked at the front;
[0061] Alternatively, construct a fitness score maximum value problem; obtain the maximum fitness score based on the new test case set, and retain the test case individuals when the maximum fitness score is obtained.
[0062] Optionally, the end condition is reaching the iteration number threshold or the fitness score of the current test case set reaching the threshold.
[0063] According to another aspect of the present invention, the present invention also provides a software test case generation system, including:[[]]
[0064] An initialization module configured to initialize the test environment to obtain the path exploration space of the target software; wherein, the path exploration space includes nodes and directed edges between the nodes, and two nodes with a directed relationship form a path branch, and each path branch includes corresponding path branch conditions;
[0065] A user behavior model construction module configured to construct a user behavior model based on the collected user historical behavior data, and the user behavior model includes one or more user parameters simulating user behaviors and their parameter values;
[0066] An input parameter space construction module, configured to construct a test parameter space for a target software, where the input parameter space includes a plurality of input variables, and the variable value of each input variable corresponds to one or more user parameter values;
[0067] An initial set construction module, configured to drive multiple ants to perform path exploration in different regions of a path exploration space respectively, and update pheromone information of path branches during the path exploration; after each ant determines a path branch during the path exploration, based on the path branch condition, obtain the corresponding input variable value from the input parameter space according to the user behavior model; after each ant finishes the path exploration, obtain the input variable values of all path branches corresponding to the explored path to form an initial test case individual, and all the initial test case individuals form an initial test case set;
[0068] An optimization module, configured to perform iterative optimization based on the initial test case set, and the optimization module includes:
[0069] A selection unit, configured to screen parent test case individuals from the current test case set based on the pheromone value and the fitness score of the test case individuals; the current test case set at the first optimization is the initial test case set;
[0070] A reproduction unit, configured to perform a reproduction operation based on the parent test case individuals to obtain new test case individuals, and merge the new test case individuals into the current test case set to obtain a new test case set;
[0071] A pheromone update unit, configured to update the pheromone value of the path branches covered by the parent test case individuals, and update the pheromone value of the path branches covered by the reproduced test case individuals; and
[0072] An evaluation unit, configured to evaluate whether the new test case set meets the optimization end condition, stop the iterative optimization and store the new test case set when the new test case set meets the optimization end condition; trigger the selection unit to start a new round of optimization when the new test case set does not meet the optimization end condition.
[0073] According to another aspect of the present invention, the present invention further provides an electronic device, where the electronic device includes a processor and a memory storing computer program instructions; when the electronic device executes the computer program instructions, the foregoing software test case generation method is implemented.
[0074] According to another aspect of the present invention, the present invention further provides a computer-readable storage medium, where computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the foregoing software test case generation method is implemented.
[0075] According to another aspect of the present invention, the present invention also provides a computer program product, including a set of computer program instructions, which, when executed by a processor, implement the foregoing software test case generation method.
[0076] The test cases generated by the present invention can truly reflect the software usage situation and improve the ability to reveal security vulnerabilities. Description of the Drawings
[0077] Next, the preferred embodiments of the present invention will be further described in detail with reference to the accompanying drawings, where:
[0078] Figure 1 is a flowchart of a software test case generation method according to an embodiment of the present invention;
[0079] Figure 2 is a schematic diagram of a path structure according to an embodiment of the present invention;
[0080] Figure 3 is a flowchart of a method for optimizing an initial test case individual to obtain an initial test case set according to an embodiment of the present invention;
[0081] Figure 4 is a flowchart of a method for selecting parent test case individuals for reproduction using a roulette wheel selection mechanism according to an embodiment of the present invention;
[0082] Figure 5 is a schematic block diagram of a software test case generation system according to an embodiment of the present invention;
[0083] Figure 6 is a schematic diagram of the architecture of a software test system according to an embodiment of the present invention; and
[0084] Figure 7 is a schematic diagram of the hardware structure principle of an electronic device according to an embodiment of the present invention. Detailed Embodiments
[0085] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0086] In the following detailed description, reference may be made to the various specification drawings that form a part of the present application and illustrate specific embodiments of the present application. In the drawings, like reference numerals generally describe substantially similar components in different figures. The various specific embodiments of the present application are described in sufficient detail below so that those of ordinary skill in the relevant art and technology can implement the technical solutions of the present application. It should be understood that other embodiments may also be utilized or structural, logical, or electrical changes may be made to the embodiments of the present application.
[0087] Figure 1 It is a flowchart of a software test case generation method according to an embodiment of the present invention. In this embodiment, the software test case generation method includes the following steps:
[0088] Step S101, initialize the test environment to obtain the path exploration space of the target software. Among them, the path exploration space includes nodes and directed edges between the nodes. Two nodes with a directed relationship form a path branch, and each path branch includes corresponding path branch conditions.
[0089] Step S102, construct a user behavior model and the input parameter space of the target software based on the collected user historical behavior data. Among them, the user behavior model includes one or more user parameters simulating user behavior and their parameter values; the input parameter space includes multiple input variables, and the variable value of each input variable corresponds to one or more user parameter values.
[0090] Step S103, drive multiple ants to perform path exploration in different regions of the path exploration space to obtain an initial test case set. Among them, the path branch pheromone is updated in real time during the path exploration process. After each ant determines a path branch during the path exploration process, it selects the corresponding input variable from the input parameter space based on the path branch condition, and then selects the corresponding input operation and input parameter value from the input parameter space according to the user behavior model as the corresponding input variable value. And so on, until the termination condition of the path exploration is reached, such as reaching a preset path length (such as at most 10-step operations), or entering a termination state (such as triggering a crash, completing the core process), or covering the target branch or meeting the coverage threshold, etc.
[0091] After each ant finishes the path exploration, obtain the input variable values of all path branches corresponding to the explored path to form an initial test case individual, and all initial test case individuals form an initial test case set.
[0092] Step S104, screen the parent test case individuals from the current test case set based on the pheromone value and the fitness score of the test case individuals. Among them, the test case set in the first optimization is the initial test case set.
[0093] Step S105: Perform a breeding operation based on the parent test case individual to obtain a new test case individual. Among them, incorporate the new test case individual into the current test case set to obtain a new test case set.
[0094] Step S106: Update the pheromone values of the path branches, including updating the pheromone values of the path branches covered by the parent test case individual and the pheromone values of the path branches covered by the test case individuals obtained through breeding.
[0095] Step S107: Evaluate whether the new test case set meets the optimization end condition. If the new test case set meets the optimization end condition, then in Step S108, store the new test case set and stop the iterative optimization; if the new test case set does not meet the optimization end condition, in Step S109, set the new test case set as the current test case set and return to Step S104 for a new round of iterative optimization.
[0096] In Step S101, in one embodiment, when initializing the test environment, first obtain the dynamic call graph matrix of the target software, which uses the matrix representation method M(t) in graph theory to describe the interaction and dependency relationships between software units. Specifically, the interaction and dependency relationships between software units are described by Expression (1-1):
[0097] M(t) = M(t - 1) + ΔM(t) (1-1)
[0098] M(t) represents the call matrix of software units within time period t, M(t - 1) represents the call matrix of software units within time period t - 1, and ΔM(t) represents the change in call data within time period t. The dynamic changes of the software during operation are represented by the matrix of the call data of software units.
[0099] Among them, software units, such as components, modules, and even more detailed functions, etc., serve as the rows and columns of the call matrix, and the matrix elements represent the call data between software units within a preset time period, such as the number of times or frequency. In one embodiment, the matrix element M(t) [i,j] represents the number of times or frequency that the i-th software unit calls (jumps / calls) the j-th software unit within time period t. For example, if "Component A" and "Component B" are mapped to the i-th row / column and the j-th column / row, then M(t) [i,j] is the number of times Component A calls Component B within time period t. When further refining "component" to the "function" level, M(t) [f1,f2] represents the number of times function f1 calls function f2 within time period t.
[0100] Over time, M(t), M(t+1), etc. can be combined or superimposed to obtain the overall call matrix M_total. Thus, it can be seen that the call graph matrix not only includes the structural composition of the software, but also can obtain whether the call relationships of certain components / modules / functions are covered, as well as their call frequencies and control flows.
[0101] By parsing the dynamic call graph matrix to obtain two software units corresponding to each matrix element; using the software units as nodes and the call relationships between the software units as edges to construct a call flow graph, the call flow graph includes nodes and edges with directions between the nodes, where two nodes connected by an edge with a direction relationship form a path branch. According to the call relationships between the parsed software units, obtain the input operations and corresponding input parameter conditions for implementing the path branch, hereinafter referred to as branch conditions.
[0102] Specifically, query the elements greater than 0 in the current call graph matrix. If the count of a certain row and a certain column in M(t) is greater than 0, it means that the two software units represented by the corresponding row and column have indeed been called during the test. Therefore, it can be obtained that these two software units have "triggered" the corresponding function call actions during the test. When modeling the internal structure of the software, the two software units that have "triggered" the call actions are used as nodes, and the directed call relationship of "software unit i calls software unit j" is used as the "edge" connecting software unit i and software unit j, thus forming a path branch. According to this method, connect all the call relationships where the matrix elements (M(t) [i,j] ) are not 0, and all the actual paths taken during the test are obtained, which constitute the call flow graph, where multiple consecutive software units and their call relationships form a dynamic call chain.. If it does not appear (or is 0) in the matrix, it means that the call relationship between the corresponding two software units has not been tested and executed.
[0103] Therefore, based on the call flow graph, the path exploration space for the ants in the present invention to explore paths is constituted. In addition, through the aforementioned dynamic call graph matrix, it is also possible to obtain the coverage of modules / functions and the coverage of paths or path branches during the test process.
[0104] Since there is a specific mapping relationship between software units and code, the code file or code line corresponding to the software unit can be obtained. Based on the matrix element M(t)[X,Y]≠0, by combining code-level coverage tools, it can be confirmed that "X calls a specific code branch or code line in Y".
[0105] In order to enable test cases to test the target software in various environments, when initializing the test environment, the present invention also obtains various software running environment simulation configuration parameters, which are used to simulate the running states of different environments and provide diverse test scenarios for testing. Among them, the running environment simulation configuration parameters include three types of parameters: the operating system version (abbreviated as O), network conditions (abbreviated as N), and hardware configuration data (abbreviated as H). Each type of running environment simulation configuration parameter may include one or more running environment simulation configuration parameters. In one embodiment, each running environment simulation configuration parameter is set as an environmental factor e i , since the software running environment (or the system environment of the target software) has different impacts on different test cases, in order to measure this impact, an environmental fitness function E(C, v) is defined. By measuring the fitness of a test case individual v to the environment under a specific system configuration, the impact of the software running environment on the test case individual v is quantified. Among them, C represents the set of software running environment simulation configuration parameters, C = {O, N, H}, O represents the operating system version, N represents network conditions, and H represents hardware configuration data. The present invention uses the set E = {e1, e2,..., e i} of environmental factors e n to represent a specific system simulation configuration. Each environmental factor e i can correspond to a running environment simulation configuration parameter or a set of multiple running environment simulation configuration parameters. Each environmental factor e i is assigned a corresponding weight w i . Based on the fitness function E(C, v), the comprehensive impact I (abbreviated as the environmental impact value) of the environmental configuration is calculated by the following formula (1-2):
[0106]
[0107] Among them, f(e i ) is the impact value of the environmental factor e i . Among them, the calculation method of the impact value f(e i ) of the environmental factor e i is a conventional technique in the field of software testing and will not be elaborated here.
[0108] In addition, security is usually also an important aspect in software testing. In one embodiment, when a software unit such as a software unit or a code segment is marked with a security risk attribute tag, the security risk attribute tag of the software unit is obtained during initialization, so that the path that needs to be subjected to security testing can be obtained. In order to evaluate the ability in terms of security testing during the test case generation process, the present invention further obtains a threat model for security testing during initialization. The threat model is a security module for analyzing and processing a certain or certain security vulnerabilities or security threats, such as the Common Vulnerability Scoring System (CVSS for short). The threat model can judge whether a vulnerability is detected based on the triggering condition corresponding to the security vulnerability, and quantitatively score or rate the risk of the vulnerability when the vulnerability is detected. For example, the threat model determines whether a test case triggers a vulnerability (such as SQL injection detection returns 1 / 0) through symbolic execution, and quantifies the severity s of the security vulnerability and the possibility a of the vulnerability being exploited through the risk scoring function R(s,a) shown in formula (1-3).
[0109] R(s,a) = s × a (1-3)
[0110] Among them, s represents the vulnerability severity, s ∈ [0, 10], a represents the possibility of the vulnerability being exploited (calculated by Monte Carlo simulation), a ∈ [0, 1], and the product of the two is used as the comprehensive quantitative risk value. For example, for a certain XSS vulnerability, R(s,a) = 0.7 × 8.5 + 0.3 × (probability 0.9) ≈ 7.2, and it is determined to be a high risk according to the corresponding rating standard.
[0111] The security vulnerability types in the present invention include SQL injection, cross-site scripting attack, buffer overflow, etc.
[0112] In step S102, each piece of data in the user historical behavior data includes at least user identity information, time information, user operation information, and corresponding input information. These parameters representing various types of information are collectively referred to as user parameters. According to the category, user parameters include user operation parameters, input parameters corresponding to user operations, user identity parameters, and environment parameters. The parameter values of user operation parameters can be various user operations, such as click, select, input, etc.; the parameter values of input parameters are, for example, the specific content corresponding to the user operation, such as the confirmation key information corresponding to the click confirmation key operation, the commodity information corresponding to the select operation, the input content corresponding to the input operation, etc.; the parameter values of user identity parameters are, for example, information such as user name and password; the environment parameters are, for example, one or more types of information such as operating system version, network condition, and hardware configuration.
[0113] For example, for an online shopping platform system, a piece of user historical behavior data collected is as follows (time information is omitted):
[0114] User operations and corresponding input information:
[0115] Search keyword - Smart phone;
[0116] Select product category - Electronics;
[0117] Select product price range - (1000 - 2000) yuan;
[0118] User account information:
[0119] User name - user123;
[0120] Password - pass456;
[0121] Environmental parameters:
[0122] User's device type - Android phone;
[0123] Operating system version - Android 12;
[0124] Network conditions - (Wi-Fi, bandwidth 10Mbps).
[0125] The parameter values of the user operation parameters extracted from the above data are "search keyword", "select product category", "select product price range", etc. The corresponding input parameters are, for example, specific keywords (such as smart phones), specific product categories (such as electronics), etc. The parameter values of the environmental parameters include the user's device type ("Android phone"), the user's device operating system version ("Android 12"), the network conditions used by the user ("Wi-Fi, bandwidth 10Mbps"), etc. The parameter values of the user identity parameters are specific account information, that is, the user name "user123" and the password "pass456".
[0126] The user behavior model in the present invention is various data that can simulate user behavior. For example, the distribution probability and usage frequency of specific user parameter values of a certain type of user, a sequence composed of multiple user parameter values representing different user behaviors; the association relationship between different categories and / or the same category of user parameter values representing user behavior patterns.
[0127] Taking the user operation parameters as an example, the parameter values of the user operation parameters are specific categories of user operations, and the distribution probability of each user operation is calculated based on formulas (1 - 4).
[0128]
[0129] Among them, x is a specific category of user operation, n xis the number of user operation x extracted from historical data in a preset time period (such as one week, three days, one month, etc.), and N is the number of all user operations of the user in the same time period. User operation x is, for example, "search keyword", "select product category", or "select product price range" in the foregoing embodiments, etc. By calculating the distribution probability of each specific user operation, it is possible to distinguish whether a user operation conforms to the user's habitual operation. Similarly, the distribution probabilities of various other types of user parameters can also be obtained, such as the distribution probabilities of various specific input parameter values, the distribution probabilities of various environmental parameter values, etc.
[0130] The usage frequency of user parameters can also reflect the corresponding user behavior patterns. For example, by counting the frequency of a user clicking a button in different time periods within their respective statistical time periods, the user's habit of clicking the button can be obtained, such as clicking 5 times per minute. Another example is that according to the average number of product pages browsed by different users during a search process, a user behavior pattern of accessing 10 pages each time can be obtained.
[0131] A sequence composed of multiple user parameter values can also represent the corresponding user behavior. For example, for a shopping platform, a user behavior model obtained is: log in - click on product 1 - click on product 2 - click on product 3 - click on product 4 - click on product 5 - click on the purchase button on the product page... Another user behavior model obtained is: log in - select product category, search for product keywords - click on product 1 - click on product 2 - click on add to cart - click on product 3 -... - click on the purchase button... The sequences composed of multiple categories of user parameter values representing the above two different user behaviors are each a user behavior model, and these two user behavior models represent different user behavior patterns.
[0132] Another example is that when the user behavior model is a sequence of specific user parameter values such as the operation of clicking on a product - specific product information - operating system in environmental parameters, the user behavior models composed of different specific parameter values represent different user behavior patterns.
[0133] The distribution probability of each of the above user parameter values, the usage frequency of each user parameter value, and the sequence of multiple user parameter values are collectively referred to as the user behavior model.
[0134] The input parameter space of the target software includes multiple input variables, and the input variables correspond to user parameters. The variable value of each input variable includes one or corresponding multiple user parameter values. In one embodiment, the input parameter space of the target software is represented by a set of variables: X = {x1, x2,..., x n}; X represents the input parameter space, and the variable x iRepresents the i-th input variable. The input variable can be an input operation or input parameter that the target software can accept. Here, the input operation corresponds to the user operation in the user behavior model, and the input parameter value is the specific information, data, etc. that the user can input, corresponding to the input parameter value corresponding to the user operation. For example, the input operation can be a login operation, inputting a search keyword, clicking on a relevant link, and the input parameter value can be the content that can be input for an input operation, such as a specific keyword, a specific link, etc. The input variable can also be a user identity parameter, such as a username and password. The input variable can also be an environmental parameter, such as a user device parameter, the user device operating system, network configuration or conditions, etc. The specific variable values of these input variables form the complete space of the input parameters of the target software. For example, the input variable x1 represents the username, and its value can be "user123" or "admin", etc.; the input variable x2 represents the user operating system, and the parameter value can be "Windows10", "Android12", etc., the input variable x3 represents the login operation, the input variable x4 represents the operation of inputting a search keyword, the input variable x5 represents the keyword, and its value is one or more keywords in the keyword table, etc.
[0135] The test case in the present invention refers to a specific instance for a complete test task, including multiple component parameters, and is the smallest unit in the test case set. In the following description, in order to highlight the relationship with the test case set, a specific instance for a complete test task is called a test case individual, and when not emphasizing the relationship with the test case set, a specific instance for a complete test task is called a test case.
[0136] To achieve a complete test task, the component parameters that make up the test case are of multiple types and need to meet certain conditions. For example, the component parameters of the test case at least include specific input operations and corresponding input parameter values. The input operation can be one or multiple and consecutive, that is, an input operation sequence. Referring to the paths and path branches in the call flow graph, an input operation and the corresponding input parameter value in the test case can implement a transition from one node to another node, that is, implement a path branch. When the test case includes multiple and consecutive input operations, multiple consecutive path branches can be implemented, that is, constitute a path.
[0137] In addition, the component parameters of the test case can also include environmental parameters and, when necessary, user identity parameters. The environmental parameter is used to define the running scenario, and the user identity parameter is a special input parameter.
[0138] In step S103, before driving multiple ants to perform path exploration in different regions of the path exploration space, ant colony parameters are first set, such as the initial pheromone values of each path branch, the pheromone evaporation rate ρ, the calculation method of the pheromone intensity constant Q, the number of ants, the number N of initial test case individuals to be obtained, the exploration termination condition, etc. Among them, according to the number of ants and the number N of test case individuals to be obtained, the exploration times of each ant are configured.
[0139] In one embodiment, the initial pheromone value is set for the paths in the path exploration space based on the user behavior model. For example, first, the corresponding user operation is determined based on the matching relationship between the paths in the path exploration space and the user operations. For example, in a page browsing path, each path branch formed by each node (such as a page or a state) among multiple nodes corresponds to a user operation. Then, the user behavior model is queried based on the user operation, and the initial pheromone value of the path branch is set according to the distribution probability of the user behavior operations in the user behavior model. The larger the distribution probability, the larger the initial pheromone value; the smaller the distribution probability, the smaller the initial pheromone value. Thus, the ants can be made to select paths that are more matched with the user behavior during path exploration. For some high-risk regions, in order to obtain test cases for effectively performing security risk testing, larger initial pheromone values are set for the paths in these regions, so as to guide the ants to explore these regions when the ants perform path exploration.
[0140] The process of each ant's path exploration specifically includes: First, a starting node is determined for each ant, and then the probability from the starting node to all possible next nodes is calculated according to the probability formula. Among them, the probability formula is shown as formula (1-5) below:
[0141]
[0142] τij: The pheromone concentration from node i to node j, which is calculated according to formula (1-6). is the probability from node i to node j, and l is one of the multiple candidate nodes, that is, one of all the allowed nodes.
[0143]
[0144] Among them, t is the time parameter in the ant colony algorithm, corresponding to the present invention, which is the iteration number when the ant performs path exploration; in this embodiment, when generating the initial test case set, it terminates when N test case individuals that meet the conditions and are preset are obtained. Before termination, the ants need to perform path exploration multiple times, that is, multiple iterations. τ ij (t + 1) is the pheromone value from node i to node j (referred to as path L i,j ) for the (t + 1)-th exploration updated after the t-th exploration, τ ij(t) represents path L i,j The pheromone concentration at the t-th exploration, where v is the evaporation rate of pheromone, is the pheromone value released (left) by the k-th ant when exploring path L i,j and is also the pheromone increment left after being explored by one ant on path L i,j where Q is the pheromone intensity constant, and L k is the path length of the k-th ant.
[0145] η ij is the heuristic factor from node i to node j. The heuristic factors in the present invention include user behavior weight and security risk weight. For example, η ij = η1·user operation probability + η2·security risk value, where the user operation probability is a specific embodiment of the user behavior simulation matching degree. Additionally, the security risk weight can be increased. Of course, other weights can also be set according to the needs of testing, such as test execution time weight, running environment weight, path length weight, business importance weight, high-frequency feature weight, etc. By introducing other concerned factors when selecting the next node, it is possible to direct to the concerned path branches during the path exploration process. Specifically, it can be flexibly used according to the nature of the target software in actual applications.
[0146] α and β are the weights controlling pheromone and heuristic factors respectively. Similarly, τ il is the pheromone from node i to node l, and η il is the heuristic factor from node i to node l.
[0147] After calculating the selection probabilities of all candidate nodes, the node with the maximum selection probability is selected as the next node, thus obtaining a path branch. After each ant determines a path branch during the path exploration process, based on the path branch conditions, the corresponding input variables are selected from the input parameter space, and then the corresponding input operations and input parameter values are selected from the input parameter space according to the user behavior model as the corresponding input variable values. And so on until the termination condition of the path exploration is reached, such as reaching the preset path length (e.g., at most 10-step operations), or entering the termination state (e.g., triggering a crash, completing the core process), or covering the target branch or meeting the coverage threshold, etc.
[0148] The complete search trajectory of each ant is the explored path, where the corresponding input variable values selected from the input parameter space corresponding to all path branches constitute a test case individual.
[0149] The user behavior models provided by the present invention have various forms, such as the distribution probability, usage frequency, specific operation sequences, etc. of the aforementioned user operations or input operations. Therefore, after determining a path branch during the path exploration process of the ant, it first determines the form of the user behavior model to be used based on all input variables corresponding to the path branch conditions. When only one input variable is involved, it selects the input operation or input parameter value with the highest distribution probability that can be used as the input variable value. For example, when the path is from the login page to the user's personal homepage, since a login variable is required, the corresponding input operation is the login operation, and the corresponding input parameters include two input parameters, namely the username and password. There are multiple usernames and multiple passwords in the input space. When there are multiple input parameter values, it selects the username and password with the highest distribution probability as the input parameter values, thereby obtaining a test case individual: click the login button; input parameters: username "user123" and password "pass456".
[0150] When there are multiple input operations corresponding to the input variable, it selects the input operation with the highest distribution probability or the input operation with the highest usage frequency in the user behavior model. Figure 2 It is a schematic diagram of the path structure according to an embodiment of the present invention. For Figure 2 the node i in, the corresponding input operations can be "click product operation a" and "search operation b". Referring to the user behavior model, the distribution probability of "search operation b" is the highest, so "search operation b" is selected. Another example is when the determined path is from node i through node j1, node j3 and then to node j, which includes two consecutive input operations "search operation" and "click operation". According to the user behavior model: "search for'smartphones', select the 'electronic products' category, and browse products with a price range of '1000 - 2000 yuan'", the specific input operation values are determined, thereby obtaining a test case individual: "input parameters corresponding to the input operation: search for'smartphones'; select the 'electronic products' category, price range '1000 - 2000 yuan'; click xxx mobile phone".
[0151] Since the user behavior model also provides the usage frequency of input parameters, further, the input parameters can be selected according to the usage frequency of input parameters. For example, according to the user behavior model, when obtaining the usage frequency of the input parameter that the user views 10 pages on average each time when browsing the product page, add the browsing operation steps of 10 product pages in the generated test case. For example, the generated test case simulates that the user uses a "wired network" on a computer with the "Windows 10" system to access the shopping platform and randomly clicks on 10 product pages.
[0152] Another example Figure 2For the path shown, based on the user behavior model, when the frequency of a user clicking on products during product browsing is 5 per time, 4 randomly selected products can be added on this basis, thus obtaining a test case individual: "Search for'smartphone'; select the 'electronics' category, with the price range being '1000 - 2000 yuan'; click on xxx phone 1; click on xxx phone 2; click on xxx phone 3; click on xxx phone 4; click on xxx phone 5".
[0153] Finally, in order to specify the running scenario of the currently generated test case individual, the running environment parameters can usually be configured as well, thereby obtaining a test case individual with complete component parameters, such as "Username: user123; Password: pass456; Search for'smartphone'; Select the 'electronics' category; Select the price range of '1000 - 2000 yuan'; Click on xxx phone 1; Click on xxx phone 2; Click on xxx phone 3; Click on xxx phone 4; Click on xxx phone 5; Operating system version: Windows10; Network configuration: Wi-Fi, latency 50ms". Of course, other more detailed running environment parameters such as hardware conditions can also be included.
[0154] During the ant's path exploration process, the pheromone values of the path branches are updated in real time. Specifically, the following several update methods are included:
[0155] (1) Adjusting pheromone update according to defect discovery priority: In actual application scenarios, some functional modules or code segments are more likely to have defects. Therefore, at this time, ants can be given priority to explore "high-risk areas". For example, before testing, the risk levels of functional modules or code segments are marked based on data analysis. During the ant's path exploration process, it is judged whether the currently explored path branch is marked with a security risk attribute mark. When the currently explored path branch is marked with a security risk attribute mark, the pheromone increment of the path branch marked with the security risk attribute mark is increased. And, the pheromone increment is adjusted according to the risk level. The higher the risk level, the greater the pheromone increment value. Furthermore, during path exploration, the probability of the path corresponding to this functional module or code segment being selected is increased, thereby generating test cases that can reveal defects in potential high-risk areas.
[0156] (2) Path optimization scenarios; Since different software systems have different test path requirements, according to the path optimization requirements of the target software and specific test scenarios, pheromone update rules or update formulas are defined for each specific scenario. The scenarios mentioned above are, for example: user login scenario, data input verification scenario, API interface call scenario. During path exploration, it is judged which scenario the currently explored path exploration space area corresponds to, and the pheromone values of the path branches are adjusted based on the matching method with the scenario.
[0157] (3) Scenarios of path coverage: During the process of driving multiple ants to explore paths in different regions of the path exploration space, calculate the coverage rate of the explored paths for the path exploration space or a preset local path exploration space. Based on the corresponding relationship where the coverage rate is negatively correlated with pheromone, adjust the pheromone values of the path branches in the path exploration space, so that the higher the path coverage rate, the lower the pheromone concentration, thereby guiding the ants to select less-covered paths as test case individuals.
[0158] (4) Performance testing scenarios: When conducting performance / stress testing, one often cares about the time that each test path may consume, the resource occupancy during concurrent execution (CPU, memory, network bandwidth, etc.), the upper limit of test execution time or throughput metrics, etc. Thus, in one embodiment, add the calculation weight of "execution time (or resource consumption) T" to the foregoing formula (1-7). For example, replace parameter L k with L1·path length + L2·execution time + L3·resource overhead, or take the execution time / resource as a penalty term Rk and introduce it into the calculation formula, that is so that when the ants explore paths, they can tend to select paths that are short, fast, and resource-saving, achieving the minimization of resource consumption.
[0159] (5) Adjustment based on environmental fitness: In actual test scenarios, the test performance of different system configurations is crucial. Thus, in one embodiment, add environmental factors to the calculation process of pheromone concentration to generate test cases that are more in line with the real usage scenarios. For example, set the pheromone increment according to the magnitude of the influence value calculated based on the operating environment parameters of the current test environment.
[0160] (6) Optimization adjustment based on user behavior: According to the frequency and habits of users' actual operations (such as login, query, submit form), increase the pheromone concentration for common paths, and preferentially generate test cases for these common paths to improve the actual value of the test cases. For example, after initializing the test environment to obtain the path exploration space of the target software and the user behavior model, query the paths in the path exploration space of the target software through the user behavior model, determine the paths corresponding to the operations with high usage frequency or high distribution probability in the user behavior model, and set larger pheromone increment values for these paths, so that the pheromone values of these paths can be increased during information update, thereby achieving the purpose of preferentially generating test case individuals for these paths, and thus improving the actual value of the test case individuals.
[0161] To achieve the above adjustment of pheromone values, the calculation weights of the foregoing various factors can be added to the pheromone intensity constant Q or the path length L k so as to adjust the pheromone values left by the ants in the foregoing various situations It serves the purpose of guiding other ants to explore the predetermined area.
[0162] Furthermore, in order to make the initialized test cases closer to real user behaviors and real operation scenarios, and be better oriented towards the final test goal, making subsequent iterative optimization easier and convergence faster, in another embodiment of the present invention, when constructing the initial test case set, a batch of initially generated test case individuals are used as candidate initial test case individuals. A maximum value problem function is constructed based on the user behavior model and environmental fitness. By adjusting the candidate initial test case individuals, or adjusting the user behavior model and / or operating environment parameters of the candidate initial test case individuals, the maximum value problem function is converged to the maximum value; the candidate initial test case individuals when converging to the maximum value are determined as the initial test case individuals.
[0163] In one embodiment, the maximum value problem function is as follows:
[0164]
[0165] where P(x i ) is the distribution probability of a specific input variable value x i in the test case, that is, the probability of being used, and is calculated according to formula (1 - 4). n is the total number of various input variable values in all test cases. E(C, v) is the environmental impact value I of the operating environment parameters in each test case on this test case, and is calculated according to formula (1 - 1). m is the total number of operating environment parameters in the current test case set V. α and β are trade - off coefficients, used to balance the influence of the user behavior model and environmental configuration on the initialization process.
[0166] The present invention measures the degree to which the test case set as a whole is close to real user behaviors and real operation scenarios based on the user behavior model and the simulated configuration parameters of the operating environment, and optimizes the test case individual v according to the measurement result.
[0167] See Figure 3 , Figure 3 is a flowchart of a method for optimizing initial test case individuals to obtain an initial test case set according to an embodiment of the present invention. In this embodiment, the test case individuals obtained during the ant path exploration process are used as candidate initial test case individuals, and the candidate initial test case individuals constitute the candidate initial test case set V. The method includes the following steps:
[0168] Step S1031, obtain a candidate initial test case individual from the candidate initial test case set V as the target test case individual.
[0169] Step S1032: Calculate the user behavior simulation score of the target test case individual based on the user behavior model used when generating the target test case individual. In this embodiment, that is, calculate the distribution probability P(x i ) of each input variable value in the target test case individual. Specifically, for each input variable value, use formula (1-4) to calculate its probability.
[0170] Step S1033: Calculate the environmental fitness score of the target test case individual based on the operating environment parameters of the target test case individual. In this embodiment, that is, calculate the influence value of the operating environment simulation configuration in the target test case individual on this target test case. Specifically, take each operating environment simulation configuration parameter in the target test case as an environmental factor e, and use formula (1-1) to calculate the environmental influence value I of the current operating environment parameter on the target test case individual.
[0171] Step S1034: Determine whether there is still a candidate initial test case individual in the candidate initial test case set V. If there is, return to Step S1031. If there is no candidate initial test case individual in the candidate initial test case set V, then execute Step S1035.
[0172] Step S1035: Calculate the weighted sum of the user behavior simulation scores and environmental fitness scores of all candidate initial test case individuals. In this embodiment, calculate the weighted sum of the distribution probabilities of all input variable values and the environmental influence value I of the current all candidate initial test cases according to formula (1-8).
[0173] Step S1036: Adjust the operating environment parameters of the candidate initial test case individual and / or the user behavior model used to generate the candidate initial test case individual to obtain a new test case individual. For example, randomly change the input variable values of one or more candidate initial test case individuals in the candidate initial test case set V in the direction of increasing the probability and / or randomly change the operating environment parameters of one or more candidate initial test case individuals in the test case set V in the way of increasing the environmental influence value I, and use it as the new target test case individual. For example, apply the genetic algorithm, take each test case individual as a chromosome in the genetic algorithm, and each input variable value and each operating environment configuration parameter in the test case individual as a gene in the genetic algorithm, and change the input variable values and / or operating environment parameters in a test case individual through gene crossover and mutation operations to obtain a new test case individual.
[0174] Step S1037: Calculate the user behavior simulation score and environmental fitness score in the new target test case. For example, calculate the probability p(x) that each input variable value in the new target test case is used and the environmental influence value I.
[0175] Step S1038, calculate the weighted sum of the user behavior simulation scores and the environmental fitness scores of all test case individuals in the current test case set.
[0176] Step S1039, determine whether the weighted sum of the current test case set V has converged to the maximum value. If it has converged to the maximum value, end the process; otherwise, return to Step S1036. Here, "converging to the maximum value" means that the weighted sum of the current test case set V no longer increases. That is, by comparing the current weighted sum with the weighted sum calculated in the previous time, when the weighted sum no longer increases, it is confirmed that it has converged to the maximum value.
[0177] Furthermore, after obtaining the initial test case set, evaluate the test effectiveness of each test case individual. In one embodiment, calculate the fitness score of each test case individual through the following fitness function expression (1-9). In this embodiment, evaluate the test effectiveness through the software metric coverage ability and the ability to discover potential defects.
[0178] f(v) = γ·Coverage (v) + δ·Defects (v) (1-9)
[0179] Coverage(v) is the software metric coverage rate of a single test case individual v, usually a percentage, ranging from (0-1). In one embodiment, the software metric is, for example, a path. When specifically calculating, determine the number of path branches included by the path corresponding to the test case individual v, calculate the proportion of the number of path branches of the test case individual v to the total number, and use the path branch proportion as the coverage rate Coverage(v). The software metric can also be, for example, the number of code lines. Based on the test path corresponding to the test case individual v, determine the relevant software structure, and based on the correspondence between the software structure and the code, determine the number of code lines. Calculate the proportion of the number of code lines corresponding to the test case individual v to the total number of code lines, and use the code line proportion as the coverage rate Coverage(v). Or, the software metric can also be, for example, the function modules of the software. Calculate the ratio of the number of function modules covered by the test case individual v to the total number of function modules, and use the function module ratio as the coverage rate Coverage(v). For example, if the target software system has 100 function modules and the test case individual v covers 20 of them, then Coverage(v) = 20%.
[0180] Defects(v) represents the number of potential defects that can be revealed by an individual test case v, indicating the defect detection ability of an individual test case v. The defect detection ability Defects(v) is a positive integer, and its specific range depends on the test scenario and system complexity. For example, its calculation method is as follows: using the aforementioned security threat model to analyze the potential defect triggering ability of the test case individual v on the target software through symbolic execution, such as predicting the types and quantities of vulnerabilities that can be exposed.
[0181] γ and δ are weight coefficients used to balance the contributions of coverage and defect detection ability to fitness.
[0182] In this embodiment, the weight coefficients can be set as needed. If coverage is given priority, γ > δ can be set. If defect discovery is emphasized, δ > γ is set.
[0183] For example, in a medium - sized software system, for a test case, its Coverage(v) may be between 0.1 and 0.5 (i.e., 10% - 50%), and its Defects(v) may be between 0 and 5 (the number of detected defects is limited).
[0184] For example, set γ = 0.6 and δ = 0.4. If the coverage of the test case individual v is 0.3 (30%) and 2 defects are detected, then the fitness score is: f(v) = 0.6·0.3 + 0.4·2 = 0.18 + 0.8 = 0.98. This score indicates that the test effectiveness performance of this individual is very good and is suitable for being retained in the current set for further optimization.
[0185] In a large and complex software system, for a test case individual v, Coverage(v) may be between 0.01 and 0.3. Defects(v) may be between 0 and 10 (it is easier to discover defects).
[0186] For example, set γ = 0.7 and δ = 0.3. If the coverage of the test case individual v is 0.1 and 5 defects are detected, then the fitness score is: f(v) = 0.7·0.1 + 0.3·5 = 0.07 + 1.5 = 1.57. This score indicates that the test case individual v is relatively effective.
[0187] The scores calculated for each test case individual through the expression (1 - 9) are used as the initial fitness scores of the test case individuals. The higher the score calculated based on the fitness function expression (1 - 9), the better the performance of the test case individual v in terms of both software function coverage and the ability to discover potential defects, that is, the better the test effectiveness, the higher the potential value, and the more effectively it can discover software defects or verify software functions.
[0188] In one embodiment, after the weighted sum of the current test case set V converges to the maximum value, candidate initial test case individuals with initial fitness scores meeting the requirements are determined as the initial test case individuals. For example, the test case individuals in the current test case set V are sorted according to the initial fitness scores, and a certain number of test case individuals ranked at the front are used as the initial test case individuals, so as to obtain the final initial test case set. Or, the test case individuals with initial fitness scores greater than the threshold are used as the initial test case individuals, so as to obtain the final initial test case set.
[0189] In addition, with reference to the evaluation of the test effectiveness of each test case individual, in another embodiment, the test effectiveness of the test case set V can also be evaluated as a whole. Specifically, refer to the fitness function expression (1-10):
[0190] f(V) = γ·Coverage(V) + δ·Defects(V) (1-10)
[0191] Where Coverage(V) represents the software metric coverage rate of the test case set V, reflecting the coverage of the test case set V for the target software, Defects(V) represents the number of potential defects that can be revealed through the test case set V, reflecting the effectiveness of the defect detection ability of the test case set, and γ and δ are trade-off coefficients used to balance the contributions of the coverage rate and the defect detection ability to the fitness.
[0192] The initial test case set of the present invention has a certain degree of diversity, can initially cover key or common use case scenarios, and to a certain extent faces or takes into account goals such as function coverage, security testing, and user behavior coverage. Therefore, in the subsequent iterative optimization (step S104), combined with fitness evaluation (such as coverage rate, security vulnerability discovery ability, user behavior matching degree, etc.), these initial test case individuals are continuously optimized to gradually approach or continuously improve the satisfaction degree of the test target. It avoids suddenly expanding the search space during optimization, thus reducing the search difficulty and also improving the optimization efficiency.
[0193] In step S104, when screening the parent test case individuals from the current test case set based on the pheromone value and the fitness of the test case individuals, first obtain the fitness score of each test case individual in the current test case set; then obtain the total pheromone value of the path branches covered by each test case individual in the current test case set to obtain the pheromone value of each test case individual; then calculate the selection probability of each test case individual in the current test case set based on the pheromone value and the fitness score. In one embodiment, the selection probability of each test case individual is calculated based on the following formula (1-11):
[0194]
[0195] Among them, P select (x) is the selection probability of a test case individual calculated based on pheromone and fitness. τ(x) represents the pheromone concentration associated with the test case individual x, and f(x) is the fitness score of the test case individual.
[0196] Then, the test case individuals with selection probabilities greater than the threshold are used as the parent test case individuals selected from the current test case set; or the test case individuals are sorted in descending order of selection probability, and the preset number of test case individuals ranked at the front are used as the parent test case individuals selected from the current test case set.
[0197] Among them, when the current test case set is the initial test case set, the fitness score is the initial fitness score calculated using formula (1-9).
[0198] If the current test case set is the optimized test case set, first, collect the corresponding running data when running each test case individual in the current test case set; then calculate the coverage rate of the test case individual for the target software metrics based on the running data of each test case individual; query the user behavior model data and calculate the matching rate between the simulated user behavior of each test case individual and the user behavior model; evaluate each test case individual based on the threat model for security testing to obtain the defect revelation ability value; finally, calculate the weighted sum of the coverage rate of each test case individual for the target software metrics, the defect revelation ability value, and the matching rate between the simulated user behavior of the test case individual and the user behavior model as the fitness score of each test case individual.
[0199] In one embodiment, the fitness score of each test case individual is calculated based on the following formula (1-12):
[0200]
[0201] Among them, covered(i,v), detect(d,v), and match(u,v) are indicator functions respectively. The three terms in formula (1-12) respectively target the corresponding coverage criteria, security criteria, and user behavior matching criteria during evaluation, and are respectively used to judge the software metrics (such as lines of code, functions, modules, path coverage, conditions, etc.) covered by the test case individual v, the ability to detect security threats or vulnerabilities, and the matching degree with user behavior. λ1, λ2, and λ3 are the weights for the three criteria respectively.
[0202] Specifically, c is a coverage criterion, C is a set of coverage criteria, and the set of coverage criteria C includes various coverage criteria for lines of code, functions, modules, as well as path coverage and condition coverage. i is a coverage unit, such as a line of code or a code item, function, branch, module, path, or condition, etc., and I c is the set of coverage units for the coverage criterion c, and w c is the weight for the coverage criterion c, |I c | is the total number of coverage units i, and covered(i, v) is an indicator function that is 1 if the test case individual v covers the coverage unit i, and 0 otherwise.
[0203] s is a security test type, S is a set of security test types, d is a vulnerability or security threat (hereinafter referred to as a security test item), and D s is the set of security test items for the security test type s, and w s is the weight for the security test type s, and detect(d, v) is an indicator function that is 1 if the test case individual v can detect the security test item v, and 0 otherwise, and T s is the detection threshold for the security test type s, and e is the base of the natural logarithm.
[0204] b is a user behavior pattern, B is a set of user behavior patterns, and U b is the set of user operation sequences for the user behavior pattern b, and w b is the weight for the user behavior pattern b, p(u) is the probability of the user operation sequence u occurring, and match(u, v) is a matching function that reflects the matching degree between the test case individual v and the user operation sequence u.
[0205] In step S105, when performing a reproduction operation based on the parent test case individuals, the reproduction operation can be performed based on each parent test case individual, or some parent test case individuals can be selected for the reproduction operation. When selecting some parent test case individuals, some test case individuals with high fitness can be selected as the parent test case individuals, or the parent test case individuals can be selected from the set of parent test cases for the reproduction operation based on the roulette wheel selection mechanism.
[0206] See Figure 4 , Figure 4 is a flowchart of a method for selecting parent test case individuals for reproduction using the roulette wheel selection mechanism according to an embodiment of the present invention. It includes the following steps:
[0207] Step S1051, obtain the fitness score of each parent test case individual.
[0208] Step S1052, calculate the selection probability p of each parent test case individual in the parent test case set select (v).
[0209] Step S1053, calculate the cumulative probability P of each parent test case individual in the parent test case set select (v).
[0210] Step S1054, generate a random number, and select the corresponding test case individual as the parent test case individual based on the position where the random number is located.
[0211] Step S1055, determine whether the number of selected parent test case individuals has reached the preset quantity. If it has reached, end. If it has not reached the preset quantity, return to Step S1054.
[0212] Among them, in Step S1051, the fitness score of the parent test case individual can be the initial fitness score calculated based on the foregoing formula (1-9) or the fitness score calculated based on the formula (1-12).
[0213] In Step S1052, in one embodiment, the selection probability p of a target parent test case individual in the parent test case set can be calculated by the following formula (1-13) select (v):
[0214]
[0215] Among them, p select (v) represents the probability that the target parent test case individual v is selected for reproduction; p select (v) ∈ [0, 1], Population represents the parent test case set, and the sum of the selection probabilities of all parent test case individuals v is 1; F(v) is the fitness score of the target parent test case individual v. In the first optimization, it is the initial fitness score. In the subsequent optimization process, it is calculated by applying the formula (1-12), and it reflects the quality of the test case individual v from three aspects: coverage rate, security test ability, and user behavior matching degree. The higher F(v) is, the better the quality of the test case is, and the greater the probability of being selected;.
[0216] ∑ v′∈Population F(v) represents the sum of the fitness scores of all parent test case individuals in the parent test case set.
[0217] In another embodiment, in order to enhance the influence of the fitness difference of the test case individuals and make the influence of the fitness difference of the test case individual v on the selection probability more significant. Calculate the selection probability p of the parent test case individual based on the formula (1-14)select (v):
[0218]
[0219] T is a regulatory parameter used to adjust the sensitivity of the selection mechanism to fitness differences. When the value of T is large, the probability distribution is smoother, and individuals with low fitness have a certain probability of being selected, which helps to maintain the species characteristics. When T is small, the selection is more concentrated on individuals with high fitness, accelerating convergence but possibly leading to premature convergence to a local optimum. By adjusting the parameter T, the smoothness of the selection probability distribution is controlled, allowing individuals with lower fitness scores to also have a certain probability of being selected. Among them, in the selection process of the same round of optimization, the same regulatory parameter T is used for all test cases for calculation, so as to ensure that a consistent selection strategy is implemented for the entire population, facilitating the comparison of the fitness of different individuals and the assignment of selection probabilities. For different iteration rounds, the value of T can be dynamically adjusted to enhance or weaken the "amplification" or "smoothing" effect on fitness differences.
[0220] In this embodiment, through the operation of e F(V) / T , the fitness scores of the test case individuals v are exponentiated, enhancing the influence of the fitness differences of the test case individuals and making the influence of the fitness differences of the test case individuals v on the selection probability more significant.
[0221] In step S1053, in order to calculate the cumulative probability P select (v) of the test case individuals, all the test case individuals are queued and numbered n, where n = 1... N, and N is the total number of test case individuals in the seed set. For any test case individual v i in the queue, the selection probability p i (v select ) of the test case individual v and the sum of the selection probabilities p i of all the test case individuals sorted in front of it are used as its cumulative probability P select (v select ), that is i Thus, N cumulative probability values are obtained, forming a cumulative probability value queue. Then, in step S1054, the range of the random number is 1 to N, and the random number is used as the sorting of the cumulative probability value queue, so as to obtain the test case individual corresponding to the position indicated by the random number.
[0222] After obtaining the parent test case individuals, the breeding operations performed on the parent test case individuals are, for example, crossover operations and / or mutation operations. Each test case individual is regarded as a chromosome, and each input variable value and each running environment configuration parameter in the test case are regarded as chromosome genes. By means of gene crossover and mutation operations, the input variable values and / or running environment parameters in a test case individual are changed to obtain new test case individuals.
[0223] For example, the crossover operation is performed through the following formula (1-15), and the mutation operation is performed through the following formula (1-16):
[0224] v(t + 1)′ i = crossover(v(t) i , v(t) r1 , cr) (1-15)
[0225] v(t + 1)″ i = mutate(v(t) i , mr) (1-16)
[0226] Among them, crossover and mutate respectively represent the crossover and mutation operations, cr and mr are respectively the crossover rate and the mutation rate. In this embodiment, they can refer to the number, ratio, position, etc. of genes that can represent crossover or mutation, so as to determine the specific input variable values (such as user identity parameters, input operations, input parameters, etc.) and / or environment parameters for crossover or mutation. v(t + 1)′ i represents the test case individual after crossover, v(t) i is the selected test case individual for crossover, v(t) r1 is another randomly selected test case individual for the crossover operation. v(t + 1)″ i represents the test case individual after mutation.
[0227] For input parameters with a parameter value range, the mutation operation can be performed using formula (1-17):
[0228] R mutated = R + μ·(rand([-1,1])·(max(R)-min(R))) (1-17)
[0229] Among them, R mutatedIt represents the value of the input parameter after mutation; R represents the original input parameter value; rand([-1,1]) represents generating a random number within the range of [-1,1]. By introducing randomness, the invention makes the mutation direction and amplitude of each test parameter different; max(R) and min(R) respectively represent the upper and lower limits of the input parameter values, which are used to control the range of the mutation operation to ensure that the input parameter values do not exceed the acceptable bounds. By changing the mutation rate μ, the purpose of diversity control can be achieved. For example: setting a higher μ: making the input parameter span larger. For example, for the username and password in the input parameters, completely different username formats can be tried, while for some numerical parameters, a lower mutation rate μ can be set: fine-tuning near the original parameters. For example, changing the time interval of the login request, fine-tuning the price range of the goods clicked by the user, etc.
[0230] Through the cross and mutation of input operations, input parameters, user identity parameters, operating environment parameters, etc., the invention can obtain new test case individuals that simulate different user behaviors, meet the triggering conditions of security tests, increase the running scenarios, increase the coverage of software functions, code, functions, branches, conditions, paths, etc.
[0231] For example, an original test case individual is as follows:
[0232] Input parameters: {"user123","pass456"};
[0233] Input operation: Click the navigation bar 3 times after logging in;
[0234] Operating environment parameters: The operating system is "Windows10" and the bandwidth is 10Mbps.
[0235] When mutating, the input parameters can be mutated. For example, mutate the username "user123" in the original input parameters to user'OR'1'='1 in the input parameter space to trigger SQL injection testing.
[0236] When mutating the input operation, the user behavior pattern can be adjusted. For example, mutate the original click on the navigation bar 3 times to click on the navigation bar 5 times and add a search operation.
[0237] When mutating the operating environment parameters to adjust the environment configuration. For example, the original configuration is a bandwidth of 10Mbps; the mutated bandwidth becomes 5Mbps, so as to simulate a network fluctuation scenario.
[0238] The finally generated mutated test case is as follows:
[0239] Input data: {"user'OR'1'='1","pass456"};
[0240] Input operation: Click the navigation bar 5 times and perform a search;
[0241] Operating environment parameters: Operating system "Windows 10", bandwidth 5 Mbps.
[0242] Therefore, it can be seen that during the mutation operation, by mutating the input parameters (such as username and password), multiple possible vulnerability trigger conditions can be generated, generating test case individuals that are more likely to trigger security vulnerabilities, thereby expanding the security test coverage. Additionally, test case individuals simulating other user behaviors can also be generated through the mutation operation, making the test process closer to the real user operation habits and improving the user behavior coverage. For example, by changing the input operations in the test cases, new behavior patterns can be generated. For instance, when mutating the number of clicks and paths, new operation habits can be obtained, increasing the diversity of test cases. Through the mutation operation on the operating environment, the running states of different environments can be simulated, providing diverse test scenarios for the execution of the test.
[0243] After step S105, a new test case set Population is obtained combined , which includes the original test case individuals that have not undergone reproduction and the new test case individuals after reproduction. In a further embodiment, the new test case set Population combined is updated. The specific update steps include: calculating the fitness scores of the new test case individuals to measure their performance in terms of coverage, security testing ability, and user behavior simulation, and then traversing the fitness scores of each new test case individual, retaining the new test case individuals whose fitness scores are greater than or equal to the threshold, and eliminating the new test case individuals whose fitness scores are less than the threshold; or, sorting the new test case individuals in descending order of fitness scores; retaining the preset number of new test case individuals ranked at the front. Or, constructing a fitness score maximum problem; obtaining the maximum fitness score according to formula (1-18) based on the new test case set, and retaining the test case individuals at the time of the maximum fitness score as the updated test case set.
[0244] where Population combined is a combination of the original test case set and the newly generated test case individuals after reproduction, and N is the size of the test case set. Population new is the updated test case set.
[0245] Up to step S105, an optimization is completed. Then, in step S106, the global pheromone value is updated according to formula (1-19).
[0246]
[0247] Among them, Δτ is the pheromone value of the path branches covered by the parent test case individual and the reproduced test case individual. By introducing the crossover and mutation operations of the genetic algorithm into the update of pheromone, it is beneficial to generate test case individuals in the direction of covering the paths in high-risk areas during the next optimization. By adding the incremental pheromone of excellent solutions while decaying the old pheromone, both the original search experience is retained and the reinforcement of key paths is highlighted, thereby guiding the generation of subsequent solutions to test case individuals that expand to high-risk or key branches, avoiding over-concentration on local optima, and ultimately improving the coverage and quality of the test case set.
[0248] The optimization end condition in step S107 is, for example, that the current test case set meets the system test objectives. The test objectives include, for example, software system test objectives, security test objectives, and user behavior test objectives, and are calculated through the fitness calculation formula of the corresponding test case set. For software system test objectives, for example, specifically determine whether the current test case set covers the specified software internal structure and reaches the specified coverage rate, whether it reaches the code coverage rate, function call coverage rate, and condition judgment coverage rate of modules or components, and whether it reaches the specified environmental fitness (E(C, V)); for security test objectives, for example, whether the specified vulnerabilities or threats are detected, or whether the specified number or proportion of vulnerabilities or threats are detected. For user behavior test objectives, for example, whether the specified user behavior coverage and matching degree are reached. In one embodiment, the fitness score of the test case set is calculated according to the following formula (1-20):
[0249]
[0250] Among them, C, S, and B respectively represent the sets of coverage criteria for software systems, security test types, and user behavior patterns, w c , w s and w b are the weight coefficients corresponding to each type; c is a coverage criterion, C is the set of coverage criteria, and the set of coverage criteria C includes various coverage criteria for lines of code, functions, modules, path coverage, and condition coverage. Coverage c (V) is the score of the software units (such as various coverage criteria such as lines of code, functions, modules, path coverage, and conditions) covered by the test case set V calculated according to formula (1-21).
[0251]
[0252] Among them, I c is the set of covered units for the coverage criterion c, |Ic | is the total number of covered units i, and covered(i, V) is an indicator function that is 1 if the test case set covers the covered unit i and 0 otherwise.
[0253] s is the type of security test, and S is the set of security test types; Safety s (V) is the score calculated based on the following formula (1-22).
[0254]
[0255] where d is a vulnerability or security threat (hereinafter referred to as a security test item, such as SQL injection, cross-site scripting attack, and buffer overflow, etc.), and D s is the set of security test items for the security test type s, and detect(d, v) is an indicator function that is 1 if the test case individual v can detect the security test item v and 0 otherwise, and T s is the detection threshold for the security test type s, and e is the base of the natural logarithm.
[0256] b is the user behavior pattern, and B is the set of user behavior patterns; S45, User Behavior Coverage Behavior b (V) is the score calculated based on formula (1-23), which is used to evaluate the degree of agreement between the test case set and the user behavior model, and reflects the effect of the test case set in simulating real user behavior.
[0257]
[0258] where, U b is the set of user operation sequences for the user behavior pattern b, p(u) is the probability of the user operation sequence u occurring, and match(u, V) is a matching function that reflects the matching degree between the test case set V and the user operation sequence u. The user behavior coverage Behavior b (V) reflects the user behavior coverage and matching degree, that is, the user behavior covered by the user behavior in the test case and the corresponding matching degree.
[0259] In step S107, calculate the fitness score of the updated test case set. When the fitness score of the updated test case set reaches the threshold, it is considered that the optimization end condition is satisfied. At this time, stop the optimization and store the current test case set V final , V final = {V|V ∈ Population at termination}, and the finally obtained test case set V finalIt includes the test case individuals with the highest fitness after multiple rounds of iterative optimization. In addition, the optimization end condition further includes the number of iterations. When the number of iterations is reached, the optimization also stops. When the optimization is terminated only by the number of iterations, the finally generated test case set V can also be final is evaluated, and the fitness score of the test case set V is calculated according to the aforementioned fitness calculation formula, so as to evaluate whether the predetermined test coverage criteria are met, including code coverage rate, the ability to discover potential security vulnerabilities, and simulating user behavior. final
[0260] As can be seen from the foregoing method, the present invention combines the ant colony algorithm and the genetic algorithm, utilizes the advantages of the ant colony algorithm in specific path selection and the ability of the genetic algorithm in global search and optimization, guides the selection process of the genetic algorithm through the pheromone of the ant colony, and at the same time introduces the crossover and mutation operations of the genetic algorithm into the pheromone update of the ant colony algorithm, explores and utilizes the information in the software test space, discovers high-risk areas and generates test cases that can efficiently cover these areas, and adjusts the probability of a path or feature being selected through the pheromone concentration.
[0261] In addition, in a further embodiment, corresponding to Figure 1 , before performing step S103, it also includes a step of scenario classification. For example, calculate the following scenario classification features and conditions: Whether the target module or functional module has a high security risk: For example, the risk score of the module under test is greater than the threshold, or there are serious security vulnerabilities, sensitive data processing logics, etc. in the test history; The current test path or area is a high-risk path or a specific security policy branch that needs to be key-tested: For example, a path where encryption / decryption may occur, a key process for authentication / permission verification; A path or area that is particularly sensitive to security defects or high-risk inputs and requires generating targeted malicious inputs or extreme parameters to detect potential vulnerabilities: For example, in the known attack vector library, there are high-priority vulnerability prompts for the current path or area. When one or more of the above conditions are met, it can be determined as a high-risk security test scenario. When it is determined as a high-risk security test scenario, perform step S103 and generate a test case set according to the foregoing method. If it is determined that it is not a high-risk security test scenario, then determine whether other preset scenarios are met. For example, determine whether the complex path coverage scenario is met. If the complex path coverage scenario is met, use a method that combines the Ant Colony Optimization (ACO) and the Particle Swarm Optimization (PSO) to generate a test case set; Another example is to determine whether the input combination explosion scenario is met. If the input combination explosion scenario is met, use a method that combines the Genetic Algorithm (GA) and the Particle Swarm Optimization (PSO) to generate a test case set. If none of the obvious test scenarios are met, any of the foregoing methods or a method applicable to the general scenario can be used to generate a test case set.
[0262] Figure 5It is a principle block diagram of a software test case generation system according to an embodiment of the present invention. The system includes an initialization module 11, a user behavior model construction module 12, an input parameter space construction module 13, an initial set construction module 14, and an optimization module 15. Among them, the initialization module 11 is used to initialize the test environment to obtain the path exploration space of the target software. Among them, the path exploration space includes nodes and directed edges between the nodes. Two nodes with a directed relationship form a path branch, and each path branch corresponds to a path branch condition for implementing the path branch. The user behavior model construction module 12 is used to construct a user behavior model based on the collected user historical behavior data. The user behavior model includes one or more user parameters simulating user behavior and their parameter values. The input parameter space construction module 13 is used to construct the test parameter space of the target software. The input parameter space includes multiple input variables, and the variable value of each input variable corresponds to one or more user parameter values. The initial set construction module 14 is used to drive multiple ants to perform path exploration in different regions of the path exploration space, and update the pheromone information of the path branches during the path exploration process. After each ant determines a path branch during the path exploration process, based on the path branch condition, the corresponding input variable value is obtained from the input parameter space according to the user behavior model. After each ant finishes the path exploration, the input variable values of all path branches corresponding to the explored path are obtained to form an initial test case individual, and all initial test case individuals form an initial test case set. The optimization module 15 is used to perform iterative optimization based on the initial test case set. The optimization module 15 includes a selection unit 151, a reproduction unit 152, a pheromone update unit 153, and an evaluation unit 154. The selection unit 151 is used to screen the parent test case individuals from the current test case set based on the pheromone value and the fitness score of the test case individuals. When optimizing for the first time, the current test case set is the initial test case set. The reproduction unit 152 is used to perform a reproduction operation based on the parent test case individuals to obtain new test case individuals. Among them, the new test case individuals are merged into the current test case set to obtain a new test case set. The pheromone update unit 153 is used to update the pheromone values of the path branches covered by the parent test case individuals, and update the pheromone values of the path branches covered by the reproduced test case individuals. The evaluation unit 154 is used to evaluate whether the new test case set meets the optimization end condition. When the new test case set meets the optimization end condition, the iterative optimization is stopped and the new test case set is stored. When the new test case set does not meet the optimization end condition, the selection unit 151 is triggered to start a new round of optimization. For the specific process, please refer to the foregoing description of the method, which will not be elaborated here.
[0263] Figure 6It is a schematic structural diagram of a software testing system according to an embodiment of the present invention. The software testing system includes an operation terminal and a server terminal. Among them, the operation terminal is located on the terminal device 102, and the server terminal is located on the server 104 or a server cluster. The terminal device 102 communicates with the server 104 through a network. The terminal device 102 includes a desktop computer, a laptop computer or a mobile intelligent terminal device, such as a mobile phone, a tablet computer, etc. The server terminal includes the Figure 5 software test case generation system shown above. The operation terminal includes an interactive interface. When test cases need to be generated, testers configure corresponding parameters through the interactive interface and send them to the server terminal via the network. The server terminal generates a test case set according to the received instructions and corresponding parameters according to the method described above in the present invention. The parameters configured by the testers through the interactive interface are, for example, the target software name, version, user historical data storage address, number of iterations, storage address of the test case set, etc. After the test case set is generated, the tester can send a test instruction to the server terminal through the interactive interface. After receiving the test instruction, the server terminal executes software testing using the test case set and records the test results.
[0264] Figure 7 It is a schematic diagram of the hardware structure principle of an electronic device according to an embodiment of the present invention. The electronic device can be implemented as a server or various other terminal devices, such as a desktop personal computer, a tablet computer, a laptop computer, a mobile phone, etc. It includes a processor 601 and a memory 602. A program instruction set is stored on the memory 602. When the processor 601 executes the program instruction set on the memory 602, the above-mentioned software test case generation method is implemented.
[0265] Specifically, the above-mentioned processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0266] The memory 602 may include a mass storage for data or instructions. By way of example and not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 602 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 602 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 602 is a non-volatile solid state memory.
[0267] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the software test case generation method provided by the present invention.
[0268] In one example, the electronic device may further include a communication interface 603 and a bus 604. The processor 601, the memory 602, and the communication interface 603 are connected through the bus 604 to complete communication with each other.
[0269] The communication interface 603 is mainly used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present invention.
[0270] The bus 604 includes hardware, software, or both, and couples the components of the online data flow charging device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 604 may include one or more buses. Although the embodiments of the present invention describe and illustrate specific buses, the present invention contemplates any suitable bus or interconnect.
[0271] The present invention also provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, they can implement any one of the software test case generation methods in the foregoing embodiments. The computer-readable storage medium may be any medium that is tangible and contains or stores computer-executable instructions for use by or in connection with an instruction execution system, apparatus, and device. The storage medium may be a transitory computer-readable storage medium or a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media may include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Corresponding embodiments of such storage devices include, for example, magnetic disks, optical discs based on CD, DVD, or Blu-ray technology, and persistent solid-state memories such as flash memories, solid-state drives, and the like.
[0272] The present invention also provides a computer program product, which includes a set of computer program instructions. When the set of computer program instructions is executed by a processor, it implements any one of the software test case generation methods in the foregoing embodiments. The computer program product includes, but is not limited to, application installation packages, application plugins, applets that can run in certain applications, etc., published on websites and in app stores.
[0273] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps, after understanding the spirit of the present invention.
[0274] The above embodiments are only for illustrating the present invention and are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can make various changes and variations without departing from the scope of the present invention. Therefore, all equivalent technical solutions should also fall within the scope of the disclosure of the present invention.
Claims
1. A method for generating software test cases, characterized in that: include: Initialize the test environment to obtain a path exploration space of the target software; wherein the path exploration space includes nodes and edges with pointing between the nodes, two nodes with pointing relationships constitute a path branch, and each path branch includes a corresponding path branch condition; Based on the collected user historical behavior data, a user behavior model and an input parameter space of the target software are constructed, wherein the user behavior model includes one or more user parameters and parameter values simulating user behavior; the input parameter space includes a plurality of input variables, and the variable value of each input variable corresponds to one or more user parameter values; Drive multiple ants to explore paths in different areas of the path exploration space, and update the pheromone values of the path branches during the path exploration process; after each ant determines the path branch during the path exploration process, it selects the corresponding input variable value from the input parameter space based on the path branch condition and the user behavior model; after each ant finishes the path exploration, the input variable values of all the path branches of the explored path constitute an initial test case individual, and all the initial test case individuals constitute an initial test case set; Based on the initial set of test cases, the following iterative optimization steps are performed: The parent test case individuals are selected from the current test case set based on the pheromone value and the fitness score of the test case individuals; the current test case set during the first optimization is the initial test case set; Performing a breeding operation based on the parent test case individual to obtain a new test case individual, wherein the new test case individual is merged into the current test case set to obtain a new test case set; Update the pheromone values of the path branches covered by the parent test case individuals, and update the pheromone values of the path branches covered by the bred test case individuals; Evaluate whether the new test case set meets the optimization end condition, stop iterative optimization and store the new test case set when the new test case set meets the optimization end condition; repeat the aforementioned iterative optimization steps for the new test case set when the new test case set does not meet the optimization end condition.
2. The software test case generation method according to claim 1, characterized in that: The steps for each ant to determine the path branch during the path exploration process include: Based on the directional relationship of the edge between the two nodes, determine one or more candidate nodes that form a path branch with the current node; Calculate the selection probability from the current node to each candidate node based on the path branch pheromone value and the heuristic factor; and The candidate node with the highest selection probability is determined as the next node that forms a path branch with the current node; wherein the heuristic factor at least includes a user behavior matching weight and a security risk weight determined based on the user behavior model.
3. The software test case generation method according to claim 2, characterized in that: The heuristic factor also includes one or more of the following weights: a detection execution time weight, an operating environment weight, and a path length weight.
4. The software test case generation method according to claim 1, characterized in that: The steps for updating the path branch pheromone value during the path exploration process include: During the path exploration process, it is determined whether the currently explored path branch is marked with a security risk attribute tag. In response to the currently explored path branch being marked with a security risk attribute tag, a corresponding incremental value is increased based on the current pheromone value of the path branch marked with the security risk attribute tag.
5. The software test case generation method according to claim 1, characterized in that: Each ant further includes: Determine the scene corresponding to the current path exploration space area; and The pheromone values of the path branches are adjusted based on a manner that matches the scenario.
6. The software test case generation method according to claim 1, characterized in that: Further including: In the process of driving multiple ants to explore paths in different areas of the path exploration space, the coverage rate of the explored paths to the path exploration space or the preset local path exploration space is calculated, and the pheromone values of the path branches in the path exploration space are adjusted based on the corresponding relationship between the coverage rate and the negative correlation between pheromones.
7. The software test case generation method according to claim 1, characterized in that: After each ant finishes the path exploration, the input variable values of all path branches of the explored path constitute candidate initial test case individuals. The method further includes: Construct the maximum problem function of the initial test case set based on the environmental fitness and user behavior model; Configure the operating environment parameters for each candidate initial test case individual; Calculate the environmental fitness score of the candidate initial test case individuals based on the individual operating environment parameters of the candidate initial test case; Calculate the user behavior simulation score based on the user behavior model used when generating the candidate initial test case individuals; Calculating the maximum problem function value of the initial test case set based on the environmental fitness score and the user behavior simulation score of each candidate initial test case individual, and adjusting the candidate initial test case individual operating environment parameters and / or the user behavior model used to generate the candidate initial test case individual to make the maximum problem function converge to a maximum value; and The candidate initial test case individuals that converge to the maximum value are determined as the initial test case individuals, and all initial test case individuals constitute the initial test case set.
8. The method for generating software test cases according to claim 1, characterized in that: The steps of selecting the parent test case individual from the current test case set based on the pheromone value and the fitness of the test case individual include: Get the fitness score of each test case in the current test case set; Obtain the sum of the pheromone values of the path branches covered by each test case individual in the current test case set to obtain the pheromone value of each test case individual; Calculate the selection probability of each test case individual in the current test case set based on the pheromone value and fitness score; The test case individuals whose selection probability is greater than a threshold are used as the parent test case individuals screened out from the current test case set; or the test case individuals are sorted in descending order according to the selection probability, and a preset number of test case individuals in the front order are used as the parent test case individuals screened out from the current test case set.
9. The software test case generation method according to claim 8, characterized in that: When the current test case set is the initial test case set, the steps of obtaining the fitness score of each individual test case in the current test case set include: Calculate the coverage of each initial test case individual on the target software indicator; Evaluate each individual initial test case based on the threat model for security testing to obtain a defect disclosure capability value; and Calculate the weighted sum of the coverage rate and defect revealing capability value of each initial test case individual for the target software indicator as the initial fitness score of the initial test case individual; When the current test case set includes an optimized test case set, the step of obtaining the fitness score of each individual test case in the current test case set includes: Run each test case in the current test case set and collect the corresponding running data; Calculate the coverage of the target software indicator by the individual test case based on the running data of each individual test case; Evaluate each test case based on the threat model for security testing to obtain the defect disclosure capability value; Querying the user behavior model data and calculating the matching rate between the user behavior simulated by each test case individual and the user behavior model; and The fitness score of each test case is calculated as the weighted sum of the coverage of the target software indicator, the defect revealing capability value, and the matching rate between the user behavior simulated by the test case and the user behavior model.
10. The method for generating software test cases according to claim 8, characterized in that: The steps of performing a breeding operation based on the parent test case individual further include: Calculate the selection probability of each parent test case individual in the parent test case set; Calculate the cumulative probability of each parent test case individual in the parent test case set; Generate random numbers; and Select the parent test case individual corresponding to the position indicated by the random number for breeding operation.
11. The software test case generation method according to claim 1 or 10, characterized in that: The reproduction operation includes a crossover operation and / or a mutation operation.
12. The method for generating software test cases according to claim 11, characterized in that: When performing mutation operations based on the parent test case individuals, input variable values corresponding to the mutant gene categories are obtained from the input parameter space in a manner of simulating new user behaviors and / or triggering security risks.
13. The software test case generation method according to claim 1, characterized in that: After performing a breeding operation based on the parent test case individual to obtain a new test case individual, further comprising: Calculate the fitness score of the new test case individual; Traverse the fitness scores of each new test case individual, retain the new test case individuals whose fitness scores are greater than or equal to the threshold, and eliminate the new test case individuals whose fitness scores are less than the threshold; Alternatively, the new test case individuals are sorted in descending order of fitness scores; a preset number of new test case individuals that are sorted first are retained; Alternatively, a maximum fitness score problem is constructed; the maximum fitness score is obtained based on a new test case set, and the test case individuals with the maximum fitness score are retained.
14. The software test case generation method according to claim 1, characterized in that: The optimization end condition is that the number of iterations reaches a threshold or the fitness score of the current test case set reaches a threshold.
15. A software test case generation system, characterized in that: include: An initialization module, configured to initialize the test environment to obtain a path exploration space of the target software; wherein the path exploration space includes nodes and edges with pointing between the nodes, two nodes with pointing relationships constitute a path branch, and each path branch includes a corresponding path branch condition; A user behavior model building module, configured to build a user behavior model based on the collected user historical behavior data, wherein the user behavior model includes one or more user parameters and parameter values simulating user behavior; An input parameter space construction module configured to construct a test parameter space of the target software, wherein the input parameter space includes a plurality of input variables, and a variable value of each input variable corresponds to one or more user parameter values; The initial set construction module is configured to drive multiple ants to perform path exploration in different areas of the path exploration space, and update the path branch pheromone value during the path exploration process; after each ant determines the path branch during the path exploration process, based on the path branch condition, the ant obtains the corresponding input variable value from the input parameter space according to the user behavior model; after each ant completes the path exploration, the ant obtains the input variable values of all path branches corresponding to the explored path to form an initial test case individual, and all the initial test case individuals constitute an initial test case set; An optimization module, configured to perform iterative optimization based on an initial test case set, the optimization module comprising: A selection unit configured to select parent test case individuals from a current test case set based on pheromone values and fitness scores of the test case individuals; the current test case set during the first optimization is an initial test case set; A breeding unit, configured to perform a breeding operation based on the parent test case individual to obtain a new test case individual, wherein the new test case individual is merged into the current test case set to obtain a new test case set; a pheromone updating unit configured to update the pheromone values of the path branches covered by the parent test case individuals and to update the pheromone values of the path branches covered by the bred test case individuals; and An evaluation unit is configured to evaluate whether a new test case set satisfies an optimization end condition, stop iterative optimization and store the new test case set when the new test case set satisfies the optimization end condition; and trigger the selection unit to start a new round of optimization when the new test case set does not satisfy the optimization end condition.
16. An electronic device, characterized in that: The electronic device comprises a processor and a memory storing computer program instructions; when the electronic device executes the computer program instructions, the method for generating software test cases as described in any one of claims 1 to 14 is implemented.
17. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the software test case generation method as described in any one of claims 1 to 14 is implemented.
18. A computer program product, characterized in that It includes a computer program instruction set, which, when executed by a processor, implements the software test case generation method described in any one of claims 1 to 14.
Citation Information
Patent Citations
A combined approach to accelerate test case generation using genetic methods and symbolic execution.
CN109344057B