Software test case generation method and system, electronic equipment and storage medium

By building user behavior model and intelligent algorithm optimization generation of test case sets, the problem of time-consuming manual writing and low automation coverage is solved, and efficient and diversified software testing is achieved.

CN120256313APending Publication Date: 2025-07-04QIAN JIN NETWORK INFORMATION TECH SHANGHAI LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510476872.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, manual writing of test cases is time-consuming and difficult to adapt to the rapid iteration of software, and the automated generation method fails to effectively reflect user behavior, resulting in low test coverage and high cost.

Method used

A user behavior model is constructed based on user historical behavior data, an initial test case set is generated, and a test case set that simulates user behavior is generated through iterative optimization, and intelligent algorithms such as ant colony algorithm, particle swarm optimization and genetic algorithm are used to improve the coverage and effectiveness of test cases.

Benefits of technology

The generated test case set can fully reflect the actual use of the software, improve test effectiveness and coverage, adapt to different operating environments, and reduce manual intervention and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256313A_ABST
    Figure CN120256313A_ABST
Patent Text Reader

Abstract

The invention relates to a software test case generation method and system, electronic equipment and a storage medium. The method comprises the following steps: initializing a test environment to obtain running environment simulation configuration parameters and target software structure parameters; constructing a user behavior model and an input parameter space of target software based on the collected user historical behavior data; an initial test case set is constructed based on the input parameter space of the target software, the user behavior model and the test environment, and the initial test case set comprises a preset number of initial test case individuals; and performing iterative optimization on the initial test case individuals in the initial test case set until the optimized test case set meets an optimization ending condition, and storing the optimized test case set. According to the method, the test case set for simulating user behaviors can be generated, and the effectiveness of target software testing can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer software testing, and particularly to a method, system, electronic device and storage medium for generating software test cases. Background Art

[0002] In the field of software engineering, software testing is a process for discovering and fixing defects in software, verifying whether the software meets the design requirements, and ensuring its quality and performance. It is a crucial link in the software development process. Test cases are the core part of software testing, directly related to the efficiency and coverage of testing, and affecting the quality and reliability of software products. A test case can be regarded as a specific test task, containing elements such as input, operation, and expected output, and is a data set in various forms, such as execution paths, input data, and execution conditions.

[0003] The methods for generating test cases can generally be divided into the following two types:

[0004] First, test cases are manually written by testers. This means that testers need to have very rich testing experience and relatively high professional levels, and there are certain drawbacks such as blindness, high cost, and difficulty in improving test coverage. Although the method of manually writing test cases is effective in some cases, with the continuous increase in the complexity of software systems, it has become increasingly difficult to manually write test cases, time-consuming and prone to missing errors, especially in cases where high coverage and in-depth detection are required. In addition, the manual method is difficult to adapt to the rapid iteration of software. Updating a function may require rewriting a large number of test cases, which significantly increases the cost of software development and maintenance.

[0005] Second, test cases are automatically generated. For example, random testing, model-based testing (MBT), symbolic execution, search-based testing, and combining intelligent swarm algorithms in automated methods.

[0006] Although the aforementioned automated test case generation methods solve some drawbacks of the manual test case writing method, such as high requirements for manual labor, time-consuming, blindness, etc., these automated test case generation methods still have many unsatisfactory aspects. For example, model-based testing requires detailed model design, which is itself a complex and time-consuming process; symbolic execution and search-based testing face the problem of state space explosion and are difficult to handle large software systems. In addition, since these methods do not consider the user's behavior patterns and the diversity of software operating environments when generating test cases, the generated test cases cannot effectively reflect the real usage situation of the software, thus limiting the effectiveness of software testing. Summary of the Invention

[0007] In view of the technical problems existing in the prior art, the present invention provides a software test case generation method, system, electronic device and storage medium, which can generate a test case set simulating user behaviors and effectively improve the effectiveness of target software testing.

[0008] To solve the above technical problems, according to one aspect of the present invention, there is provided a software test case generation method, including the following steps:

[0009] Initialize a test environment to obtain operation environment simulation configuration parameters and target software structure parameters;

[0010] Construct a user behavior model and an input parameter space of the target software based on the collected user historical behavior data;

[0011] Construct an initial test case set based on the input parameter space of the target software, the user behavior model and the test environment, wherein the initial test case set includes a preset number of initial test case individuals; and

[0012] Iteratively optimize the initial test case individuals in the initial test case set until the optimized test case set meets the optimization end condition, and store the optimized test case set.

[0013] Optionally, the target software structure parameters include a dynamic call graph matrix of the target software; after obtaining the dynamic call graph matrix by initializing the test environment, it further includes:

[0014] Parse the dynamic call graph matrix to obtain two software units corresponding to each matrix element;

[0015] Construct a call flow graph with the software units as nodes and the call relationships between the software units as edges, wherein two nodes connected by an edge form a path branch; and

[0016] Parse the call relationships between the software units to obtain path branch conditions. The input operations and corresponding input parameter conditions.

[0017] Optionally, after obtaining the call flow graph by initializing the test environment, it further includes:

[0018] Obtain the security risk attribute marks of the software units.

[0019] Optionally, the step of constructing a user behavior model and an input parameter space of the target software based on the collected user historical behavior data includes:

[0020] Extract multiple types of user parameter values from the user historical behavior data, and the categories of the user parameters include user operation parameters, input parameters corresponding to the user operations, user identity parameters or environment parameters;

[0021] Generate a user behavior model based on the extracted user parameter values; and

[0022] Construct an input parameter space for the target software, where the input parameter space includes multiple input variables, and the variable value of each input variable includes one or more user parameter values.

[0023] Optionally, the step of generating a user behavior model based on the extracted user parameter values includes:

[0024] Calculate the distribution probability of each user parameter based on the extracted user parameter values and all user parameter values of the same user;

[0025] Alternatively, count the usage frequency of each user parameter value of the same user within a statistical time period;

[0026] Alternatively, count multiple user parameter value sequences representing different user behavior patterns.

[0027] Optionally, the step of constructing an initial test case set based on the input parameter space, user behavior model, and test environment of the target software includes:

[0028] Configure the operating environment parameters based on the operating environment simulation configuration parameters obtained from the initialized test environment;

[0029] Determine the initial test objective based on the target software structure parameters and user behavior model obtained from the initialized test environment;

[0030] Determine the input variables in the input parameter space that can achieve the initial test objective; and

[0031] Select input variable values that can simulate user behavior for each input variable from the input parameter space according to the user behavior model;

[0032] Among them, the operating environment parameters and one or more input variable values corresponding to an initial objective constitute an initial test case individual;

[0033] The initial test objective is a path branch in the call flow graph or a path composed of multiple path branches.

[0034] Optionally, when determining the initial test objective, select the path branch corresponding to the software unit with a security risk attribute mark at the node; correspondingly, after determining the input variables in the input parameter space that can achieve the initial test objective, it includes: determining the input variable values from the input parameter space based on the security test trigger conditions.

[0035] Optionally, the step of constructing an initial test case set based on the input parameter space, user behavior model, and test environment of the target software further includes:

[0036] Construct a maximum problem function for the initial test case set based on the environmental fitness and user behavior model;

[0037] Calculate the environmental fitness score of each initial test case individual based on the operating environment parameters of the initial test case individual;

[0038] Calculate the user behavior simulation score based on the user behavior model for generating the initial test case individual;

[0039] Adjust the operating environment parameters of one or more initial test case individuals in the initial test case set and / or the user behavior model for generating the initial test case individual to make the maximum problem function converge to the maximum value;

[0040] Determine the test case set when converging to the maximum value as the initial test case set.

[0041] Optionally, during the iterative optimization of the initial test case individuals in the initial test case set, the steps for optimizing the current test case set once include:

[0042] Select some test case individuals from the current test case set as seed test case individuals, and multiple seed test case individuals form a seed set; and

[0043] Reproduce the seed test case individuals in the seed set to generate new test case individuals, and the seed test case individuals and the new test case individuals form the test case set for the next optimization.

[0044] Optionally, the steps for selecting some test case individuals from the current test case set as seed test case individuals include:

[0045] Obtain the fitness scores of each test case individual in the current test case set;

[0046] Traverse the fitness scores of each test case individual, and use the test case individuals with fitness scores greater than the score threshold as seed test case individuals or use the test case individuals ranked before the sorting threshold as seed test case individuals.

[0047] Optionally, when the current test case set is the initial test case set, the steps for obtaining the fitness scores of each test case individual in the current test case set include:

[0048] Calculate the coverage rate of each initial test case individual for the target software metrics;

[0049] Evaluate each initial test case individual based on the threat model for security testing to obtain the defect revelation ability value of each initial test case individual; and

[0050] Calculate the weighted sum of the coverage rate and defect revelation ability value of each initial test case individual for the target software metrics as the initial fitness score of the initial test case individual;

[0051] When the current test case set is the optimized test case set, the steps to obtain the fitness score of each test case individual in the current test case set include:

[0052] Run each test case individual in the current test case set and collect the corresponding running data;

[0053] Calculate the coverage rate of the test case individual for the target software metrics based on the running data of each test case individual;

[0054] Evaluate each test case individual based on the threat model for security testing to obtain the defect revelation ability value of each test case individual;

[0055] Query the user behavior model data and calculate the matching rate between the user behavior simulated by each test case individual and the user behavior model; and

[0056] Calculate the weighted sum of the coverage rate of each test case individual for the target software metrics, the defect revelation ability value, and the matching rate between the user behavior simulated by the test case individual and the user behavior model as the fitness score of each test case individual.

[0057] Optionally, the end condition is to reach the iteration count threshold or the fitness score of the current test case set reaches the threshold.

[0058] Optionally, when breeding the seed test case individuals in the seed set to generate new test case individuals, it includes one or more of the following steps:

[0059] Change one or more path branches in the path corresponding to the seed test case to generate a new test case individual;

[0060] Change the running environment parameters of the seed test case to generate a new test case individual;

[0061] Change the input operations or input operation sequences and / or corresponding input parameters used to simulate user behavior in the seed test case.

[0062] According to another aspect of the present invention, the present invention also provides a software test case generation system, including:

[0063] An initialization module configured to initialize the test environment to obtain the running environment simulation configuration parameters and the target software structure parameters;

[0064] A parameter construction module, configured to construct a user behavior model and an input parameter space of a target software based on the collected user historical behavior data;

[0065] An initial set construction module, configured to construct an initial test case set based on the input parameter space of the target software, the user behavior model, and a test environment, where the initial test case set includes a preset number of initial test case individuals; and

[0066] An optimization module, configured to iteratively optimize the initial test case individuals in the initial test case set until the optimized test case set meets the optimization end condition, and store the optimized test case set.

[0067] According to another aspect of the present invention, the present invention also provides an electronic device, which includes a processor and a memory storing computer program instructions; when the electronic device executes the computer program instructions, the foregoing any software test case generation method is implemented.

[0068] According to another aspect of the present invention, the present invention also provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing any software test case generation method is implemented.

[0069] According to another aspect of the present invention, the present invention also provides a computer program product, including a set of computer program instructions, and when the set of computer program instructions is executed by a processor, the foregoing any software test case generation method is implemented.

[0070] Embodiments of the present invention construct a user behavior model based on user historical behavior data. When automatically generating test cases, the selection of input parameters is guided by the user behavior model, so that the generated test case individuals can simulate user behavior, can comprehensively reflect the actual usage of the software, and improve the effectiveness of target software testing; the test case individuals in the test case set have different running environment parameters, so the test case set can adapt to different software running environments and provide diverse test scenarios for testing. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Next, the preferred embodiments of the present invention will be further described in detail with reference to the drawings, where:

[0072] Figure 1 is a flowchart of a software test case generation method according to an embodiment of the present invention;

[0073] Figure 2 is a flowchart of a method for constructing an initial test case set according to an embodiment of the present invention;

[0074] Figure 3 It is a schematic diagram of the path from the product home page to a specific product page according to an embodiment of the present invention;

[0075] Figure 4 It is a flowchart of a method for optimizing test case individuals to obtain an initial test case set according to an embodiment of the present invention;

[0076] Figure 5 It is a flowchart of a method for iteratively optimizing test case individuals in a test case set according to an embodiment of the present invention;

[0077] Figure 6 It is a flowchart of a method for selecting seed test case individuals for reproduction using a roulette wheel selection mechanism according to an embodiment of the present invention;

[0078] Figure 7 It is a flowchart of a method for generating test cases according to another embodiment of the present invention;

[0079] Figure 8 It is a flowchart of a method for generating test cases by fusing the ant colony algorithm (ACO) and the particle swarm optimization (PSO) algorithm according to an embodiment of the present invention;

[0080] Figure 9 It is a flowchart of a method for generating test cases by fusing the genetic algorithm (GA) and the particle swarm optimization (PSO) algorithm according to an embodiment of the present invention;

[0081] Figure 10 It is a flowchart of a method for generating test cases by fusing the ant colony algorithm (ACO) and the genetic algorithm (GA) according to an embodiment of the present invention;

[0082] Figure 11 It is a schematic block diagram of the principle of a software test case generation system according to an embodiment of the present invention;

[0083] Figure 12 It is a schematic block diagram of the principle of a software test case generation system according to another embodiment of the present invention;

[0084] Figure 13 It is a schematic diagram of the architecture of a software test system according to an embodiment of the present invention; and

[0085] Figure 14 It is a schematic diagram of the hardware structure principle of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0086] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0087] In the following detailed description, reference can be made to the various specification drawings that form a part of this application and illustrate specific embodiments of the application. In the drawings, like reference numerals generally describe substantially similar components in different figures. The various specific embodiments of the present application have been described in sufficient detail below to enable those of ordinary skill in the relevant art and technology to implement the technical solutions of the present application. It should be understood that other embodiments can also be utilized or structural, logical, or electrical changes can be made to the embodiments of the present application.

[0088] Figure 1 is a flowchart of a software test case generation method according to an embodiment of the present invention. In this embodiment, the software test case generation method includes the following steps:

[0089] Step S101, initialize the test environment to obtain the running environment simulation configuration parameters and the target software structure parameters.

[0090] Step S102, construct a user behavior model and the input parameter space of the target software based on the collected user historical behavior data.

[0091] Step S103, construct an initial test case set. Specifically, construct an initial test case set based on the input parameter space of the target software, the user behavior model, and the test environment, where the initial test case set includes a preset number of initial test case individuals.

[0092] Step S104, perform iterative optimization on the initial test case individuals in the initial test case set until the optimization end condition is met.

[0093] Among them, the running environment simulation configuration parameters in step S101 include three types of parameters: the operating system version (abbreviated as O), the network condition (abbreviated as N), and the hardware configuration data (abbreviated as H). Each type of running environment simulation configuration parameter can include one or more running environment simulation configuration parameters. In order to enable test cases to test the target software in various environments, the present invention thus simulates the running states of different environments through the software running environment simulation configuration parameters to provide diverse test scenarios for testing. In one embodiment, each running environment simulation configuration parameter is set as an environmental factor e i, since the software running environment (or the system environment of the target software) has different impacts on different test cases, in order to measure this impact, an environment fitness function E(C, v) is defined. The fitness of a test case individual v to the environment under a specific system configuration is measured by the environment fitness function E(C, v), so as to quantify the impact of the software running environment on the test case individual v. Among them, C represents the set of software running environment simulation configuration parameters, C = {O, N, H}, O represents the operating system version, N represents the network condition, and H represents the hardware configuration data. The present invention adopts the set of environmental factors e i E = {e1, e2,..., e n} to represent a specific system simulation configuration. Each environmental factor e i can correspond to a set of one or more running environment simulation configuration parameters. Each environmental factor e i is assigned a corresponding weight w i . Based on the comprehensive impact I of the fitness function E(C, v) on the environment configuration (abbreviated as the environmental impact value), it is calculated by the following formula (1-1):

[0094]

[0095] where f(e i ) is the impact value of the environmental factor e i . Among them, the calculation method of the impact value f(e i ) of the environmental factor e i is a conventional technique in the field of software testing and will not be elaborated here.

[0096] In one embodiment, the target software structure parameters in step 101 include a dynamic call graph matrix. When the present invention executes the initialization step in step S101, it initializes the dynamic call graph of the internal structure of the target software system and uses the matrix representation method M(t) in graph theory to describe the interaction and dependency relationships between software units. Specifically, the interaction and dependency relationships between software units are described by formula (1-2):

[0097] M(t) = M(t - 1) + ΔM(t); (1-2)

[0098] M(t) represents the call matrix of software units within the time period t, M(t - 1) represents the call matrix of software units within the time period t - 1, and ΔM(t) represents the change in call data within the time period t. The dynamic change situation of the software during operation is represented by the matrix of the call data of software units.

[0099] Among them, software units, such as components, modules, and even more detailed functions, etc., serve as the rows and columns of the call matrix, and the matrix elements represent the call data between software units within a preset time period, such as the number of times or frequency. In one embodiment, the matrix element M(t) [i,j] represents the number of times or frequency that the i-th software unit calls (jumps / calls) the j-th software unit within the time period t. For example, if "Component A" and "Component B" are mapped to the i-th row / column and the j-th column / row, then M(t) [i,j] is the number of times Component A calls Component B within the time period t. When further refining "component" to the "function" level, M(t) [f1,f2] represents the number of times function f1 calls function f2 within the time period t.

[0100] As time accumulates, M(t), M(t + 1),... can be merged or superimposed to obtain the overall call matrix M_total. Thus, it can be seen that the call graph matrix not only includes the structural composition of the software, but also can obtain whether the call relationships of certain components / modules / functions are covered, and can also obtain their call frequencies and control flows.

[0101] By parsing the dynamic call graph matrix to obtain two software units corresponding to each matrix element; using the software units as nodes and the call relationships between software units as edges to construct a call flow graph, the call flow graph includes nodes and edges with directions between the nodes, where two nodes connected by an edge with a direction relationship form a path branch. According to the call relationships parsed between software units, obtain the input operations and corresponding input parameter conditions for implementing the path branch, hereinafter referred to as branch conditions.

[0102] Specifically, query the elements greater than 0 in the current call graph matrix. If the count of a certain row and a certain column in M(t) is greater than 0, it means that the two software units represented by the corresponding row and column have indeed made calls during the test. Therefore, it can be obtained that these two software units have "triggered" the corresponding function call actions during the test. When modeling the internal structure of the software, use the two software units that have "triggered" the call actions as nodes, and use the directed call relationship of "software unit i calls software unit j" as the "edge" connecting software unit i and software unit j, thus forming a path branch. According to this method, connect all the call relationships where the matrix elements (M(t) [i,j] ) are not 0 in series, then all the paths actually traversed during the test are obtained, which constitute the call flow graph, and a continuous series of software units and their call relationships constitute a dynamic call chain. If it does not appear (or is 0) in the matrix, it means that the call relationship between the corresponding two software units has not been tested and executed.

[0103] Therefore, based on the dynamic call graph matrix and the call flow graph, the coverage of modules / functions and the coverage of paths or path branches during the testing process can be obtained.

[0104] Since there is a specific mapping relationship between software units and code, the code file or code line corresponding to the software unit can be obtained. Based on the matrix element M(t) [X,Y] ≠0, by combining with the coverage tool at the code level, it can be confirmed that "X calls a specific code branch or code line in Y".

[0105] In addition, security is usually also an important aspect in software testing. In one embodiment, when software units such as software units or code segments are marked with security risk attribute marks, the security risk attribute marks of the software units are obtained during initialization, so as to determine test case individuals that can perform security tests at this location when generating test cases.

[0106] In step S102, each piece of data in the user historical behavior data includes at least user identity information, time information, user operation information, and corresponding input information. These parameters representing various types of information are collectively referred to as user parameters. According to the category, user parameters include user operation parameters, input parameters corresponding to user operations, user identity parameters, and environment parameters. The parameter values of user operation parameters can be various user operations, such as click, select, input, etc.; the parameter values of input parameters are, for example, the specific content corresponding to the user operation, such as the confirmation key information corresponding to the click confirmation key operation, the commodity information corresponding to the select operation, the input content corresponding to the input operation, etc.; the parameter values of user identity parameters are, for example, information such as username and password; the environment parameters are, for example, one or more types of information among the operating system version, network conditions, and hardware configuration.

[0107] For example, for an online shopping platform system, a piece of user historical behavior data collected is as follows (time information is omitted):

[0108] User operations and corresponding input information:

[0109] Search keyword - smartphone;

[0110] Select commodity category - electronics;

[0111] Select commodity price range - (1000 - 2000) yuan;

[0112] User account information:

[0113] Username - user123;

[0114] Password - pass456;

[0115] Environment parameters:

[0116] Device type of the user - Android phone;

[0117] Operating system version - Android 12;

[0118] Network condition - (Wi-Fi, bandwidth 10Mbps).

[0119] The parameter values of the user operation parameters extracted from the above data are "search keywords", "select product category", "select product price range", etc. The corresponding input parameters are, for example, specific keywords (such as smart phones), specific product categories (such as electronic products), etc. The parameter values of the environmental parameters include the user device type ("Android phone"), the user device operating system version ("Android 12"), the network condition used by the user ("Wi-Fi, bandwidth 10Mbps"), etc. The parameter value of the user identity parameter is specific account information, that is, the user name "user123" and the password "-pass456".

[0120] The user behavior model in the present invention is various data or combinations that can simulate user behavior. For example, the distribution probability and usage frequency of specific types of user parameter values, the sequence composed of multiple user parameter values representing different user behaviors; the association relationship between different types and / or the same type of user parameter values representing user behavior patterns.

[0121] Taking the user operation parameter as an example, the parameter value of the user operation parameter is a specific type of user operation, and the distribution probability of each user operation is calculated based on formula (1-3).

[0122]

[0123] Among them, x is a specific type of user operation, and n x is the number of user operations x extracted from the historical data in a preset time period (such as one week, three days, one month, etc.), and N is the total number of all user operations of the user in the same time period. User operation x is, for example, "search keywords", "select product category" or "select product price range" in the foregoing embodiments, etc. By calculating the distribution probability of each specific user operation, it is possible to distinguish whether a user operation conforms to the user's habitual operation. Similarly, the distribution probabilities of various other types of user parameters can also be obtained, such as the distribution probabilities of various specific input parameter values, the distribution probabilities of various environmental parameter values, etc.

[0124] The usage frequency of user parameters can also reflect corresponding user behavior patterns. For example, by counting the button click frequency of the same user within their respective statistical time periods at different times, the user's button clicking habit can be obtained, such as clicking 5 times per minute. Another example is that according to the average number of product pages viewed by different users during a single search process, a user behavior pattern of viewing 10 pages per visit can be obtained.

[0125] Sequences composed of multiple user parameter values can also represent corresponding user behaviors. For example, for an e-commerce platform, a user behavior model obtained is: login - click on product 1 - click on product 2 - click on product 3 - click on product 4 - click on product 5 - click on the purchase button on the product page... Another user behavior model obtained is: login - select product category and search for product keywords - click on product 1 - click on product 2 - click on add to cart - click on product 3 -... - click on the purchase button... The sequences composed of multiple categories of user parameter values representing the above two different user behaviors are each a user behavior model, and these two user behavior models represent different user behavior patterns.

[0126] Another example is that when the user behavior model is a sequence of specific user parameter values such as the operation of clicking on a product - specific product information - the operating system in the environmental parameters, the user behavior models composed of different specific parameter values represent different user behavior patterns.

[0127] The distribution probability of each of the above user parameter values, the usage frequency of each user parameter value, and sequences of multiple user parameter values are collectively referred to as the user behavior model.

[0128] The input parameter space of the target software includes multiple input variables, and the input variables correspond to user parameters. The variable value of each input variable includes one or corresponding multiple user parameter values. In one embodiment, the input parameter space of the target software is represented by a set of variables: X = {x1, x2,..., x n}; X represents the input parameter space, and the variable x iRepresents the i-th input variable. The input variable can be an input operation or input parameter that the target software can accept. Here, the input operation corresponds to the user operation in the user behavior model, and the input parameter value is the specific information, data, etc. that the user can input, corresponding to the input parameter value corresponding to the user operation. For example, the input operation can be a login operation, inputting a search keyword, clicking on a relevant link, and the input parameter value can be the content that can be input for an input operation, such as a specific keyword, a specific link, etc. The input variable can also be a user identity parameter, such as a username and password. The input variable can also be an environmental parameter, such as a user device parameter, the user device operating system, network configuration or conditions, etc. The specific variable values of these input variables form the complete space of the input parameters of the target software. For example, the input variable x1 represents the username, and its value can be "user123" or "admin", etc.; the input variable x2 represents the user operating system, and the parameter value can be "Windows10", "Android12", etc., the input variable x3 represents the login operation, the input variable x4 represents the operation of inputting a search keyword, the input variable x5 represents the keyword, and its value is one or more keywords in the keyword table, and so on.

[0129] The test case in the present invention refers to a specific instance for a complete test task, including multiple component parameters, and is the smallest unit in the test case set. In the following description, in order to highlight the relationship with the test case set, a specific instance for a complete test task is referred to as a test case individual, and when not emphasizing the relationship with the test case set, a specific instance for a complete test task is referred to as a test case.

[0130] To achieve a complete test task, the component parameters that make up the test case are of multiple types and need to meet certain conditions. For example, the component parameters of the test case should at least include specific input operations and corresponding input parameter values. The input operation can be one or multiple and consecutive, that is, an input operation sequence. Referring to the paths and path branches in the call flow graph, an input operation and the corresponding input parameter value in the test case can implement a transition from one node to another node, that is, implement a path branch. When the test case includes multiple and consecutive input operations, multiple consecutive path branches can be implemented, that is, constitute a path.

[0131] In addition, the component parameters of the test case can also include environmental parameters and, if necessary, user identity parameters. The environmental parameter is used to define the running scenario, and the user identity parameter is a special input parameter.

[0132] In step S103, when generating initial test case individuals to construct an initial test case set, according to the categories and specific variable values of input variables that can be provided in the input parameter space, various user behavior models, and the current test environment, a preset number of initial test case individuals are generated according to the conditions that need to be satisfied by the composition of the test cases.

[0133] See Figure 2 , Figure 2 FIG. is a flowchart of a method for constructing an initial test case set according to an embodiment of the present invention. In this embodiment, the method for constructing an initial test case set based on the input parameter space, user behavior model, and test environment of the target software includes the following steps:

[0134] Step S1031, configure the running environment parameters according to the running environment simulation configuration parameters obtained based on the initialized test environment. Define the running scenario of the target software when running the test case individuals by configuring the running environment parameters for the initial test case set. Among them, the running environment parameters include one or more of the operating system O, network condition N, and hardware configuration H. For example, a specific running scenario is obtained by setting the network condition to "10 Mbps with a delay of 200 ms". Specific running scenarios can be configured for each initial test case individual, or multiple running scenarios can be configured for the initial test case individuals in the initial test case set, and multiple initial test case individuals are run in each running scenario. When the running environment parameters are configured, the test environment is deployed based on the running environment parameters in the subsequent optimization process.

[0135] Step S1032, determine the initial test target based on the target software structure parameters and user behavior model obtained from the initialized test environment. The initial test target is, for example, a path branch in the call flow graph or a path composed of multiple path branches. The user behavior models provided by the present invention have various forms, such as the aforementioned specific user operations or the distribution probability of input parameter values, the usage frequency of user operations, specific operation sequences, etc. And since each user operation can usually implement a path branch, when determining the path as the initial test target, the path or path branch is determined with reference to the user behavior model. For example, see Figure 3 , Figure 3 FIG. is a schematic diagram of a path from the product home page to a specific product page according to an embodiment of the present invention. From Figure 3It can be seen that there are multiple paths corresponding from the product home page (node i) to a specific product page (node j). By clicking on the specific product operation (input operation a), the path branch from node i to node j is realized. It can also be indirectly realized from node i to node j through the operation of searching for products (input operation b) by combining different search conditions (conditions c, d, etc.) through other nodes (node j1, node j2, or node j3). In one embodiment, different paths are first formed according to the starting node and the ending node, and then the paths are decomposed to obtain multiple consecutive path branches. The corresponding input operation and the corresponding input parameters are determined according to the corresponding path branch conditions, so as to determine the input operation sequence, and then compared with the operation sequence provided by the user behavior model, and a path that best matches the user behavior model is selected as the initial test target.

[0136] Step S1033, determine the input variables that can achieve the initial test target from the input parameter space. Each path branch has a corresponding branch condition. Through the branch condition, the required input operation and the input parameter condition can be determined. The input parameter space is queried based on the input operation to obtain the corresponding input variables.

[0137] Step S1034, select the input variable values that can simulate the user behavior for each input variable from the input parameter space according to the user behavior model.

[0138] The user behavior model provided by the present invention has various forms, such as the distribution probability of the aforementioned user operations or input operations, the usage frequency of user operations, specific operation sequences, etc. Therefore, first, based on all the input variables corresponding to the initial test target, the form of the user behavior model to be used is determined. When only one input variable is involved, the input operation or input parameter value with the highest distribution probability that can be used as the input variable value is selected. For example, when the path is from the login page to the user's personal home page, since a login variable is required, the corresponding input operation is the login operation, and the corresponding input parameters include two input parameters, namely the username and the password. There are multiple usernames and multiple passwords in the input space. At this time, when there are multiple input parameter values, the username and password with the highest distribution probability are selected as the input parameter values, so that the obtained test case individual is: [click the login button; input parameters: username "user123" + password "pass456"].

[0139] When there are multiple input operations corresponding to the input variable, select the input operation with the highest distribution probability or the highest usage frequency in the user behavior model. For example, for the aforementioned Figure 3Node i, the corresponding input operations can be "click product operation a" and "search operation b". Referring to the user behavior model, the distribution probability of "search operation b" is the highest, so "search operation b" is selected. For another example, when the determined path is the path from node i through node j1 to node j, it includes two consecutive input operations "search operation" and "click operation". According to the user behavior model: [search for "smartphone", select the "electronic product" category, browse products with a price range of "1000 - 2000 yuan"], the specific input operation values are determined, and the obtained test case individual is: [search for "smartphone"; select the "electronic product" category, price range "1000 - 2000 yuan"; click "xxx phone"].

[0140] Since the user behavior model also provides the usage frequency of input parameters, further, the input parameters can be selected according to the usage frequency of input parameters. For example, according to the user behavior model, when the usage frequency of the input parameter that the user visits 10 pages on average each time when browsing the product page is obtained, add the browsing operation steps of 10 product pages in the generated test case. For example, the generated test case simulates that the user uses a "wired network" on a computer with the "Windows10" system to access the shopping platform and randomly clicks on 10 product pages.

[0141] For another example Figure 3 , obtained based on the call flow graph Figure 3 the path shown, but based on the user behavior model, when the user clicks on a product, and the frequency of clicking on a product when the user browses the product is 5 each time, 4 randomly selected products can be added on this basis, that is, the obtained test case is: [search for "smartphone"; select the "electronic product" category, price range "1000 - 2000 yuan"; click "xxx phone 1"; click "xxx phone 2"; click "xxx phone 3"; click "xxx phone 4"; click "xxx phone 5"].

[0142] Finally, in order to specify the running scenario of the currently generated test case individual, usually the above-mentioned set running environment parameters can also be added, so as to obtain a test case with complete composition parameters, such as [username: user123; password: pass456; search for "smartphone"; select the "electronic product" category; select the price range "1000 - 2000 yuan"; click "xxx phone 2"; click "xxx phone 3"; click "xxx phone 4"; click "xxx phone 5"; operating system version: Windows10; network configuration: Wi-Fi, latency 50ms]. Of course, other more detailed running environment parameters such as hardware conditions can also be included.

[0143] The present invention generates an initial test case set by generating a preset number of initial test case individuals according to the foregoing steps. From the foregoing process of generating initial test case individuals, it can be seen that the input operations, operation sequences, and selection of input parameters for constructing test case individuals are guided by the user behavior model, so that the generated test case individuals can simulate the real operations of users and reflect the real usage of the target software during testing. In addition, in a further embodiment, when guiding the input operations, operation sequences, and selection of input parameters of test case individuals by the user behavior model, high-frequency or high-distribution probability user input operations and input parameter values are preferentially used as the input operations and input parameter values of test cases; or high-frequency or high-distribution probability input operations and low-frequency or low-distribution probability but highly valuable user operations or input parameter values are respectively used in a certain proportion, so as to achieve diversification and comprehensiveness of testing.

[0144] Further, when generating a preset number of initial test case individuals, paths in different regions of the call flow graph are selected as initial test targets, so that the initial test case set can evenly cover different regions for subsequent optimization processing.

[0145] Further, when a security risk attribute mark in a path branch or software unit is obtained during initialization, when determining the initial test target in step S1032, a path branch or software unit with a security risk attribute mark is selected to generate a part of the initial test case individuals, so that the generated test cases can cover some path regions with security risks. Among them, in order to generate test cases that can cover paths with security risks and test whether they contain security vulnerabilities or security threats, when selecting input parameters in step S1033, input parameters are selected based on the triggering conditions of security testing (such as a specific vulnerability). For example, for the login operation, in order to obtain a security test case for SQL injection, the user name in the corresponding input parameters is selected as "admin'OR'1'='1". If the output is login failure when performing the login operation with this user name, it means that the system of the target software can prevent SQL injection.

[0146] During initialization, the present invention further obtains a threat model for security testing. The threat model is a security module used to analyze and process certain security vulnerabilities or security threats, such as the Common Vulnerability Scoring System (CVSS for short). The threat model can determine whether a vulnerability is detected based on the triggering conditions corresponding to the security vulnerability, and quantitatively score or rate the risk of the vulnerability when the vulnerability is detected. For example, the threat model determines whether a test case triggers a vulnerability through symbolic execution (such as SQL injection detection returning 1 / 0), and quantifies the severity s of the security vulnerability and the likelihood a of the vulnerability being exploited through the risk scoring function R(s,a) shown in Formula 1-4.

[0147] R(s,a) = s × a (1-4)

[0148] Among them, s represents the vulnerability severity, s ∈ [0, 10], a represents the likelihood of the vulnerability being exploited (calculated through Monte Carlo simulation), a ∈ [0, 1], and the product of the two is used as the comprehensive quantitative risk value. For example, for a certain Cross Site Scripting (XSS) vulnerability, R(s,a) = 0.7 × 8.5 + 0.3 × (probability 0.9) ≈ 7.2, which is determined to be a high-risk according to the corresponding rating criteria.

[0149] The types of security vulnerabilities in the present invention include SQL injection, cross-site scripting attacks, buffer overflows, etc.

[0150] Furthermore, in order to make the initialized test cases closer to real user behaviors and real operating scenarios, better target the final test objectives, and make subsequent iterative optimization easier and converge more quickly, in another embodiment of the present invention, when constructing the initial test case set, a batch of initially generated test case individuals are used as candidate initial test case individuals, a maximum value problem function is constructed based on the user behavior model and environmental fitness, and by adjusting the candidate initial test case individuals, or adjusting the user behavior model and / or operating environment parameters of the candidate initial test case individuals, the maximum value problem function is made to converge to the maximum value; the candidate initial test case individuals when converging to the maximum value are determined as the initial test case individuals.

[0151] In one embodiment, the maximum value problem function is as shown in Formula 1-5 below:

[0152]

[0153] Among them, P(x i ) is a specific input variable value x in the test case iThe distribution probability, that is, the probability of being used, is calculated according to formula (1-3). n is the total number of various input variable values in all test cases. E(C,v) is the environmental impact value I of the operating environment parameters in each test case on the test case, which is calculated according to formula (1-1). m is the total number of operating environment parameters in the current test case set V. α and β are trade-off coefficients used to balance the impacts of the user behavior model and environmental configuration on the initialization process.

[0154] The present invention measures the degree to which the overall test case set approximates real user behavior and real operating scenarios based on the user behavior model and simulated operating environment configuration parameters, and optimizes the individual test case v according to the measurement result.

[0155] See Figure 4 , Figure 4 is a flowchart of a method for optimizing test cases to obtain an initial test case set according to an embodiment of the present invention.

[0156] Step S1031a, obtain an individual test case from the test case set V as the target test case individual.

[0157] Step S1032a, calculate the distribution probability P(x i ) of each input variable value in the target test case individual. Specifically, for each input variable value, its distribution probability is calculated using formula (1-3).

[0158] Step S1033a, calculate the impact value of the simulated operating environment configuration in the target test case individual on the target test case individual. Specifically, each simulated operating environment configuration parameter in the target test case is taken as an environmental factor e, and the environmental impact value I of the current operating environment parameter on the target test case individual is calculated using formula (1-1).

[0159] Step S1034a, determine whether there are still individual test cases in the test case set V. If so, return to step S1031a. If there are no individual test cases in the test case set V, then execute step S1035a.

[0160] Step S1035, calculate the weighted sum of the probability sum of all input variable values of all current test case individuals and the environmental impact value I.

[0161] Step S1036a, optimize the test case individuals to obtain new test case individuals. For example, randomly change the input variable values of one or more test case individuals in the test case set V in the direction of increasing probability and / or randomly change the running environment parameters of one or more test case individuals in the test case set V in the way of increasing the environmental impact value I, and use them as the new target test case individuals. For example, applying a genetic algorithm, each test case individual is regarded as a chromosome in the genetic algorithm, and each input variable value and each running environment configuration parameter in the test case individual are regarded as a gene in the genetic algorithm respectively. In the process of optimizing the test case individuals, the input variable values and / or running environment parameters in a test case individual are changed through gene crossover and mutation operations to obtain new test case individuals.

[0162] Step S1037a, calculate the distribution probability P(x i ) and the environmental impact value I of each input variable value in the new target test case.

[0163] Step S1038a, calculate the weighted sum of the probabilities P(x i ) of all input parameters of all test case individuals in the current test case set V and the environmental impact value I.

[0164] Step S1039a, determine whether the weighted sum of the current test case set V has converged to the maximum value. If it has converged to the maximum value, end the process; otherwise, return to Step S1036a. Here, "converging to the maximum value" means that the weighted sum of the current test case set V no longer increases, that is, by comparing the size of the current weighted sum with the weighted sum calculated last time, when the weighted sum no longer increases, it is confirmed that it has converged to the maximum value.

[0165] Furthermore, after obtaining the initial test case set, evaluate the test effectiveness of each test case individual. In one embodiment, calculate the fitness score of each test case individual through the following fitness function expression (1-6). In this embodiment, the test effectiveness is evaluated by the software metric coverage ability and the ability to discover potential defects.

[0166] f(v) = γ·Coverage(v) + δ·Defects(v) (1-6)

[0167] Coverage(v) is the software metric coverage rate of a single test case individual v, usually a percentage, ranging from (0 - 1). In one embodiment, the software metric is, for example, a path. When specifically calculating, the number of path branches included is determined through the path corresponding to the test case individual v, and the ratio of the number of path branches of the test case individual v to the total number is calculated. The path branch ratio is used as the coverage rate Coverage(v). The software metric can also be, for example, lines of code. Based on the test path corresponding to the test case individual v, the relevant software structure is determined, and the number of lines of code is determined based on the correspondence between the software structure and the code. The ratio of the number of lines of code corresponding to the test case individual v to the total number of lines of code is calculated, and the line of code ratio is used as the coverage rate Coverage(v). Or, the software metric can also be, for example, a functional module of the software. The ratio of the number of functional modules covered by the test case individual v to the total number of functional modules is calculated, and the functional module ratio is used as the coverage rate Coverage(v). For example, if the target software system has 100 functional modules and the test case individual v covers 20 of them, then Coverage(v) = 20%.

[0168] Defects(v) is the number of potential defects that a single test case individual v can reveal, representing the defect detection ability of a single test case individual v. The defect detection ability Defects(v) is a positive integer, and its specific range depends on the test scenario and system complexity. For example, its calculation method is: through the aforementioned security threat module, symbolic execution is used to analyze the potential defect triggering ability of the test case individual v for the target software, such as predicting the types and quantities of vulnerabilities that can be exposed.

[0169] γ and δ are weight coefficients, used to balance the contributions of the software metric coverage rate and the defect detection ability to the fitness.

[0170] In this embodiment, the weight coefficients can be set as needed. If the coverage rate is given priority, γ > δ can be set. If defect discovery is emphasized, δ > γ is set.

[0171] For example, in a medium - sized or small - sized software system, for a test case, its Coverage(v) can be set between 0.1 and 0.5 (i.e., 10% - 50%), and its Defects(v) can be set between 0 and 5 (the number of detected defects is limited).

[0172] For example, set γ = 0.6 and δ = 0.4. If the coverage rate of the test case individual v is 0.3 (30%) and 2 defects are detected, then the fitness score is: f(v) = 0.6·0.3 + 0.4·2 = 0.18 + 0.8 = 0.98. This score indicates that the test effectiveness performance of this individual is very good and is suitable for being retained in the current set for further optimization.

[0173] In a large and complex software system, for a test case individual v, Coverage(v) may be between 0.01 and 0.3. Defects(v) may be between 0 and 10 (it is easier to discover defects).

[0174] For example, set γ = 0.7 and δ = 0.3. If the coverage rate of the test case individual v is 0.1 and 5 defects are detected, then the fitness score is: f(v) = 0.7·0.1 + 0.3·5 = 0.07 + 1.5 = 1.57. This score indicates that the test case individual v is relatively effective.

[0175] Take the score calculated for each test case individual through expression (1-6) as the initial fitness score of the test case individual. The higher the score calculated based on the fitness function expression (1-6), the better the performance of the test case individual v in terms of both the software function coverage degree and the ability to discover potential defects, that is, the better the test effectiveness, the higher the potential value, and the more effectively it can discover software defects or verify software functions.

[0176] In one embodiment, after the weighted sum of the current test case set V converges to the maximum value, determine the candidate initial test case individuals whose initial fitness scores meet the requirements as the initial test case individuals. For example, sort the test case individuals in the current test case set V according to the initial fitness scores, and take a certain number of test case individuals ranked at the front as the initial test case individuals, so as to obtain the final initial test case set. Or, take the test case individuals whose initial fitness scores are greater than the threshold as the initial test case individuals, so as to obtain the final initial test case set.

[0177] In addition, referring to the evaluation of the test effectiveness of each test case, in another embodiment, it is also possible to conduct an overall evaluation of the test effectiveness of the current test case set V. Specifically, see the fitness function expression (1-7):

[0178] f(V) = γ·Coverage(V) + δ·Defects(V) (1-7)

[0179] Among them, Coverage(V) represents the software metric coverage rate of the test case set V, reflecting the coverage of the test case set V for the target software. Defects(V) represents the number of potential defects that can be revealed through the test case set V, reflecting the effectiveness of the defect detection ability of the test case set. γ and δ are trade-off coefficients used to balance the contributions of the coverage rate and the defect detection ability to the fitness.

[0180] The initial test case set of the present invention has a certain degree of diversity and can initially cover key or common use case scenarios. To a certain extent, it targets or takes into account goals such as function coverage, security testing, and user behavior coverage. Therefore, in subsequent iterative optimization (step S104), in combination with fitness evaluation (such as coverage rate, security vulnerability discovery ability, user behavior matching degree, etc.), these initial test case individuals are continuously subjected to optimization steps such as crossover, mutation, and screening, gradually approaching or continuously improving the satisfaction degree of the test target, avoiding a sudden expansion of the search space during optimization, thus reducing the search difficulty and also improving the optimization efficiency.

[0181] See Figure 5 , Figure 5 which is a flowchart of a method for iteratively optimizing test case individuals in a test case set according to an embodiment of the present invention. In this embodiment, the method includes:

[0182] Step S1041: Select some test case individuals from the current test case set as seed test case individuals, and multiple seed test case individuals form a seed set.

[0183] Step S1042: Breed the seed test case individuals in the seed set to generate new test case individuals, and the seed test case individuals and the new test case individuals form a new test case set.

[0184] Step S1043: Evaluate whether the newly obtained test case set in the current optimization meets the optimization end condition. If it meets, store the test case set obtained in the current optimization in step S1044 and stop iterative optimization. If it does not meet, execute step S1045.

[0185] Step S1045: Use the new test case set as the test case set for the next round of optimization, and return to step S1041.

[0186] Among them, in step S1041, some test case individuals with high fitness scores can be selected from the current test case set as seed test case individuals according to the fitness scores, or seed test case individuals can be selected from the current test case set based on the roulette wheel selection mechanism.

[0187] In the first optimization, the current test case set is the initial test case set, and the fitness score of the test case individual is the initial fitness score f(v) of the initial test case individual, that is, the initial fitness score calculated according to the foregoing formula (1-6). In the subsequent optimization process, when the current test case set is the test case set updated after the previous round of optimization, the fitness score of the test case individual is the weighted sum of the coverage rate of the target software metrics, the defect revelation ability value, and the user behavior matching rate calculated based on the running data obtained by running the test cases in the previous round of optimization.

[0188] For example, the fitness score F(v) is calculated based on the following formulas (1-8):

[0189]

[0190] where covered(i,v), detect(d,v), and match(u,v) are indicator functions respectively. The three terms in formula (1-8) respectively correspond to the three criteria for evaluation: the coverage criterion, the security criterion, and the user behavior matching criterion. The three criteria are respectively used to judge the coverage rate of the test case individual v on software metrics (such as lines of code, functions, modules, and path coverage and conditions), the ability to detect security threats or vulnerabilities, and the degree of matching with user behavior. λ1, λ2, and λ3 are the weights for the three criteria respectively.

[0191] Specifically, c is a coverage criterion, C is the set of coverage criteria. The set of coverage criteria C includes various coverage criteria for lines of code, functions, modules, and path coverage and condition coverage. i is a coverage unit, such as a line of code or a code item, a function, a branch, a module, a path, or a condition, etc. I c is the set of coverage units for the coverage criterion c, w c is the weight for the coverage criterion c, |I c | is the total number of coverage units i, and covered(i,v) is an indicator function. If the test case individual v covers the coverage unit i, it is 1, otherwise it is 0.

[0192] s is the type of security test, S is the set of security test types, d is a vulnerability or security threat (hereinafter referred to as a security test item), D s is the set of security test items for the security test type s, w s is the weight for the security test type s, detect(d,v) is an indicator function. If the test case individual v can detect the security test item v, it is 1, otherwise it is 0, T s is the detection threshold for the security test type s, and e is the base of the natural logarithm.

[0193] b is the user behavior pattern, B is the set of user behavior patterns, U b is the set of user operation sequences for the user behavior pattern b, w b is the weight for the user behavior pattern b, p(u) is the probability of the user operation sequence u occurring, and match(u,v) is a matching function that reflects the matching degree between the test case individual v and the user operation sequence u.

[0194] In step S1041, test case individuals that do not meet the requirements are eliminated by the fitness score. The conditions to be met are, for example, that the fitness score is greater than the threshold, or, when sorted by the fitness score, it needs to be ranked among the top N. The test case individuals that meet the conditions are used as test case individuals that can reproduce, and those that do not meet the conditions are eliminated, which not only avoids a large number of test cases but also retains high-quality test cases, accelerating the convergence speed of optimization.

[0195] When breeding the seed test case individuals in the seed set to generate new test case individuals in step S1042, each seed test case individual can be directly bred, or some seed test case individuals can be selected for breeding. When selecting some seed test case individuals, a certain number of seed test case individuals can be randomly selected for breeding, or the roulette wheel selection mechanism can be used to select seed test case individuals from the current test case set.

[0196] See Figure 6 , Figure 6 is a flowchart of a method for selecting seed test case individuals for breeding using the roulette wheel selection mechanism according to an embodiment of the present invention. It includes the following steps:

[0197] Step S10421, obtain the fitness score of each test case individual in the seed set.

[0198] Step S10422, calculate the selection probability p select (v) of each test case individual in the seed set.

[0199] Step S10423, construct the cumulative probability P select (v) of each test case individual.

[0200] Step S10424, generate a random number, and select the corresponding test case individual based on the position where the random number is located.

[0201] Step S10425, determine whether the number of selected test case individuals has reached the preset quantity. If it has reached, end. If it has not reached the preset quantity, return to step S10424.

[0202] Among them, in step S10421, the fitness score of the test case individual can be the initial fitness score calculated based on formula (1-6) or the fitness score calculated based on formula (1-8) as described above.

[0203] In step S10422, in one embodiment, the selection probability p select (v) of the test case individual can be calculated by the following formula (1-9):

[0204]

[0205] Among them, p select (v) represents the probability that the test case individual v is selected for reproduction; p select (v) ∈ [0, 1], and the sum of the selection probabilities of all test case individuals v is 1; F(v) is the fitness score of the test case individual v. In the first optimization, it is the initial fitness score, and in subsequent optimization processes, it is the fitness score calculated by applying formula (1-8) and reflecting the quality of the test case individual v from three aspects: coverage rate, security testing ability, and user behavior matching degree. The higher F(v) is, the better the quality of the test case is, and the greater the probability of being selected; Population represents the current test case set, that is, the seed set obtained after being selected in step S1043.

[0206] ∑ v′∈Population F(v) represents the sum of the fitness scores of all test case individuals in the seed set.

[0207] In another embodiment, in order to enhance the influence of the fitness difference of test case individuals and make the influence of the fitness difference of test case individual v on the selection probability more significant. Calculate the selection probability p select (v) of the test case individual based on formula (1-10):

[0208]

[0209] T is a regulation parameter used to adjust the sensitivity of the selection mechanism to the fitness difference. When the value of T is large, the probability distribution is smoother, and individuals with low fitness have a certain probability of being selected, which helps to maintain the species characteristics. When T is small, the selection is more concentrated on individuals with high fitness, accelerating convergence, but it may lead to premature convergence to a local optimum. Control the smoothness of the selection probability distribution by adjusting the parameter T, allowing individuals with lower fitness scores to also have a certain probability of being selected. Among them, in the selection process of the same round of optimization, the same regulation parameter T is used for all test cases to calculate, so as to ensure the implementation of a consistent selection strategy for the entire population, facilitate the comparison of the fitness of different individuals and the assignment of selection probabilities. For different iteration rounds, the value of T can be dynamically adjusted to enhance or weaken the "amplification" or "smoothing" effect on the fitness difference.

[0210] In this embodiment, through the operation of e F(V) / T , the fitness score of the test case individual v is exponentiated to enhance the influence of the fitness difference of the test case individual and make the influence of the fitness difference of the test case individual v on the selection probability more significant.

[0211] In step S10423, in order to calculate the cumulative probability P of the test case individualselect (v), queue all the test case individuals and label them with serial numbers n, where n = 1... N, and N is the total number of test case individuals in the seed set. For any test case individual v in the queue i , take the test case individual v i 's selection probability p select (v i ) and the sum of the selection probabilities p select (v) of all the test case individuals sorted before it as its cumulative probability P select (v i ), that is Thus, N cumulative probability values are obtained, forming a cumulative probability value queue. Then in step S10424, the range of the random number is 1 to N. Take the random number as the sorting of the cumulative probability value queue, so as to obtain the test case individual corresponding to the position indicated by the random number.

[0212] In step S1042, when breeding the seed test case individuals in the seed set to generate new test case individuals, there can be various methods. For example, changing the operating environment parameters of the seed test case to generate new test case individuals; changing one or more input variable values of the seed test case to generate new test case individuals, or changing one or more path branches in the path corresponding to the seed test case to generate new test case individuals.

[0213] In some other specific implementation manners, the mutation operation or crossover operation in the genetic algorithm can also be used to generate new test case individuals, and the advantage that high-quality genes can be inherited to the next generation in the genetic algorithm can be utilized to generate higher-quality test case individuals, thereby accelerating the convergence of the optimization, that is, meeting the end condition of the optimization. When using the genetic algorithm, similar to constructing the initial test case set, each test case individual is regarded as a chromosome, and each input variable value and each operating environment configuration parameter in the test case are regarded as chromosome genes. By gene crossover and mutation operations, the input variable values and / or operating environment parameters in a test case individual are changed to obtain new test case individuals.

[0214] For example, the crossover operation is performed through the following formula (1-11), and the mutation operation is performed through the following formula (1-12):

[0215] v(t + 1)' i = crossover(v(t) i , v(t) r1 , cr) (1-11)

[0216] where crossover represents the crossover operation, cr is the crossover rate, and v(t + 1)' iRepresents the test case individual after crossover, v(t) i Is the selected test case individual for crossover, v(t) r1 Is another randomly selected test case individual for the crossover operation.

[0217] v(t + 1)″ i = mutate(v(t) i , mr) (1 - 12)

[0218] Where mutate represents the mutation operation, mr is the mutation rate, and v(t) i Is the selected test case individual for mutation. v(t + 1)″ i Represents the test case individual after mutation.

[0219] In this embodiment, the number, ratio, position, etc. of genes for crossover or mutation can be specified, so that the specific input variable values (such as user identity parameters, input operations, input parameters, etc.) and / or environmental parameters for crossover or mutation can be determined.

[0220] For input parameters with a parameter value range, the mutation operation of the input parameters can be performed using formula (1 - 13):

[0221] R mutated = R + μ·(rand([-1, 1])·(max(R) - min(R))) (1 - 13)

[0222] Where R mutated Represents the mutated input parameter value; R represents the original input parameter value; rand([-1, 1]) represents generating a random number within the range of [-1, 1]. By introducing randomness in the present invention, the mutation direction and amplitude of each test parameter are different; max(R) and min(R) respectively represent the upper and lower limits of the input parameter value range, which are used to control the range of the mutation operation to ensure that the input parameter value does not exceed the acceptable limit. By changing the mutation rate μ, the purpose of diversity control can be achieved. For example: setting a higher μ: making the input parameter span larger. For example, for the user name and password in the input parameters, completely different user name formats can be tried, while for some numerical parameters, a lower mutation rate μ can be set: fine-tuning near the original parameter. For example, changing the time interval of the login request, fine-tuning the price range of the goods clicked by the user, etc.

[0223] By mutating the input operation, input parameters, user identity parameters, operating environment parameters, etc., the present invention can simulate user behavior, meet the triggering conditions of security testing, increase the running scenarios, increase the coverage of software functions, code, functions, branches, conditions, paths, etc.

[0224] For example, an original test case individual is as follows:

[0225] Input parameters: {"user123","pass456"};

[0226] Input operations: Click the navigation bar three times after logging in;

[0227] Operating environment parameters: The operating system is "Windows10" and the bandwidth is 10Mbps.

[0228] During mutation, the input parameters can be mutated. For example, mutate the username "user123" in the original input data to user'OR'1'='1 for testing SQL injection.

[0229] When mutating the input operations, the user behavior pattern can be adjusted. For example, mutate the original three clicks on the navigation bar to five clicks on the navigation bar and add a search operation.

[0230] Mutate the operating environment parameters to adjust the environment configuration. For example, the original configuration has a bandwidth of 10Mbps; after mutation, the bandwidth becomes 5Mbps, so as to simulate a network fluctuation scenario.

[0231] The final mutated test case is as follows:

[0232] Input data: {"user'OR'1'='1", "pass456"};

[0233] Input operations: Click the navigation bar five times and perform a search;

[0234] Operating environment parameters: Operating system "Windows10", bandwidth 5Mbps.

[0235] Therefore, it can be seen that during the mutation operation, by mutating the input parameters (such as username and password), multiple possible vulnerability trigger conditions can be generated, generating test case individuals that are more likely to trigger security vulnerabilities, thereby expanding the security test coverage. In addition, test case individuals simulating other user behaviors can also be generated through the mutation operation, making the test process closer to the real user operation habits and improving the user behavior coverage rate. For example, by changing the input operations in the test case, new behavior patterns can be generated. For example, when mutating the number of clicks and paths, new operation habits can be obtained, increasing the diversity of test cases. Through the mutation operation of the operating environment, the running states of different environments can be simulated, providing diverse test scenarios for the execution of the test.

[0236] After breeding new test case individuals, update the test case set to obtain a new test case set, and the updated test case set reflects the changes in breeding. Among them, when updating the test case set, in one embodiment, for the newly added test case individuals, calculate their fitness scores F(v) according to formula (1-8) to measure their performance in terms of coverage, security testing ability, and user behavior simulation. Then, sort them according to the fitness scores F(v), and retain the top N test case individuals with higher fitness scores for the next round of optimization, and eliminate the test case individuals with lower fitness. Or, in another embodiment, form a candidate test case set by the seed test case individuals and the new test case individuals, obtain the maximum value through the following formula (1-14), and use the test case individual corresponding to the obtained maximum value as the updated test case set.

[0237]

[0238]

[0239] Among them, Population combined is a combination of the original test case set and the newly generated test case individuals after breeding, and the total number of test case individuals in it is greater than the preset total number N. Population new is the updated test case set.

[0240] As can be seen from the foregoing steps, as the number of iterations increases, those test case individuals that can more effectively discover defects, cover more paths, and more conform to real user behaviors will be retained and further propagated, gradually eliminating the test case individuals with low fitness until the iteration end condition is finally met and test indicators such as "coverage reaches a certain threshold", "successfully detect specified or certain risk vulnerabilities", and "user behavior coverage meets the standard" are achieved.

[0241] The optimization end conditions in step S104 and step S1043 are, for example, that the current test case set meets the system test objectives, and the system test objectives include, for example, software system test objectives, security test objectives, and user behavior test objectives. For the software system test objectives, their corresponding coverage criteria, for example, specifically determine whether the current test case set covers the specified software internal structure and reaches the specified coverage rate, whether it reaches the code coverage rate, function call coverage rate, and conditional judgment coverage rate of modules or components, and whether it reaches the specified environmental adaptability (E(C, V)); for the security test objectives, they correspond to security criteria, for example, whether the specified vulnerabilities or threats are detected, or whether the specified number or proportion of vulnerabilities or threats are detected. For the user behavior test objectives, they correspond to user behavior matching criteria, for example, whether the specified user behavior coverage and matching degree are reached. In one embodiment, the fitness score of the test case set is calculated according to the following formula (1-15):

[0242]

[0243] where C, S, and B respectively represent the set of coverage criteria for the software system, security test types, and user behavior patterns, and w c , w s and w b are the weight coefficients corresponding to each type; c is a coverage criterion, C is the set of coverage criteria, and the set of coverage criteria C includes various coverage criteria for lines of code, functions, modules, path coverage, and conditional coverage. Coverage c (V) is the score of the software units (such as various coverage criteria for lines of code, functions, modules, path coverage, and conditions) covered by the test case set V calculated according to formula (1-16).

[0244]

[0245] where I c is the set of covered units for the coverage criterion c, |I c | is the total number of covered units i, and covered(i, V) is an indicator function that is 1 if the test case set covers the covered unit i and 0 otherwise.

[0246] s is the security test type, S is the set of security test types; Safety s (V) is the score calculated based on the following formula (1-17).

[0247]

[0248] where d is a vulnerability or security threat (hereinafter referred to as a security test item, such as SQL injection, cross-site scripting attack, buffer overflow, etc.), and D s is a set of security test items for the security test type s. detect(d, v) is an indicator function that is 1 if the test case individual v can detect the security test item v, and 0 otherwise. T s is the detection threshold for the security test type s, and e is the base of the natural logarithm.

[0249] b is the user behavior pattern, and B is the set of user behavior patterns; the user behavior coverage degree Behavior b (V) is the score calculated based on formula (1-18), which is used to evaluate the degree of coincidence between the test case set and the user behavior model, and reflects the effect of the test case set in simulating real user behavior.

[0250]

[0251] where U b is the set of user operation sequences for the user behavior pattern b, p(u) is the probability of the user operation sequence u occurring, and match(u, V) is a matching function that reflects the matching degree between the test case set V and the user operation sequence u. The user behavior coverage degree Behavior b (V) reflects the user behavior coverage degree and matching degree, that is, the user behavior covered by the user behavior in the test case and the corresponding matching degree.

[0252] When the fitness score of the test case set reaches the threshold, it is considered that the optimization end condition is satisfied. At this time, the optimization is stopped, and the current test case set V final is stored, and V final ={V|V∈Populationat termination}. At this time, the finally obtained test case set V final contains the test case individuals with the highest fitness after multiple rounds of iterative optimization. In addition, the optimization end condition also includes the number of iterations. When the number of iterations is reached, the optimization is also stopped. When the optimization is terminated only by the number of iterations, the finally generated test case set V final can also be evaluated to ensure that it meets the preset system test objectives.

[0253] See Figure 7 , Figure 7 is the flowchart of the method for generating test cases according to another embodiment of the present invention. In this embodiment, the software test case generation method includes the following steps:

[0254] Step S201: Initialize the test environment to obtain the operating environment simulation configuration parameters, the target software structure parameters, and the threat model for security testing.

[0255] Step S202: Construct a user behavior model and the input parameter space of the target software based on the collected historical user behavior data.

[0256] Step S203: Obtain the scenario classification features.

[0257] Step S204: Determine whether the conditions formed by the classification features meet the classification conditions of Scenario 1; if they meet the conditions of Scenario 1, execute Step S301; if they do not meet the classification conditions of Scenario 1, execute Step S205.

[0258] Step S205: Determine whether the conditions formed by the classification features meet the classification conditions of Scenario 2; if they meet the conditions of Scenario 2, execute Step S401; if they do not meet the classification conditions of Scenario 2, execute Step S206.

[0259] Step S206: Determine whether the conditions formed by the classification features meet the classification conditions of Scenario 3; if they meet the conditions of Scenario 3, execute Step S501; if they do not meet the classification conditions of Scenario 3, it is determined as Scenario 4, and then execute Step S103.

[0260] Among them, in the processing flow starting from Step S301, a method that combines the Ant Colony Optimization (ACO) algorithm and the Particle Swarm Optimization (PSO) algorithm is used to generate test cases.

[0261] In the processing flow starting from Step S401, a method that combines the Genetic Algorithm (GA) and the Particle Swarm Optimization (PSO) is used to generate test cases.

[0262] In the processing flow starting from Step S501, a method that combines the Ant Colony Optimization (ACO) algorithm and the Genetic Algorithm (GA) is used to generate test cases.

[0263] In this embodiment, in order to give full play to the advantages of various intelligent swarm algorithms, the present invention classifies the test scenarios into four categories: the first category is the complex path coverage scenario, the second category is the input combination explosion scenario, the third category is the high-risk security testing scenario, and the fourth category is the general scenario. By obtaining the scenario classification features in Step S203 and determining the category of the current scenario according to the classification conditions of various scenarios, and then using the corresponding specific algorithms to generate a test case set.

[0264] Among them, the features for scenario classification include, for example: path complexity (such as metrics like McCabe's Cyclomatic Complexity, maximum branch depth, number of loop nesting levels, etc.), scale of the input parameter space (such as the size of the Cartesian product of the value ranges of input parameter values, total number of parameter dimensions), importance level of functions / modules, security level (such as high-risk security modules, core business logic modules), known risks, distribution of vulnerabilities (such as risk scores obtained from historical defect data or security vulnerability databases), time / resource constraints (such as available test execution time, hardware resource limitations, concurrency limitations, etc.), historical coverage or test case execution results (whether there is already a high coverage, or whether high-priority defects are frequently triggered, etc.).

[0265] The scenario classification features and conditions for determining the complex path coverage scenario (hereinafter referred to as Scenario 1) include, for example: the McCabe complexity is greater than or equal to the threshold or the average branch depth of the software module is greater than or equal to the threshold;

[0266] The combined quantity of input variable values corresponding to a software unit in the input parameter space is less than the threshold (for example, the input variable dimension is low, or the value range of the parameter value is small); the number of path branches in the target area of the call flow graph is greater than the threshold; the area where the existing test case coverage is not ideal; the area where the coverage is less than the threshold during the historical test process, or the known high-risk paths that have not been fully tested. When any two or more of the above classification conditions are met, it is determined as the complex path coverage scenario. Among them, the McCabe complexity is also known as the Cyclomatic Complexity, and its measurement method is well-known to those of ordinary skill in the art and will not be elaborated here.

[0267] The scenario classification features and conditions for determining the input combination explosion scenario (hereinafter referred to as Scenario 2) include, for example: the dimension or combination quantity of the input parameters is greater than the threshold or the Cartesian product number of all variable parameter values is greater than the threshold; the path complexity is not very high, but a higher "breadth" coverage requirement for the input parameter space is needed: such as the McCabe complexity is less than the threshold and conventional process / function coverage needs to be satisfied; the time / resource is relatively loose, allowing multiple generations of iteration to explore the huge input space: such as the available test execution time is greater than or equal to the threshold. When one or more of the above conditions are met, it can be determined as the input combination explosion scenario.

[0268] Scenario classification features and conditions for determining high-risk security test scenarios (Scenario Three) include, for example: the target module or function has high security risks: for example, the risk score of the module under test is greater than the threshold, or there are serious security vulnerabilities, sensitive data processing logics, etc. in the test history; the current test path or area is a high-risk path or a specific security policy branch that needs to be focused on for testing: for example, paths where encryption / decryption may occur, key processes for authentication / permission verification; paths or areas that are particularly sensitive to security defects or high-risk inputs and require generating targeted malicious inputs or extreme parameters to detect potential vulnerabilities: for example, in the known attack vector library, there are high-priority vulnerability prompts for the current path or area. When one or more of the above conditions are met, it can be determined as a high-risk security test scenario.

[0269] See Figure 8 , Figure 8 is a method flow chart for generating test cases by a method that combines the ant colony algorithm (ACO) and the particle swarm optimization (PSO) algorithm according to an embodiment of the present invention. In this embodiment, the software test case generation method includes the following steps:

[0270] Step S301, drive multiple ants to perform path exploration in different regions of the path exploration space to obtain an initial test case set.

[0271] Step S302, map the initial test case set to an initial particle swarm. Among them, each initial test case individual corresponds to the initial position of a particle.

[0272] Step S303, move the particle position to obtain a new particle position, and the new particle position is a new test case individual.

[0273] Step S304, eliminate test case individuals whose fitness scores do not meet the requirements, and the test case individuals that meet the fitness score requirements form a new test case set.

[0274] Step S305, determine whether the new test case set meets the optimization end condition. If it meets the optimization end condition, end the optimization and store the current test case set. If it does not meet, return to Step S303.

[0275] In Step S301, the path exploration space is, for example, constructed from all path branches in the call flow graph in the foregoing embodiment. Before driving multiple ants to perform path exploration in different regions of the path exploration space, it is first necessary to set various required ant colony parameters, such as the initial pheromone values of each path branch, the pheromone evaporation rate ρ, the calculation method of the pheromone intensity constant Q, the number of ants, the number N of initial test case individuals to be obtained, the exploration termination condition, etc. Among them, according to the number of ants and the number N of test cases to be obtained, configure the exploration times of each ant.

[0276] In one embodiment, pheromone initial values are set for paths in the path exploration space based on a user behavior model. For example, first, the corresponding user operation is determined based on the matching relationship between the paths in the path exploration space and the user operations. For example, each path branch formed by each node (such as a page or a state) among multiple nodes in a page browsing path corresponds to a user operation. Then, the user behavior model is queried based on the user operation, and the pheromone initial values of the path branches are set according to the distribution probability of the user behavior operations in the user behavior model. If the distribution probability is large, the pheromone initial value is large; if the distribution probability is small, the pheromone initial value is small. Thus, when ants explore paths, they can select paths that match the user behavior better.

[0277] The path exploration process of each ant specifically includes: First, a starting node is determined for each ant, and then the probabilities of moving from the starting node to all possible next nodes are calculated according to the probability formula. Among them, the probability formula is shown as the following formula (1 - 19):

[0278]

[0279] τ ij : The pheromone concentration from node i to node j, which is calculated according to formula (1 - 20). is the probability from node i to node j, and l is one of the multiple candidate nodes, that is, one of all the allowed nodes.

[0280]

[0281] Among them, t is the time parameter in the ant colony algorithm. Corresponding to the present invention, it is the iteration number when the ant conducts path exploration; in this embodiment, when generating the initial test case set, among the termination conditions for ending the process, the number of test case individuals can be used as the termination condition, or the aforementioned condition of meeting the maximum value convergence can be used as the termination condition, or the aforementioned condition that the initial fitness score is greater than the threshold can be used as the termination condition. Therefore, it is necessary for the ant to conduct path exploration multiple times, that is, multiple iterations. τ ij (t + 1) is the pheromone value from node i to node j (referred to as path L i,j ) after being updated through the t-th exploration and used for the (t + 1)-th exploration, τ ij (t) represents the pheromone concentration of path L i,j at the t-th exploration, ρ is the pheromone evaporation rate, is the pheromone value released (left) by the k-th ant when exploring path L i,j , and it is also the pheromone increment left after one ant explores on path L i,j . Q is the pheromone intensity constant, L kis the path length of the k-th ant.

[0282] η ij is the heuristic factor from node i to node j. The heuristic factor includes at least the user behavior weight determined based on the user behavior model. For example, η ij = α · user operation probability. Additionally, to also consider guiding the ant to choose a path with more security risks, a security risk weight can be added, such as η ij = α · user operation probability + β · security risk value. Of course, other weights can also be set according to the needs of testing, such as test execution time weight, running environment weight, path length weight, business importance weight, high-frequency feature weight, etc., so as to introduce other concerned factors when selecting the next node. Specifically, it can be flexibly used according to the nature of the target software in actual applications. α and β are the weights of the pheromone and the heuristic factor respectively. Similarly, τ il is the pheromone from node i to node l, and η il is the heuristic factor from node i to node l.

[0283] One node with the largest selection probability is selected as the next node to obtain a path branch. After each ant determines the path branch during the path exploration process, it selects the corresponding input variable from the input parameter space, and then selects the corresponding input operation and input parameter value from the input parameter space according to the user behavior model as the input variable value. And so on until the termination condition of the path exploration is reached, such as reaching the preset path length (such as at most 10-step operations), or entering the termination state (such as triggering a crash, completing the core process), or covering the target branch or meeting the coverage threshold, etc.

[0284] The complete search trajectory of each ant is the explored path, and the input operations and input parameter values corresponding to all path branches constitute a test case individual.

[0285] Update the pheromone value of the path branch during the ant's path exploration process. Specifically, in the case where the current explored path branch is marked with a security risk attribute mark, increase the pheromone increment of the path branch marked with the security risk attribute mark; or it can also judge the scenario corresponding to the current path exploration space area and adjust the pheromone value of the path branch based on the matching method with the scenario; or it can also calculate the coverage rate of the explored path to the path exploration space or the preset local path exploration space during the process of driving multiple ants to conduct path exploration in different areas of the path exploration space, and adjust the pheromone value of the path branch in the path exploration space based on the corresponding relationship between the coverage rate and the pheromone being negatively correlated. To implement the above adjustment of the pheromone value, it can be in the pheromone intensity constant Q or the path length L kIncrease the calculation weights of the foregoing various factors, so as to adjust the value of the pheromone left by the ants in the foregoing various situations. Thus, the purpose of guiding other ants to explore the predetermined area is achieved.

[0286] In step S302, map the initial test case set to an initial particle swarm. Among them, each initial test case individual corresponds to the initial position of the particle. When mapping, first map the test case to a multi-dimensional position vector of the particle, and each dimension corresponds to an input variable of the test case. Among them, the position of the particle corresponds to the specific test case individual. For example, the dimension parameter configuration of a particle x1 is: [input field length, number of concurrent requests, timeout threshold], and the dimension parameter configuration of a particle x2 is: [user name length, password complexity, network delay, operating system].

[0287] Then perform position encoding on each particle, that is, encode the specific dimension parameter values into numbers. If the dimension parameter value is a continuous numerical type, such as network delay, input value, etc., directly use the continuous parameter value. If the dimension parameter value is a discrete type, such as operating system type, operation sequence, encode it into a discrete value. When encoding into a discrete value, integer encoding can be performed. For example, for the operating system type, different numbers are used to represent different operating systems, such as: 0 = Windows, 1 = macOS, 2 = Linux, 3 = Android. Multi-segment encoding can also be performed. For example, divide the operation sequence dimension into multiple sub-dimensions, and each sub-dimension represents the selection of an operation step.

[0288] For example, for a particle x2 with a specific dimension parameter configuration of: [user name: User123; password: 456789; network delay: 200ms; operating system: Android], according to its position encoding rule, the user name length is 6 characters, the password complexity is a weak password, and the corresponding encoding is 0. Among them, the encoding of a medium-strength password pair is 1, and the encoding of a strong password pair is 2. The network delay is the corresponding number 200, and the encoding of the operating system is 3. Therefore, the particle x2 [user name: User123; password: 456789; network delay: 200ms; operating system: Android] is encoded as [6; 0; 200; 3], which is the current position of the particle x2.

[0289] Then, determine its position range, speed range, and initial speed according to the dimension parameter type. The position range is, for example, the value range of each dimension parameter. For example, for the first dimension parameter "username length" of particle x2, its value range is 1 - 20. The speed of each particle is a vector, and its dimension corresponds one-to-one with the dimension of the particle's position vector (i.e., the input variable values in the test case individual). Each component of the speed represents the change amount of the input variable value corresponding to the dimension. When the data type of the dimension parameter is continuous numerical type, such as input value, network latency, etc., the speed component is a real number, representing the increase or decrease amplitude of the dimension parameter value. When the range of the input value is [0, 100], the speed range can be set to ±10%, then the maximum adjustment amplitude of the speed component in each iteration is ±10. When the data type of the dimension parameter is discrete numerical type, such as operating system type, operation steps, etc., the speed component needs to be converted into the probability of discrete selection or adjustment strategy. For example, for the operating system type [0 = Windows, 1 = macOS, 2 = Linux, 3 = Android], the probability of selecting the current operating system (such as macOS) is reduced and the probability of adjacent options (such as Windows) is increased by using the Softmax function or roulette wheel selection method. Another example is to use the method of integer speed conversion to round the speed represented by a decimal to an integer, so as to select the parameter value corresponding to the integer. For example, when v = 0.7, it is converted to v = 1, and then the corresponding macOS is selected.

[0290] In step S303, based on the current position of the particle, the individual best position of the particle, and the global pheromone highest position, calculate the current moving speed of the particle according to the following formula (1 - 22).

[0291] v i (t + 1) = w·v i (t) + c1·rand1·(pbest i -x i (t)) + c2·rand2·(τ gbest -x i (t))(1 - 22)

[0292] Wherein, v i (t) and x i (t) respectively represent the speed and position of particle i at the t-th optimization. w is the inertia weight, c1 and c2 are learning factors, rand1 and rand2 generate random numbers in the range of [0, 1], and pbest i is the individual best position of particle i. In one embodiment, in order to obtain the individual best position pbest iDuring the movement of particle i, the fitness score of the test case corresponding to the new position obtained after each movement is calculated, and the test case individual with the highest fitness score is taken as the individual best position pbest i τ gbest It is the global best position selected according to the pheromone strength. In one embodiment, based on the global best reinforcement strategy, after each pheromone update, the global best position with the largest pheromone strength in the current optimization process can be obtained and recorded. Then the global best position in all optimization processes is queried, that is, the historical best position can be obtained, that is, the global best position with the largest pheromone strength recorded in all iterations is used.

[0293] Then, based on formula (1-23), the new position x(t+1) of the particle is calculated.

[0294] x(t+1)=v(t+1)+x(t) (1-23)

[0295] That is, the current moving speed v(t+1) is added to the current position x(t) of the particle to obtain the new position x(t+1) of the particle, and the new position x(t+1) corresponds to a new test case individual.

[0296] It can be seen from formula 1-22 that when calculating the new velocity of the particle, a part of the current velocity is retained through the inertia term of the first term to maintain the historical inertia of the search direction, and the particle itself is adjusted to its own historical optimal position through the second term. The historical optimal position of the group is adjusted to the historical optimal position of the group through the third term. The historical optimal position of the group in this embodiment is the position with the strongest pheromone concentration in the ant colony algorithm.

[0297] As can be seen from the foregoing, the position of the particle is closely related to the particle velocity. In order to meet test cases with different focuses, in one embodiment, the learning factors c1, c2 and / or inertia weights are as shown in the following formulas (1-24), (1-25) and (1-26), including one or more influencing factors and weights.

[0298]

[0299] Among them, g k and They are the k-th influence factor and its weight respectively, and m is the total number of influence factors. The influence factors mentioned above are, for example, one or more of a security risk factor, an environmental impact factor, a user behavior matching rate factor, a business scenario factor, a coverage rate influence factor, an execution time influence factor, and a system load influence factor. Therefore, when calculating the new velocity of a particle, if necessary, first calculate the learning factors c1, c2 and / or the inertia weight required in formula (1-22). For example, when the learning factor includes the weight of the security risk factor, query whether the software unit identifier corresponding to the current position in the software structure diagram includes a security risk identifier. If so, determine the value of the security risk factor according to the preset weight calculation method for defect priorities based on the type and quantity of the security risk identifier. In another embodiment, query the fitness calculation formula of the test case corresponding to the current position, and use the score calculated from the security test standard therein as the value of the security risk factor. For example, use the value calculated from the second term in formula (1-8) as the value of the security risk factor. By incorporating the security test standard into the global search strategy of the particle swarm algorithm, the purpose of preferentially searching for modules with higher risks can be achieved. The value calculated as the value of the security risk factor.

[0300] Similarly, when the environmental impact factor is added to the learning factors c1, c2 and / or the inertia weight, when calculating the particle velocity, the environmental impact value calculated according to formula (1-1) based on the aforementioned fitness function E(C, V) can be used as the environmental impact factor. When the user behavior matching degree is added to the learning factors c1, c2 and / or the inertia weight, the value calculated from the third term in formula (1-8) can be used as the user behavior matching degree factor. By incorporating environmental factors and the user behavior model into the calculation process of the particle position, test cases that are more in line with the actual usage scenario can be generated. In addition, in the ant colony algorithm, for some scenarios corresponding to test cases, the corresponding pheromone will be increased. For example, when the software is mainly used in scenarios with frequent user interactions (e-commerce platforms), for some paths in these scenarios, such as paths related to user input and order processing, the pheromone value is relatively high. Therefore, in order to be able to preferentially search these critical paths when passing through the particle swarm optimization algorithm, a pheromone factor can be added to the learning factors c1, c2 and / or the inertia weight, that is, the pheromone of the test case corresponding to this position is used as one of the influence factors. When the pheromone value is relatively high, it can move preferentially towards this path. The value calculated as the user behavior matching degree factor. By incorporating environmental factors and the user behavior model into the calculation process of the particle position, test cases that are more in line with the actual usage scenario can be generated. In addition, in the ant colony algorithm, for some scenarios corresponding to test cases, the corresponding pheromone will be increased. For example, when the software is mainly used in scenarios with frequent user interactions (e-commerce platforms), for some paths in these scenarios, such as paths related to user input and order processing, the pheromone value is relatively high. Therefore, in order to be able to preferentially search these critical paths when passing through the particle swarm optimization algorithm, a pheromone factor can be added to the learning factors c1, c2 and / or the inertia weight, that is, the pheromone of the test case corresponding to this position is used as one of the influence factors. When the pheromone value is relatively high, it can move preferentially towards this path.

[0301] In another embodiment, when evaluating pbest and gbest, one or more of a security risk factor, an environmental impact factor, a user behavior matching degree factor, a business scenario factor, a coverage rate influence factor, an execution time influence factor, and a system load influence factor can also be added to the evaluation formula, and test cases for multiple scenarios can also be generated accurately and efficiently.

[0302] As can be seen from the foregoing method, according to the update rule of path pheromone in the ant colony algorithm, during the multiple particle swarm iteration optimization processes in steps S303 to S305, it is possible to optimize in the direction of low path coverage rate, high defect risk level, many security vulnerabilities, and better matching with the real user usage scenario, thereby achieving the purpose of guiding particle optimization by the ant colony algorithm.

[0303] This embodiment makes full use of the path optimization ability of the ant colony algorithm (ACO) and the global search ability of the particle swarm optimization algorithm (PSO), and can realize the rapid exploration of software test paths and the optimized generation of test cases in complex path coverage scenarios, and improve the probability of the generated test cases discovering potential defects.

[0304] See Figure 9 , Figure 9 is a method flow chart for generating test cases by a method that combines the genetic algorithm (GA) and the particle swarm optimization (PSO) algorithm according to an embodiment of the present invention.

[0305] Step S401, construct an initial test case set based on the input parameter space, user behavior model, and test environment of the target software, wherein the initial test case set includes a preset number of initial test case individuals, and each initial test case individual includes multiple types of composition parameters, and the types of composition parameters at least include input operations, corresponding input parameters, and operating environment parameters.

[0306] Step S402, reproduce the test case individuals in the current test case set based on the genetic algorithm to obtain new test case individuals. Among them, the test case individual is used as a chromosome, and the composition parameters of the test case are used as genes. Among them, in the first optimization, the current test case set is the initial test case set, and the test case individuals with high fitness scores in the initial test case set are selected as the parent generation for crossover or mutation to reproduce new test case individuals (or called offspring), for example, the crossover operation is performed through formula (1-11), and the mutation operation is performed through formula (1-12). Then, the offspring with low fitness scores are eliminated to obtain a new test case set.

[0307] Step S403, map the new test case set to an initial particle swarm.

[0308] Step S404, move the particles to obtain new particle positions. Among them, the new test case individuals obtained by reproduction are used as the current positions of the particles, and the new moving speed of the particles is calculated based on the current position of the particles, the current speed, the individual best position of the particles, and the global best position, as shown in the following formula (1-27):

[0309] v i(t + 1) = w·v i (t) + c1·rand1·(pbest i - x″ i (t)) + c2·rand2·(gbest - x″ i (t))(1 - 27)

[0310] where, v i (t + 1) is the new velocity, x″ i (t) is the position corresponding to the test case individual obtained through the crossover and / or mutation operation of the gene. pbest i is the best position of this particle individual, and gbest is the global best position.

[0311] Then, the new velocity is added to the current position of the particle to obtain the new position of the particle. For example, the new position of the particle is obtained by using the following formula (1 - 28).

[0312] x i (t + 1) = x″ i (t) + v i (t + 1) (1 - 28)

[0313] The new position of the particle corresponds to an optimized test case individual.

[0314] Step S405, eliminate the test case individuals whose fitness scores do not meet the requirements. Specifically, after obtaining the new test case individuals, calculate their fitness scores, and eliminate the test case individuals whose fitness scores are lower than the threshold, so as to obtain a new optimized test case set.

[0315] Step S406, evaluate whether the current optimized new test case set meets the end condition. When the current optimized test case set meets the end condition, stop the iterative optimization; when the current optimized test case set does not meet the optimization end condition, in step S407, use the optimized test case set as the current test case set for the new round of optimization and return to step S402.

[0316] In this embodiment, a combination of genetic algorithm (GA) and particle swarm optimization algorithm (PSO) is adopted. By combining the social learning behavior of the particle swarm algorithm with the crossover and mutation mechanisms of the genetic algorithm, and using the velocity and position update mechanism of the particle to guide the genetic algorithm operation, it can quickly explore the high-dimensional input parameter space, which can not only improve the population diversity but also accelerate the optimization convergence, effectively solving the problems of difficult generation of appropriate test cases and slow optimization convergence caused by input combination explosion.

[0317] See Figure 10 , Figure 10It is a method flow chart for generating test cases by a method that combines the ant colony algorithm (ACO) and the genetic algorithm (GA) according to an embodiment of the present invention. In this embodiment, the method includes the following steps:

[0318] Step S501, drive multiple ants to perform path exploration in different regions of the path exploration space to generate an initial test case set. Specifically, after each ant determines a path branch during the path exploration process, it selects input operations and input parameters from the input parameter space according to the user behavior model; after each ant finishes the path exploration, obtain the input operations and input parameters of all path branches corresponding to the explored path and configure the running environment parameters to form an initial test case individual; all initial test case individuals form an initial test case set; update the path branch pheromone information during the path exploration process. Among them, the paths corresponding to high-risk security test scenarios in the path exploration space are marked with safety risk attribute identifiers, and when setting the initial pheromone value, increase the initial pheromone value of the path branches marked with safety risk attribute identifiers, so as to explore towards the safety risk area when generating the initial test case individuals.

[0319] Step S502, select parent test case individuals from the current test case set. Among them, the current test case set at the first optimization is the initial test case set. Specifically, first calculate the fitness score f(x) of each test case individual in the current test case set, and then calculate the selection probability of each test case individual respectively. The calculation formula is as shown in the following formula (1-29):

[0320]

[0321] Among them, P select (x) is the selection probability of a test case individual x calculated based on the pheromone and the fitness score. τ(x) represents the pheromone concentration associated with the test case individual x, for example, the sum of all pheromone values corresponding to the path of the test case individual x, and f(x) is the fitness score of the test case individual x.

[0322] Then, take the test case individuals with a selection probability greater than the threshold as the parent test case individuals, or sort according to the selection probability, and take the multiple test case individuals ranked at the front as the parent test case individuals.

[0323] After determining the parent test case individuals, update the path pheromone value according to formula (1-30).

[0324]

[0325] Among them, τ newRepresents the updated pheromone value in a path branch. Δτ corresponds to the pheromone increment value of a parent test case individual covering the current path branch.

[0326] Step S503, perform a reproduction operation on the parent test case individuals based on the genetic algorithm. The reproduction operation is, for example, a crossover operation or a mutation operation. For example, in the foregoing embodiment, each test case individual is regarded as a chromosome, and each constituent parameter of the test case individual (such as the input operation and its corresponding input parameters, and each running environment configuration parameter) is respectively regarded as a chromosome gene. By gene crossover and mutation operations, one or more constituent parameters in a test case individual are changed to obtain a new test case individual. The crossover operation and the mutation operation are as shown in formulas (1-11) and (1-12), which will not be elaborated here. Among the constituent parameters of the test case individual, each input operation and its input parameter value correspond to a path branch, and a path branch corresponds to a specific pheromone value. Therefore, the pheromone value corresponding to the input operation selected for the crossover operation or the mutation operation can be obtained through the global pheromone matrix. In this embodiment, after the crossover operation or the mutation operation, the pheromone value is updated according to formula (1-30). At this time, Δτ is the pheromone increment value corresponding to the path branch corresponding to the input operation selected for the crossover operation or the mutation operation. Through the update of the pheromone, it is beneficial to generate test case individuals in the direction of covering the paths in high-risk areas during the next optimization.

[0327] Step S504, eliminate the test case individuals whose fitness scores do not meet the requirements. The remaining test case individuals form a new test case set.

[0328] Step S505, determine whether the optimized new test case set meets the optimization end condition. If it meets, end the optimization process. If it does not meet, return to step S502.

[0329] In this embodiment, the pheromone of the ant colony guides the parent selection process of the genetic algorithm, and the crossover and mutation operations of the genetic algorithm are introduced into the path update of the ant colony algorithm, which can effectively explore and utilize the input parameter space of the target software to quickly discover high-risk areas and generate test cases that can efficiently cover these areas.

[0330] In summary, by combining the ant colony algorithm, the particle swarm optimization algorithm, and the genetic algorithm, this embodiment can intelligently select or generate test cases, cover paths that are difficult to reach in the software system, generate test cases for complex interactions, significantly improve the test coverage and depth, and ensure that all aspects of the software system can be fully tested. The method proposed in this embodiment can automatically adjust the test cases according to changes in the software system, update the test case set without manual intervention, thereby maintaining the effectiveness and relevance of the test cases, reducing the demand for human resources, and lowering the time and cost of software testing.

[0331] On the other hand, the present invention also provides a software test case generation system. Refer to Figure 11 , Figure 11 which is a schematic block diagram of a software test case generation system according to an embodiment of the present invention. The system includes an initialization module 11, a parameter construction module 12, an initial set construction module 13, and an optimization module 14. Among them, the initialization module 11 is configured to initialize the test environment to obtain operation environment simulation configuration parameters and target software structure parameters; the parameter construction module 12 is configured to construct a user behavior model and an input parameter space of the target software based on the collected user historical behavior data; the initial set construction module 13 is configured to construct an initial test case set based on the input parameter space of the target software, the user behavior model, and the test environment, where the initial test case set includes a preset number of initial test case individuals; the optimization module 14 is configured to iteratively optimize the initial test case individuals in the initial test case set until the optimized test case set meets the optimization end condition, and store the optimized test case set.

[0332] Among them, the operation environment simulation configuration parameters obtained by the initialization module 11 after initializing the test environment include three types of parameters: the operating system version (abbreviated as O), the network condition (abbreviated as N), and the hardware configuration data (abbreviated as H). The target software structure parameters include a dynamic call graph matrix and a call flow graph obtained based on the dynamic call graph matrix. Based on the call flow graph, paths composed of software units of the target software can be obtained. Software units are, for example, components, modules, functions, etc.

[0333] There are various user behavior models constructed by the parameter construction module 12, such as the distribution probability of each user parameter, the usage frequency of user parameter values within a statistical time period, the user parameter value sequence, etc. The categories of user operation parameters include user operation parameters, input parameters corresponding to user operations, user identity parameters, or environment parameters. The input parameter space of the target software includes multiple input variables, and the variable values of each input variable include one or more user parameter values.

[0334] When constructing the initial test case set, the initial set construction module 13 configures the running environment parameters based on the running environment simulation configuration parameters obtained from the initialized test environment; determines the initial test objectives, such as a certain path, based on the target software structure parameters and user behavior model obtained from the initialized test environment; determines the input variables that can achieve the initial test objectives from the input parameter space; selects the input variable values that can simulate user behavior for each input variable from the input parameter space according to the user behavior model; the running environment parameters and one or more input variable values corresponding to an initial objective constitute an initial test case individual. In a specific embodiment, the ant colony algorithm can be used to drive multiple ants to explore paths in different regions of the path exploration space to obtain an initial path. During the exploration process, the input operations and input parameter values corresponding to all path branches in the initial path constitute a test case individual.

[0335] During the iterative optimization process of the initial test case individuals in the initial test case set, the optimization module 14 selects high-quality test case individuals for reproduction to generate new test case individuals, then eliminates the test case individuals with fitness scores not meeting the requirements from all the current test case individuals, and then evaluates whether the optimization end condition is met. When the fitness score of the test case set reaches the threshold or the number of iterations reaches the threshold, it is determined that the optimization end condition is met. At this time, the optimization stops and the current test case set is output and saved. The present invention calculates the fitness score of the test case set based on formula (1-15), thereby evaluating the test case set from three aspects: the coverage criterion, the security criterion, and the user behavior matching criterion, so that the generated test case set can meet the coverage rate of relevant software metrics (function modules, functions, lines of code, etc.) and the depth and complexity of the paths; has sufficient security vulnerability revelation ability; since the test case individuals have a high degree of matching with user behavior and can cover a sufficient number of user behavior patterns, the generated test cases can comprehensively reflect the real usage situation of the software; the test case individuals in the test case set generated by the present invention have different running environment parameters, so as to be able to adapt to different software running environments and provide diverse test scenarios for testing.

[0336] In another embodiment, refer to Figure 12 , Figure 12 is the principle block diagram of the software test case generation system according to another embodiment of the present invention. The system shown in this embodiment includes in addition to Figure 11In addition to the initialization module 11, parameter construction module 12, initial set construction module 13, and optimization module 14, it further includes a scenario classification module 15. The scenario classification module 15 is respectively connected to the initialization module 11, parameter construction module 12, initial set construction module 13, and optimization module 14, and is used to obtain scenario classification features based on the standard software structure parameters and the input parameter space of the standard software, determine the category of the current scenario based on the preset scenarios and their classification conditions, and then send the current scenario category to the initial set construction module 13 and the optimization module 14. The initial set construction module 13 and the optimization module 14 generate a test case set using corresponding algorithms based on the corresponding scenario categories. For specific details, refer to Figures 7 to 10 the process shown in and the corresponding descriptions above, which will not be elaborated here.

[0337] Figure 13 is a schematic diagram of the architecture of a software testing system according to an embodiment of the present invention. The software testing system includes an operation end and a service end. Among them, the operation end is located in the terminal device 102, and the service end is located in the server 104 or a server cluster. The terminal device 102 communicates with the server 104 through a network. The terminal device 102 includes a desktop computer, a laptop computer, or a mobile intelligent terminal device, such as a mobile phone, a tablet computer, etc. The service end includes the Figure 11 or Figure 12 software test case generation system shown in. The operation end includes an interactive interface. When it is necessary to generate test cases, the tester configures corresponding parameters through the interactive interface and sends them to the service end through the network. The service end generates a test case set according to the received instructions and the corresponding parameters according to the method described above in the present invention. The parameters configured by the tester through the interactive interface are, for example, the target software name, version, user historical data storage address, number of iterations, storage address of the test case set, etc. After the test case set is generated, the tester can send a test instruction to the service end through the interactive interface. After receiving the test instruction, the service end executes software testing using the test case set and records the test results.

[0338] Figure 14 is a schematic diagram of the hardware structure principle of an electronic device according to an embodiment of the present invention. The electronic device can be implemented as a server or other various terminal devices, such as a desktop personal computer, a tablet computer, a laptop computer, a mobile phone, etc., which includes a processor 601 and a memory 602. A program instruction set is stored on the memory 602, and when the processor 601 executes the program instruction set on the memory 602, the foregoing software test case generation method is implemented.

[0339] Specifically, the above-mentioned processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present invention.

[0340] The memory 602 may include a mass storage for data or instructions. By way of example and not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 602 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 602 may be internal or external to the integrated gateway disaster recovery device. In a specific embodiment, the memory 602 is a non-volatile solid state memory.

[0341] The memory may include a read only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to execute the software test case generation method provided by the present invention.

[0342] In one example, the electronic device may further include a communication interface 603 and a bus 604. The processor 601, the memory 602, and the communication interface 603 are connected via the bus 604 and complete communication with each other.

[0343] The communication interface 603 is mainly used to implement communication between various modules, devices, units, and / or devices in the embodiments of the present invention.

[0344] The bus 604 includes hardware, software, or both, and couples the components of the online data flow metering device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, the bus 604 may include one or more buses. Although embodiments of the present invention describe and illustrate specific buses, the present invention contemplates any suitable bus or interconnect.

[0345] The present invention also provides a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, any one of the software test case generation methods in the foregoing embodiments can be implemented. The computer-readable storage medium may be tangible and any medium that contains or stores computer-executable instructions for use by or in connection with an instruction execution system, apparatus, and device. The storage medium may be a transient computer-readable storage medium or a non-transient computer-readable storage medium. Non-transient computer-readable storage media may include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Corresponding embodiments of such storage devices include, for example, magnetic disks, optical discs based on CD, DVD, or Blu-ray technology, and persistent solid-state memories such as flash memory, solid-state drives, and the like.

[0346] The present invention also provides a computer program product, which includes a set of computer program instructions, and when the set of computer program instructions is executed by a processor, any one of the software test case generation methods in the foregoing embodiments can be implemented. The computer program product includes, but is not limited to, application installation packages, application plugins published on websites, application stores, and applets that can run in certain applications, and the like.

[0347] It should be clear that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present invention is not limited to the specific steps described and illustrated, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0348] The above embodiments are only for illustrating the present invention and are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can still make various changes and modifications without departing from the scope of the present invention. Therefore, all equivalent technical solutions should also fall within the scope of the disclosure of the present invention.

Claims

1. A software test case generation method, characterized in that Including: Initializing a test environment to obtain operating environment simulation configuration parameters and target software structure parameters; Constructing a user behavior model and an input parameter space of the target software based on the collected user historical behavior data; Constructing an initial test case set based on the input parameter space of the target software, the user behavior model, and the test environment, wherein the initial test case set includes a preset number of initial test case individuals; And Iteratively optimizing the initial test case individuals in the initial test case set until the optimized test case set meets the optimization end condition, and storing the optimized test case set.

2. The software test case generation method according to claim 1, wherein The target software structure parameters include a dynamic call graph matrix of the target software; After obtaining the dynamic call graph matrix by initializing the test environment, it further includes: Parsing the dynamic call graph matrix to obtain two software units corresponding to each matrix element; Constructing a call flow graph with software units as nodes and call relationships between software units as edges, wherein two nodes connected by an edge form a path branch; and Parsing the call relationships between software units to obtain path branch conditions.

3. The software test case generation method according to claim 2, wherein After obtaining the call flow graph by initializing the test environment, it further includes: Obtaining security risk attribute marks of software units.

4. The software test case generation method according to claim 1 or 2 or 3, characterized in that, The steps of constructing a user behavior model and an input parameter space of the target software based on the collected user historical behavior data include: Extracting multiple types of user parameter values from the user historical behavior data, and the categories of user parameters include user operation parameters, input parameters corresponding to user operations, user identity parameters, or environment parameters; Generating a user behavior model based on the extracted user parameter values; and Constructing an input parameter space of the target software, wherein the input parameter space includes multiple input variables, and the variable value of each input variable includes one or more user parameter values.

5. The software test case generation method according to claim 4, wherein The steps of generating a user behavior model based on the extracted user parameter values include: Calculating the distribution probability of each user parameter value based on the extracted user parameter values and all user parameter values of the same user; Alternatively, counting the usage frequency of each user parameter value of the same user within a statistical time period; Alternatively, counting multiple user parameter value sequences representing different user behavior patterns.

6. The software test case generation method according to claim 4, wherein The steps of constructing an initial test case set based on the input parameter space of the target software, the user behavior model, and the test environment include: Configuring operating environment parameters based on the operating environment simulation configuration parameters obtained by initializing the test environment; Determining an initial test target based on the target software structure parameters and the user behavior model obtained by initializing the test environment; Determining input variables in the input parameter space that can achieve the initial test target; and Selecting input variable values that can simulate user behavior for each input variable from the input parameter space according to the user behavior model; Wherein, the operating environment parameters and one or more input variable values corresponding to an initial target constitute an initial test case individual; The initial test target is a path branch in the call flow graph or a path composed of multiple path branches; 7. The software test case generation method according to claim 6, wherein When determining the initial test target, select a path branch where the software unit corresponding to the node has a security risk attribute mark; Correspondingly, after determining the input variables in the input parameter space that can achieve the initial test objective, it includes: determining the input variable values from the input parameter space based on the security test trigger condition.

8. The software test case generation method according to claim 6, characterized in that The step of constructing the initial test case set based on the input parameter space of the target software, the user behavior model, and the test environment further includes: Constructing a maximum problem function for the initial test case set based on the environmental fitness and the user behavior model; Calculating the environmental fitness score of each initial test case individual based on the operating environment parameters of the initial test case individual; Calculating the user behavior simulation score based on the user behavior model for generating the initial test case individual; Adjusting the operating environment parameters of one or more initial test case individuals in the initial test case set and / or the user behavior model for generating the initial test case individual to make the maximum problem function converge to the maximum value; Determining the test case set when converging to the maximum value as the initial test case set.

9. The software test case generation method according to claim 6, characterized in that During the process of iteratively optimizing the initial test case individuals in the initial test case set, the step of performing one optimization on the current test case set includes: Selecting some test case individuals from the current test case set as seed test case individuals, and multiple seed test case individuals form a seed set; and Reproducing the seed test case individuals in the seed set to generate new test case individuals, and the seed test case individuals and the new test case individuals form the test case set for the next optimization.

10. The software test case generation method according to claim 9, characterized in that The step of selecting some test case individuals from the current test case set as seed test case individuals includes: Obtaining the fitness score of each test case individual in the current test case set; Traversing the fitness scores of each test case individual, and taking the test case individuals with fitness scores greater than the score threshold as seed test case individuals or taking the test case individuals ranked before the sorting threshold as seed test case individuals.

11. The software test case generation method according to claim 10, wherein When the current test case set is the initial test case set, the step of obtaining the fitness score of each test case individual in the current test case set includes: Calculating the coverage rate of each initial test case individual for the target software metrics; Evaluating each initial test case individual based on the threat model for security testing to obtain the defect revelation ability value of each initial test case individual; and Calculating the weighted sum of the coverage rate of each initial test case individual for the target software metrics and the defect revelation ability value as the initial fitness score of the initial test case individual; When the current test case set is the test case set including the optimized test case set, the step of obtaining the fitness score of each test case individual in the current test case set includes: Running each test case individual in the current test case set and collecting the corresponding running data; Calculating the coverage rate of each test case individual for the target software metrics based on the running data of each test case individual; Evaluating each test case individual based on the threat model for security testing to obtain the defect revelation ability value of each test case individual; Querying the user behavior model data and calculating the matching rate between the user behavior simulated by each test case individual and the user behavior model; and Calculate the weighted sum of the coverage rate of each test case individual for the target software metric, the defect revelation ability value, and the matching rate between the user behavior simulated by the test case individual and the user behavior model as the fitness score of each test case individual.

12. The software test case generation method according to claim 11, wherein The end condition described above is reaching the iteration number threshold or the fitness score of the current test case set reaching the threshold.

13. The software test case generation method according to claim 9, wherein When breeding the seed test case individuals in the seed set to generate new test case individuals, it includes one or more of the following steps: Changing one or more path branches in the path corresponding to the seed test case to generate a new test case individual; Changing the running environment parameters of the seed test case to generate a new test case individual; Changing the input operations or input operation sequences and / or corresponding input parameters used to simulate user behavior in the seed test case.

14. A software test case generation system, characterized in that, It includes: An initialization module, configured to initialize the test environment to obtain the running environment simulation configuration parameters and the target software structure parameters; A parameter construction module, configured to construct a user behavior model and the input parameter space of the target software based on the collected user historical behavior data; An initial set construction module, configured to construct an initial test case set based on the input parameter space of the target software, the user behavior model, and the test environment, where the initial test case set includes a preset number of initial test case individuals; And An optimization module, configured to iteratively optimize the initial test case individuals in the initial test case set until the optimized test case set meets the optimization end condition, and store the optimized test case set.

15. An electronic device, characterized in that, The electronic device includes a processor and a memory storing computer program instructions; when the electronic device executes the computer program instructions, it implements the software test case generation method described in any one of claims 1-13.

16. A computer-readable storage medium, on which computer program instructions are stored, characterized in that, When the computer program instructions are executed by the processor, it implements the software test case generation method described in any one of claims 1-13.

17. A computer program product, characterized in that, It includes a set of computer program instructions, and when the set of computer program instructions is executed by the processor, it implements the software test case generation method described in any one of claims 1-13.

Citation Information

Cited By

  • A method and system for generating label-based automated use cases

    CN122547695A