Low-code generation method and system for security vulnerability test case

By introducing similar code retrieval and effectiveness evaluation into the security vulnerability test case generation system, the problem of test case invalidity in the existing technology is solved, and more efficient test case screening and sending is achieved, which improves testing efficiency and resource utilization.

CN120029916AActive Publication Date: 2025-05-23河源市尖锋信息科技有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510120602.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-05-23
Estimated Expiration
2045-01-25

AI Technical Summary

Technical Problem

The prior art cannot effectively verify the feasibility of generated security vulnerability test cases, resulting in a large number of invalid use cases, wasting test resources and reducing testing efficiency.

Method used

By obtaining vulnerability types, function descriptions, input constraints and expected results, input test cases to generate preliminary test cases in the network, and search similar codes based on the similarity threshold, count the proportion of valid codes of the similar use case set, set as the effective probability of the test case, and send the test case to the test terminal when the effective probability is greater than or equal to the preset threshold.

Benefits of technology

Improve the availability of test cases, filter out truly valuable test cases through effectiveness evaluation criteria, reduce the generation of invalid cases, and improve test efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029916A_ABST
    Figure CN120029916A_ABST
Patent Text Reader

Abstract

The invention relates to a security vulnerability test case low-code generation method and system, and relates to the technical field of security vulnerability testing, and the method comprises the steps: obtaining vulnerability types, function description, input constraints and expected results, inputting a test case generation network, obtaining test cases, and retrieving a similar case set based on a similarity threshold; counting an effective code proportion of the similar case set, and setting the effective code proportion as a test case effective probability; and when the effective probability of the test case is greater than or equal to an effective probability threshold, sending the test case to a test terminal. The technical problem that in the prior art, feasibility verification of the security vulnerability test case cannot be achieved, and consequently a large number of invalid cases exist in the generated test case is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of security vulnerability testing technology, and in particular to a low-code generation method and system for security vulnerability test cases. Background Art

[0002] In the research on automated generation of security vulnerability test cases, although AI models can quickly generate a large number of test cases, the feasibility of these cases cannot be verified in advance, resulting in a large number of invalid cases in the generated test cases, which not only wastes test resources but also reduces test efficiency, requiring additional manual screening to remove redundant cases in actual applications. Summary of the invention

[0003] The present invention aims to solve the technical problem that the feasibility verification of security vulnerability test cases cannot be implemented in the prior art, resulting in a large number of invalid cases in the generated test cases. A low-code generation method and system for security vulnerability test cases are provided to solve the problem.

[0004] The technical solution of the present invention to solve the above technical problems is as follows:

[0005] In a first aspect, the present invention provides a low-code generation method for security vulnerability test cases, including: obtaining vulnerability type, function description, input constraints and expected results, inputting a test case generation network, and obtaining a test case; for the test case, performing similar code retrieval based on a similarity threshold to obtain a similar case set; counting the proportion of valid code in the similar case set, and setting it as the test case valid probability; when the test case valid probability is greater than or equal to the valid probability threshold, sending the test case to a test terminal.

[0006] In the second aspect, the present invention provides a low-code generation system for security vulnerability test cases, including: a test case generation unit, used to obtain vulnerability type, function description, input constraints and expected results, input a test case generation network, and obtain a test case; a similar code retrieval unit, used to perform similar code retrieval on the test case based on a similarity threshold, and obtain a similar case set; an effective code analysis unit, used to count the effective code ratio of the similar case set, and set it as the effective probability of the test case; a test case feedback unit, used to send the test case to the test terminal when the effective probability of the test case is greater than or equal to the effective probability threshold.

[0007] The beneficial effects of the present invention are: by providing a method of inputting vulnerability types, function descriptions, input constraints and expected results into a test case generation network to generate preliminary test cases; then performing similar code retrieval on the generated test cases based on a similarity threshold to obtain a similar case set; then counting the proportion of valid codes in the similar case set, defining it as the effective probability of the test case; and finally, when the effective probability of the test case is greater than or equal to a preset effective probability threshold, sending the test case to a test terminal. By performing similar code retrieval on the test cases generated automatically to obtain a similar case set, and then counting the proportion of valid codes in the similar case set, setting it as the effective probability of the test case, and using it as an evaluation standard for sorting test cases, the technical effect of improving the availability of test cases is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 A schematic diagram of a process for low-code generation of security vulnerability test cases provided by the present invention;

[0009] Figure 2 A structural diagram of a low-code generation system for security vulnerability test cases provided by the present invention.

[0010] In the accompanying drawings, the components represented by the reference numerals are described as follows:

[0011] Test case generating unit 100, similar code retrieving unit 200, effective code analyzing unit 300, test case feedback unit 400. DETAILED DESCRIPTION

[0012] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0013] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0014] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or explanation". Any embodiment described as "for example" in the present invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any technician in the field to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in the present invention.

[0015] Embodiment 1:

[0016] like Figure 1 As shown, an embodiment of the present invention provides a low-code generation method for security vulnerability test cases, comprising the steps of:

[0017] S10: Obtain vulnerability type, function description, input constraints and expected results, input test case generation network, and obtain test cases;

[0018] Specifically, vulnerability type refers to the type of security defects existing in the software system, such as SQL injection, cross-site scripting (XSS) or buffer overflow; functional description: a detailed description of the software function point, describing the expected behavior and operation logic of the function point; input constraints: restrictions on test case input data, such as data type, range or format, etc., used to ensure that the input data meets the test requirements; expected results: the expected results after the test case is executed, used to verify whether the software function is normal; test case generation network: a model based on artificial intelligence or machine learning, used to generate test cases based on the input vulnerability type, functional description, input constraints and expected results; test case: a specific test scenario used to verify whether the software function is normal or whether there is a vulnerability, including input data and expected output. The vulnerability type, functional description, input constraints and expected results can be user-defined input content, which is used to constrain the generation of test cases.

[0019] Preferably, the test case generation network construction steps are as follows:

[0020] A generative adversarial network is constructed, including a generator and an adversary. First, a test case data set is collected and mixed with a false test case data set to obtain an adversary training data set. The supervisory data value of the test case data is 1, indicating that it is true, and the supervisory data value of the false test case data is 0, indicating that it is false. The supervisory data is used to call the adversary training data set as input to train the adversary, which can achieve the distinction between false test cases and real test cases. Then, a vulnerability type record data set, a function description record data set, an input constraint record data set, an expected result record data set, and a test case identification data set are collected. The test case identification data set is used as supervision, and the vulnerability type record data set, the function description record data set, the input constraint record data set, and the expected result record data set are used as input to train the generator. Finally, the output of the generator is used as the input of the adversary. Whenever the adversary output is 1, it means that the test case generated by the generator can be mistaken for the real test case, and it can be added to the test case.

[0021] By obtaining vulnerability types and related description information, the generated test cases can be ensured to be targeted, providing high-quality input for subsequent similar code retrieval and effectiveness evaluation, thereby improving the efficiency and accuracy of the entire test case generation system.

[0022] S20: performing similar code retrieval on the test case based on a similarity threshold to obtain a similar case set;

[0023] Further, for the test case, similar code retrieval is performed based on a similarity threshold to obtain a similar case set, including:

[0024] Obtaining the benchmark code of the test case to perform AST tree configuration and obtain a benchmark abstract syntax tree;

[0025] Obtain the comparison code of the use case to be analyzed, configure the AST tree, and obtain the comparison abstract syntax tree;

[0026] By using the Tree-LSTM model, the reference abstract syntax tree and the comparison abstract syntax tree are respectively encoded to obtain a reference encoding vector and a comparison encoding vector;

[0027] According to a similarity evaluation function, performing a weighted mean evaluation on the structural similarity between the reference abstract syntax tree and the comparison abstract syntax tree and the cosine similarity between the reference coding vector and the comparison coding vector to obtain a similarity evaluation value;

[0028] When the similarity evaluation value is greater than or equal to the similarity threshold, the use case to be analyzed is added to the similar use case set.

[0029] Specifically, similarity threshold: a preset value used to measure the similarity between codes; two code segments are considered similar when their similarity is greater than or equal to the threshold. For example, setting the similarity threshold to 0.8 means that two code segments are considered similar only when their similarity is more than 80%; similar code retrieval: compare the similarity between the test case code and other known codes to find code snippets or use cases that are similar to the target test case; similar use case set: a set of all use cases that are similar to the target test case, which have a high similarity with the target use case in terms of code structure or function.

[0030] The specific process is: first, configure the benchmark code of the test case into an abstract syntax tree (AST), and then perform the same AST configuration on the codes of other use cases to be analyzed. For example, assuming that the benchmark code is a simple login function code, its AST tree will reflect the structure and logic of the code; then, use the Tree-LSTM model to encode the benchmark AST and the comparison AST to generate an encoding vector; through the similarity evaluation function, combine the structural similarity of the AST and the cosine similarity of the encoding vector to calculate the similarity evaluation value; if the value is greater than or equal to the similarity threshold (for example, 0.8), the use case to be analyzed is added to the set of similar use cases.

[0031] S30: Counting the effective code ratio of the similar test case set and setting it as the effective probability of the test case;

[0032] S40: When the effective probability of the test case is greater than or equal to the effective probability threshold, the test case is sent to the test terminal.

[0033] Furthermore, the effective code ratio of the similar use case set is counted and set as the effective probability of the test case, including:

[0034] Obtaining a first similar use case in the set of similar use cases;

[0035] Construct validity evaluation function:

[0036]

[0037] Among them, w 1 、w 2 、w 3 、w 4 and w 5 Characterize the weight parameter, which is equal to 1. R characterizes the use case repeatability parameter, C characterizes the use case coverage parameter, Cl characterizes the use case clarity parameter, Co characterizes the use case completeness parameter, Ru characterizes the use case reusability parameter, and θ characterizes the threshold parameter, which is used to adjust the offset of the function;

[0038] analyzing, according to the effectiveness evaluation function, a first similar use case effectiveness parameter of the first similar use case;

[0039] When the validity parameter of the first similar use case is greater than or equal to a validity parameter threshold, marking the first similar use case with a valid code;

[0040] The proportion of similar use cases with the valid code identifier in the similar use case set is counted and set as the validity probability of the test case.

[0041] Furthermore, the similarity evaluation function is:

[0042] S(T 1 ,T 2 ) 0 =αS(T 1 ,T 2 ) 1 +(1-α)S(E(T 1 ),E(T 2 )) 2 ,

[0043]

[0044] Among them, S(T 1 ,T 2 ) 0 Represents the comprehensive similarity, S(T 1 ,T 2 ) 1 Characterize the structural similarity between the reference abstract syntax tree and the comparison abstract syntax tree, S(E(T 1 ),E(T 2 )) 2 Characterize the cosine similarity between the reference code vector and the comparison code vector, E(T 1 ) represents the benchmark encoding vector, E(T 2 ) represents the alignment encoding vector, T 1 Representation of the benchmark abstract syntax tree, T 2 Representation comparison abstract syntax tree, V 1 The set of nodes representing the base abstract syntax tree, V 2 represents the node set of the abstract syntax tree for comparison, λ is the regularization parameter, v represents the intersection node, sim(v) represents the intersection node similarity, children(v) represents the intersection child nodes of the intersection node, n[children(v)] 1 Characterize the intersection node in V 1 and V 2 The number of child nodes whose similarity is greater than or equal to the similarity threshold, n[children(v)] 2The total number of intersection child nodes, α represents the preset weight.

[0045] Furthermore, it also includes:

[0046] Count the length of arithmetic operator code snippets at intersection nodes;

[0047] Get the total length of the code snippet of the intersection node;

[0048] The ratio of the length of the arithmetic operator code snippet to the total length of the code snippet is calculated and set as the preset weight.

[0049] Specifically, the effective code ratio refers to the proportion of use cases corresponding to code snippets that can effectively trigger vulnerabilities or meet test requirements in a set of similar use cases. The effective code is the code corresponding to similar use cases that has been verified and can accurately reflect the characteristics of vulnerabilities or test objectives. The effective probability of a test case is a probability value calculated based on the effective code ratio, which is used to evaluate the overall effectiveness of the test case and is an important indicator for measuring whether a test case is worth adopting.

[0050] The specific process is as follows: First, select a use case from the set of similar use cases as the analysis object, for example, select the first similar use case; then, construct an effectiveness evaluation function Analyze the effectiveness parameters of the use case, such as repeatability = number of executions with consistent results / total number of executions; coverage = (number of statements executed / number of executable statements + number of branches tested / total number of branches + number of conditions tested / total number of conditions) / 3; clarity = number of clear test steps and expected results / total number of test steps and expected results; completeness = number of covered function points and boundary conditions / total number of function points and boundary conditions; reusability = number of reusable test cases / total number of test function points, etc.; combine them into a comprehensive effectiveness index through weight allocation; if the index is greater than or equal to the preset effectiveness parameter threshold, mark the use case as valid code; finally, count the proportion of all use cases marked as valid code in the set of similar use cases, and set this proportion value as the effectiveness probability of the test case.

[0051] For example, assuming that there are 100 use cases in a set of similar use cases, after effectiveness evaluation, it is found that the effectiveness parameters of 70 of them are higher than the threshold and are identified as valid use cases. Then, the effective code accounts for 70%, and the effective probability of the test case is also 70%. By counting the effective code proportion, the quality of the test case can be quantified, so as to screen out the truly valuable test cases. For example, if the effective probability of a test case is low (such as less than 50%), it may be judged as redundant or invalid, thereby avoiding sending it to the test terminal and saving test resources. On the contrary, if the effective probability is high (such as more than 80%), it means that the test case has higher feasibility and value and can be adopted preferentially.

[0052] Furthermore, the similar use case screening process is as follows:

[0053] The similarity evaluation function is constructed as: S(T 1 ,T 2 ) 0 =αS(T 1 ,T 2 ) 1 +(1-α)S(E(T 1 ),E(T 2 )) 2 ,

[0054] Then according to the compared V 1 and V 2 ,Dynamically configure the weight α. After the configuration is completed, the use case similarity is analyzed through the similarity evaluation function. The comprehensiveness of the analysis results is guaranteed by analyzing the use case similarity from both structural and semantic aspects. The dynamic configuration weight can be dynamically assigned according to the compared real-time code data, which improves the individual characteristics.

[0055] Further, the vulnerability type, function description, input constraints and expected results are obtained, and step S10 includes the following steps:

[0056] S110: Traversing the test environment function points to count frequent vulnerabilities of the same function points, and obtaining a first frequent vulnerability type;

[0057] S120: traverse the test environment code to scan and obtain abnormal functions and error types;

[0058] S130: Count frequent vulnerabilities in the same state according to the abnormal function and the error type to obtain a second frequent vulnerability type;

[0059] S140: traverse the test environment function points, take the union of the first frequent vulnerability type and the second frequent vulnerability type, and obtain a comprehensive frequent vulnerability type, wherein the comprehensive frequent vulnerability type corresponds to the function point one by one;

[0060] S150: Setting the comprehensive frequent vulnerability type as the test case, and configuring the function description, the input constraints and the expected results according to the function point.

[0061] Specifically, the test environment function point refers to the module or operation point with independent functions in the software system, such as the user login module, file upload function, etc.; the statistics of frequent vulnerabilities in the same function point: by analyzing historical vulnerability data, find out the types of vulnerabilities that frequently appear in the same type of function points; abnormal functions and error types: functions or code snippets that may cause vulnerabilities detected by code scanning tools, as well as the types of errors that may occur at runtime; statistics of frequent vulnerabilities in the same state: based on abnormal functions and error types, statistics on frequently occurring vulnerability types caused by the same function and the same error type; comprehensive frequent vulnerability types: a set of vulnerability types obtained by merging the vulnerability types obtained through function point statistics and code scanning statistics; functional description, input constraints and expected results: detailed information configured for the generated test cases, used to guide the generation and verification of test cases.

[0062] The specific process is as follows:

[0063] Traversing the test environment function points: The system will check the function points in the software one by one, such as the user login module, data query module, etc. For each function point, the frequently occurring vulnerability types in history are counted to obtain the first frequent vulnerability type. For example, in the user login module, SQL injection vulnerabilities may frequently occur.

[0064] Scan the test environment code: Use static code analysis tools to scan abnormal functions and error types in the code. For example, it is detected that a certain function may cause a buffer overflow, or a certain code snippet may trigger an XSS attack.

[0065] Counting frequent vulnerabilities in the same state: Based on the abnormal function and error type, further count the vulnerability types that frequently appear in the same running state or code structure to obtain the second most frequent vulnerability type.

[0066] Merge vulnerability types: Take the union of the first frequent vulnerability type and the second frequent vulnerability type to get the comprehensive frequent vulnerability type. These vulnerability types correspond to function points one by one to ensure that each function point has a targeted vulnerability type.

[0067] Configure test cases: Use the comprehensive frequent vulnerability type as the target vulnerability type of the test case, and configure the function description, input constraints, and expected results according to the characteristics of the function point. For example, for the SQL injection vulnerability of the user login module, the function description is "user login function", the input constraint is "user name and password field", and the expected result is "return an error prompt when a SQL injection attack is detected."

[0068] By traversing the function points and scanning the code, we can dig out the high-risk vulnerability types in the actual environment and avoid blindly generating test cases. At the same time, we configure the detailed information of the test cases according to the characteristics of the function points to ensure that the generated test cases are targeted and effective.

[0069] Further, the test environment function points are traversed to perform frequent vulnerability statistics of the same function points to obtain a first frequent vulnerability type. Step S110 includes the following steps:

[0070] S111: Obtain a first function point of the test environment function points, wherein the first function point has a function description tag;

[0071] S112: Retrieving historical vulnerability test data having a function description tag identical to the function description tag;

[0072] S113: Counting vulnerability types and trigger frequencies of the historical vulnerability test data, screening the vulnerability types whose trigger frequencies are greater than or equal to a trigger frequency threshold, and adding them to the frequent vulnerability types of the first function point;

[0073] S114: Add the first function point frequent vulnerability type into the first frequent vulnerability type.

[0074] Specifically, the process of counting frequent vulnerabilities in the same state and the same function point is the same, and the same function point frequent vulnerabilities statistics are used as an example:

[0075] Function description label: a label used to identify the characteristics or functions of a test environment function point, such as "user login", "file upload" or "data query", etc., used to quickly locate and classify function points; historical vulnerability test data: vulnerability data recorded in the same or similar function points in the past, including vulnerability type, trigger conditions, occurrence frequency and other information; trigger frequency: the number of times or frequency a vulnerability appears in historical vulnerability test data, used to assess the commonness of the vulnerability; trigger frequency threshold: a preset value used to filter out frequently occurring vulnerability types. For example, if the threshold is set to 10 times, only vulnerability types with a trigger frequency greater than or equal to 10 times will be included in the statistics.

[0076] The specific process is as follows:

[0077] Get function points and their description labels: Extract the first function point (such as "user login module") from the test environment and get its function description label (such as "user login").

[0078] Retrieve historical vulnerability data: According to the function description tag, retrieve the historical vulnerability test data with the same function point. For example, find all vulnerability records related to "user login" from the vulnerability database.

[0079] Statistics of vulnerability types and triggering frequencies: Analyze the retrieved historical vulnerability data and count the occurrence frequency of each vulnerability type. For example, statistics show that SQL injection vulnerabilities appeared 20 times, XSS vulnerabilities appeared 15 times, and other vulnerability types appeared less frequently.

[0080] Filter high-frequency vulnerability types: According to the preset trigger frequency threshold (for example, 10 times), filter out the vulnerability types whose trigger frequency is greater than or equal to the threshold. For example, if the threshold is 10 times, SQL injection and XSS vulnerability types are included in the frequent vulnerability types of the first functional point.

[0081] Add to the first frequent vulnerability type set: Add the filtered high-frequency vulnerability types to the first frequent vulnerability type set for subsequent vulnerability test case generation.

[0082] By analyzing historical vulnerability data, it is possible to accurately identify the high-risk vulnerability types in a certain functional point, rather than relying on a general vulnerability list.

[0083] A low-code generation method for security vulnerability test cases provided by an embodiment of the present invention has at least the following technical effects:

[0084] The invention provides a technical solution for inputting vulnerability type, function description, input constraint and expected result into the test case generation network to generate preliminary test cases; then, based on the similarity threshold, similar code retrieval is performed on the generated test cases to obtain a similar case set; then, the proportion of valid code in the similar case set is counted and defined as the effective probability of the test case; finally, when the effective probability of the test case is greater than or equal to the preset effective probability threshold, the test case is sent to the test terminal. By performing similar code retrieval on the test cases generated automatically, a similar case set is obtained, and then the proportion of valid code in the similar case set is counted and set as the effective probability of the test case, which is used as the evaluation standard for sorting test cases, thereby achieving the technical effect of improving the availability of test cases.

[0085] Embodiment 2:

[0086] like Figure 2 As shown, based on the same inventive concept as the method for generating a security vulnerability test case with low code provided in Embodiment 1, an embodiment of the present invention further provides a system for generating a security vulnerability test case with low code, including:

[0087] The test case generation unit 100 is used to obtain the vulnerability type, function description, input constraints and expected results, input the test case generation network, and obtain the test case;

[0088] A similar code retrieval unit 200 is used to perform similar code retrieval on the test case based on a similarity threshold to obtain a similar case set;

[0089] The effective code analysis unit 300 is used to count the effective code ratio of the similar test case set and set it as the effective probability of the test case;

[0090] The test case feedback unit 400 is used to send the test case to the test terminal when the validity probability of the test case is greater than or equal to the validity probability threshold.

[0091] Furthermore, the test case generation unit 100 performs the following steps:

[0092] Traverse the test environment function points to count the frequent vulnerabilities of the same function points and obtain the first frequent vulnerability type;

[0093] Traverse the test environment code to scan and obtain the abnormal function and error type;

[0094] According to the abnormal function and error type, the frequent vulnerabilities in the same state are counted to obtain the second most frequent vulnerability type;

[0095] Traversing the test environment function points, taking the union of the first frequent vulnerability type and the second frequent vulnerability type, and obtaining a comprehensive frequent vulnerability type, wherein the comprehensive frequent vulnerability type corresponds to the function point one by one;

[0096] The comprehensive frequent vulnerability type is set as the test case, and the function description, the input constraint and the expected result are configured according to the function point.

[0097] Furthermore, the test case generation unit 100 executes the steps further including:

[0098] Obtaining a first function point of the test environment function points, wherein the first function point has a function description tag;

[0099] Retrieving historical vulnerability test data having the same functional description tag as the functional description tag;

[0100] Counting vulnerability types and trigger frequencies of the historical vulnerability test data, screening the vulnerability types whose trigger frequencies are greater than or equal to a trigger frequency threshold, and adding them to the frequent vulnerability types of the first functional point;

[0101] Add the first function point frequent vulnerability type into the first frequent vulnerability type.

[0102] Furthermore, the similar code retrieval unit 200 performs the following steps:

[0103] Obtaining the benchmark code of the test case to perform AST tree configuration and obtain a benchmark abstract syntax tree;

[0104] Obtain the comparison code of the use case to be analyzed, configure the AST tree, and obtain the comparison abstract syntax tree;

[0105] By using the Tree-LSTM model, the reference abstract syntax tree and the comparison abstract syntax tree are respectively encoded to obtain a reference encoding vector and a comparison encoding vector;

[0106] According to a similarity evaluation function, performing a weighted mean evaluation on the structural similarity between the reference abstract syntax tree and the comparison abstract syntax tree and the cosine similarity between the reference coding vector and the comparison coding vector to obtain a similarity evaluation value;

[0107] When the similarity evaluation value is greater than or equal to the similarity threshold, the use case to be analyzed is added to the similar use case set.

[0108] Furthermore, the similarity evaluation function is:

[0109] S(T 1 ,T 2 ) 0 =αS(T 1 ,T 2 ) 1 +(1-α)S(E(T 1 ),E(T 2 )) 2 ,

[0110]

[0111]

[0112] Among them, S(T 1 ,T 2 ) 0 Represents the comprehensive similarity, S(T 1 ,T 2 ) 1 Characterize the structural similarity between the reference abstract syntax tree and the comparison abstract syntax tree, S(E(T 1 ),E(T 2 )) 2 Characterize the cosine similarity between the reference code vector and the comparison code vector, E(T 1 ) represents the benchmark encoding vector, E(T 2 ) represents the alignment encoding vector, T 1 Representation of the benchmark abstract syntax tree, T 2 Representation comparison abstract syntax tree, V 1 The set of nodes representing the base abstract syntax tree, V2 Represents the node set of the comparison abstract syntax tree, λ is the regularization parameter, v represents the intersection node, sim(v) represents the intersection node similarity, children(v) represents the intersection child node of the intersection node, n[children(v)] 1 Characterize the intersection node in V 1 and V 2 The number of child nodes whose similarity is greater than or equal to the similarity threshold, n[children(v)] 2 The total number of intersection child nodes, α represents the preset weight.

[0113] Furthermore, the similar code retrieval unit 200 performs the following steps:

[0114] Count the length of arithmetic operator code snippets at intersection nodes;

[0115] Get the total length of the code snippet of the intersection node;

[0116] The ratio of the length of the arithmetic operator code snippet to the total length of the code snippet is calculated and set as the preset weight.

[0117] Furthermore, the effective code analysis unit 300 performs the following steps:

[0118] Obtaining a first similar use case in the set of similar use cases;

[0119] Construct validity evaluation function:

[0120]

[0121] Among them, w 1 、w 2 、w 3 、w 4 and w 5 Characterize the weight parameter, which is equal to 1. R characterizes the use case repeatability parameter, C characterizes the use case coverage parameter, Cl characterizes the use case clarity parameter, Co characterizes the use case completeness parameter, Ru characterizes the use case reusability parameter, and θ characterizes the threshold parameter, which is used to adjust the offset of the function;

[0122] analyzing, according to the effectiveness evaluation function, a first similar use case effectiveness parameter of the first similar use case;

[0123] When the validity parameter of the first similar use case is greater than or equal to a validity parameter threshold, marking the first similar use case with a valid code;

[0124] The proportion of similar use cases with the valid code identifier in the similar use case set is counted and set as the validity probability of the test case.

[0125] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0126] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0127] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0128] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0129] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0130] Although preferred embodiments of the present invention have been described, additional changes and modifications may occur to these embodiments once those skilled in the art understand the basic inventive concepts.

[0131] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention belong to the scope of the present invention and its equivalent technologies, the present invention is also intended to include these changes and variations.

Claims

1. A low-code generation method for security vulnerability test cases, characterized in that: include: Obtain vulnerability type, function description, input constraints and expected results, input test case generation network, and obtain test cases; For the test case, similar code retrieval is performed based on a similarity threshold to obtain a similar case set; Count the percentage of valid codes in the similar test case set and set it as the test case validity probability; When the effective probability of the test case is greater than or equal to the effective probability threshold, the test case is sent to the test terminal.

2. The method according to claim 1, characterized in that Get vulnerability type, functional description, input constraints, and expected results, including: Traverse the test environment function points to count the frequent vulnerabilities of the same function points and obtain the first frequent vulnerability type; Traverse the test environment code to scan and obtain the abnormal function and error type; According to the abnormal function and error type, the frequent vulnerabilities in the same state are counted to obtain the second most frequent vulnerability type; Traversing the test environment function points, taking the union of the first frequent vulnerability type and the second frequent vulnerability type, and obtaining a comprehensive frequent vulnerability type, wherein the comprehensive frequent vulnerability type corresponds to the function point one by one; The comprehensive frequent vulnerability type is set as the test case, and the function description, the input constraint and the expected result are configured according to the function point.

3. The method according to claim 1, characterized in that Traverse the test environment function points to count the frequent vulnerabilities of the same function points and obtain the first frequent vulnerability type, including: Obtaining a first function point of the test environment function points, wherein the first function point has a function description tag; Retrieving historical vulnerability test data having the same functional description tag as the functional description tag; Counting vulnerability types and trigger frequencies of the historical vulnerability test data, screening the vulnerability types whose trigger frequencies are greater than or equal to a trigger frequency threshold, and adding them to the frequent vulnerability types of the first functional point; Add the first function point frequent vulnerability type into the first frequent vulnerability type.

4. The method according to claim 1, characterized in that For the test case, similar code retrieval is performed based on a similarity threshold to obtain a similar case set, including: Obtaining the benchmark code of the test case to perform AST tree configuration and obtain a benchmark abstract syntax tree; Obtain the comparison code of the use case to be analyzed, configure the AST tree, and obtain the comparison abstract syntax tree; By using the Tree-LSTM model, the reference abstract syntax tree and the comparison abstract syntax tree are respectively encoded to obtain a reference encoding vector and a comparison encoding vector; According to a similarity evaluation function, performing a weighted mean evaluation on the structural similarity between the reference abstract syntax tree and the comparison abstract syntax tree and the cosine similarity between the reference coding vector and the comparison coding vector to obtain a similarity evaluation value; When the similarity evaluation value is greater than or equal to the similarity threshold, the use case to be analyzed is added to the similar use case set.

5. The method according to claim 4, characterized in that The similarity evaluation function is: S(T1,T2)0=αS(T1,T2)1+(1-α)S(E(T1),E(T2))2, Among them, S(T1, T2)0 represents the comprehensive similarity, S(T1, T2)1 represents the structural similarity between the benchmark abstract syntax tree and the comparison abstract syntax tree, S(E(T1), E(T2))2 represents the cosine similarity between the benchmark coding vector and the comparison coding vector, E(T1) represents the benchmark coding vector, E(T2) represents the comparison coding vector, T1 represents the benchmark abstract syntax tree, T2 represents the comparison abstract syntax tree, V1 represents the node set of the benchmark abstract syntax tree, V2 represents the node set of the comparison abstract syntax tree, λ is the regularization parameter, v represents the intersection node, sim(v) represents the intersection node similarity, children(v) represents the intersection child nodes of the intersection node, n[children(v)]1 represents the number of child nodes of the intersection node whose similarity in V1 and V2 is greater than or equal to the similarity threshold, n[children(v)]2 is the total number of intersection child nodes, and α represents the preset weight.

6. The method according to claim 5, characterized in that Also includes: Count the length of arithmetic operator code snippets at intersection nodes; Get the total length of the code snippet of the intersection node; The ratio of the length of the arithmetic operator code snippet to the total length of the code snippet is calculated and set as the preset weight.

7. The method according to claim 1, characterized in that The effective code ratio of the similar test case set is counted and set as the effective probability of the test case, including: Obtaining a first similar use case in the set of similar use cases; Construct validity evaluation function: Among them, w1, w2, w3, w4 and w5 represent weight parameters, which sum to 1, R represents the use case repeatability parameter, C represents the use case coverage parameter, Cl represents the use case clarity parameter, Co represents the use case completeness parameter, Ru represents the use case reusability parameter, and θ represents the threshold parameter, which is used to adjust the offset of the function; analyzing, according to the effectiveness evaluation function, a first similar use case effectiveness parameter of the first similar use case; When the validity parameter of the first similar use case is greater than or equal to a validity parameter threshold, marking the first similar use case with a valid code; The proportion of similar use cases with the valid code identifier in the similar use case set is counted and set as the validity probability of the test case.

8. A low-code generation system for security vulnerability test cases, characterized in that: include: A test case generation unit is used to obtain vulnerability types, functional descriptions, input constraints and expected results, input test case generation networks, and obtain test cases; A similar code retrieval unit, used to perform similar code retrieval on the test case based on a similarity threshold to obtain a similar case set; An effective code analysis unit, used to count the effective code ratio of the similar test case set and set it as the effective probability of the test case; The test case feedback unit is used to send the test case to the test terminal when the effective probability of the test case is greater than or equal to the effective probability threshold.

Citation Information

Patent Citations

  • Case generation model training method, test case generation method and storage medium

    CN115168218A

  • Test case reusing method based on code similarity

    CN115237758A

  • Intelligent contract vulnerability detection method and system based on Tree-LSTM and BiLSTM

    CN117195220A

  • Automatic vulnerability mining method and system for security protection

    CN118312968A