Vulnerability mining method based on grammar parse tree

By using the vulnerability mining method based on syntax parsing tree in web applications, the problem of inefficient testing of traditional fuzz testing methods is solved, and more efficient vulnerability mining capabilities are achieved, reducing the generation of invalid test cases.

CN120046156APending Publication Date: 2025-05-27BEIJING INST OF COMP TECH & APPL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510084559.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional fuzz testing methods are inefficient in testing in web applications, easily generate a large number of invalid test cases, and it is difficult to effectively explore deep code vulnerabilities.

Method used

A vulnerability mining method based on the syntax parsing tree is adopted to syntax parsing the input data of the Web interface, a syntax parsing tree is generated, and a seed pool is established for each data structure segment. Test cases are generated through data type, structured and semantic variants, and code coverage is monitored and testing strategies are adjusted.

Benefits of technology

It improves testing efficiency and vulnerability mining capabilities, reduces the generation and duplicate testing of invalid test cases, and can more effectively explore deep-level code vulnerabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046156A_ABST
    Figure CN120046156A_ABST
Patent Text Reader

Abstract

The invention relates to a vulnerability mining method based on a grammar parse tree, and belongs to the technical field of network security. According to the method and the device, the grammar analysis tree is generated by performing grammar analysis on the input data, and targeted variation is performed on different data structures based on the analysis tree, so that the test efficiency and the vulnerability mining capability are improved. The method can be widely applied to safety testing of Web application programs, and developers are helped to discover and repair safety holes in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network security technology, and particularly relates to a vulnerability mining method based on a syntax parse tree. Background Art

[0002] With the rapid development of the Internet, the security of Web applications has received increasing attention. Web applications usually interact with clients through interfaces, and the data formats transmitted by these interfaces are mainly JSON and XML. In Web applications, different business request data will trigger different execution paths, and there may be security vulnerabilities in these execution paths. Fuzz testing is a commonly used vulnerability mining technology, which triggers potential vulnerabilities in the program by generating a large number of mutated input data.

[0003] However, traditional fuzz testing methods often mutate the input data as a whole, ignoring the internal structure and semantics of the data, resulting in low testing efficiency, easy generation of a large number of invalid test cases, and even repeated testing of the same code path. Summary of the Invention

[0004] (1) Technical Problems to be Solved

[0005] The technical problem to be solved by the present invention is to provide a vulnerability mining method to improve testing efficiency and vulnerability mining ability.

[0006] (2) Technical Solutions

[0007] To solve the above technical problems, the present invention provides a vulnerability mining method based on a syntax parse tree, including the following steps:

[0008] Step 1: Construction of the syntax parse tree

[0009] Perform syntax parsing on the input data of the Web interface to generate a syntax parse tree, and each leaf node of the syntax parse tree corresponds to a data structure segment of the input data;

[0010] Step 2: Establishment of the seed pool

[0011] Establish a corresponding seed pool for each leaf node of the syntax parse tree, and the initial seeds of the data structure segments are stored in the seed pool. The initial seeds are generated by test cases or read from a configuration file;

[0012] Step 3: Generation of test cases

[0013] Select seeds from the seed pools of each leaf node according to the structure of the syntax parse tree, and mutate the seeds according to the data type to generate new test cases; The mutation methods include:

[0014] Data type mutation: adding, deleting, modifying numerical data, and fuzzifying strings;

[0015] Structured mutation: inserting, deleting, and sorting operations on arrays;

[0016] Semantic mutation: mutating according to the semantics of data, including incrementing or decrementing timestamps;

[0017] Step 4, Monitoring code coverage

[0018] Monitor the code coverage of the program under test through instrumentation technology. When the generated test cases trigger new code paths, add the corresponding seeds to the seed pool and update the weights of the seeds;

[0019] Step 5, Exception detection and reporting

[0020] Monitor the logs and returned status codes of the program under test. When an anomaly is detected, record the anomaly information and generate a report;

[0021] Step 6, Adjustment of test strategy

[0022] When no new code paths are triggered or no anomalies are detected in the program under test within a preset long time, adjust the test strategy according to the weights of the seeds and the code coverage information.

[0023] Preferably, each leaf node is a separate structure that mounts the corresponding seed pool, and the seeds stored in the seed pool are some seeds corresponding to this data part. Each seed has a structure that records its own relevant information for subsequent calculation of seed weights. The relevant information includes: the weight of the seed, the code blocks that the seed can cover, and the execution path; in step 3, the node data blocks and structure information in the leaf nodes generate new test cases through the data packet splicing algorithm.

[0024] Preferably, in step 3, the data packet splicing algorithm for generating new test cases is as follows:

[0025] 1) Obtain the root node of the syntax analysis tree;

[0026] 2) Traverse the leaf nodes of the syntax tree in level order, select seeds from the leaf node data structure, judge the data type, and transfer to one of steps 3) - 5) according to the data type;

[0027] 3) If the input data requires carrying the system time or contains CRC, generate correct data to replace the subsequent extended values and add them to the data queue, then transfer to step 6);

[0028] 4) If the data delimiter remains unchanged, directly add it to the data queue and transfer to step 6);

[0029] 5) If it is a string, number, or logical value, perform mutation of the corresponding data type, add it to the data queue, and go to step 6);

[0030] 6) Select data from the data queue in sequence for data splicing to generate an initial test case;

[0031] 7) Perform rationality verification of the test case and output the final test case.

[0032] Preferably, in step 4, first create a structure corresponding to the seed, correct the code block that can be covered by it, and then add it to the seed pool. The seed here records the code block that it can cover for subsequent weight calculation.

[0033] Preferably, the algorithm for adding the seed to the seed pool in step 4 is as follows:

[0034] 1’) Perform a level-order traversal of the created syntax parsing tree, determine whether the traversed node is a leaf node. If it is a leaf node, go to step 2’); if it is not a leaf node, continue the level-order traversal;

[0035] 2’) If the traversed node is a leaf node, perform the code block coverage correction process; replace the value of the leaf node with other random values in the corresponding seed pool, and splice them into a new test case and send it to the target test program. Determine whether this new code path can still be covered. If it can, it means that the seed value of this data part has no necessary connection with the generation of the new path, so the seed value of this data part is not added to the seed pool; if not, it means that the seed value is related to the generation of this new path, and the correction process is also a test of the target program and is not a redundant operation, so no additional consumption is added;

[0036] 3’) Create a structure corresponding to the seed, fill in the values of the structure and the code blocks that it can cover, and add the seed to the seed pool.

[0037] Preferably, the data structure section includes key-value pairs and arrays in JSON.

[0038] Preferably, adjusting the test strategy includes preferentially testing code paths with a coverage lower than a preset threshold.

[0039] The present invention also provides a system for implementing the above method.

[0040] The present invention also provides an application of the above method in the field of network security technology.

[0041] The present invention also provides an application of the above system in the field of network security technology.

[0042] (III) Beneficial effects

[0043] A vulnerability mining method based on a syntax parse tree proposed by the present invention has the following advantages:

[0044] 1) Improve test efficiency: By structurally processing the input data through a syntax parse tree, blind mutation of the entire data packet is avoided, and the generation of invalid test cases is reduced.

[0045] 2) Reduce repeated testing: By independently mutating different data structures, repeated testing of shared code blocks is avoided.

[0046] 3) Improve vulnerability mining ability: By monitoring code coverage and adjusting the test strategy, deeper code vulnerabilities can be more effectively mined. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic diagram showing different execution processes triggered by different business data in a Web program;

[0048] Figure 2 It is a schematic diagram of the syntax parse tree structure and the data packet splicing process. DETAILED DESCRIPTION OF THE INVENTION

[0049] To make the objectives, contents, and advantages of the present invention clearer, the following further describes in detail the specific embodiments of the present invention with reference to the drawings and embodiments.

[0050] The present invention proposes a vulnerability mining method based on a syntax parse tree. By performing syntax parsing on the input data, a syntax parse tree is generated, and targeted mutation of different data structures is performed based on the parse tree, thereby improving test efficiency and vulnerability mining ability.

[0051] The formats of data transmitted through Web interfaces are mainly JSON and XML. Taking the Json format as an example, the writing format of JSON data is key / value pairs; JSON values can be: strings, numbers (integers or floating-point numbers), logical values, etc.; it has two structures, objects and arrays; various complex data can be represented through these two structures. From the perspective of the fields in the data packet, it contains multiple key-value pairs, and their corresponding data types are also different. Different processing methods are adopted for these different business data in the target Web program, triggering different execution paths, as Figure 1 shown.

[0052] Business request data of different types and contents will generate different execution paths in a Web program, and these execution paths may also contain shared code blocks with common functions (d and e in the above figure). When the Web program processes business request data, different processing methods and code logics are given for different business requests (α, β, γ, δ in the above figure). For operations such as system login and password modification, although the processing processes of business request data are independent of each other, the same code segment may be called in different processing processes. Both the α and β request programs will directly or indirectly call the code segment d. Therefore, in the fuzz testing process, when testing the processing programs α and β, the d part of the code will be repeatedly tested, resulting in duplicate testing work. Therefore, if the fuzzer treats the input message data part as a whole, it will affect the efficiency of mutant data testing.

[0053] Therefore, the present invention adopts different processing methods and execution paths for different data parts of the request data content. By setting corresponding seed pools for different types and functions of data blocks through a parse tree, and combining mutants into input data packets, the overall quality of test cases can be more effectively improved, and deeper code areas can be tested.

[0054] Based on the above ideas, the present invention proposes a vulnerability mining method based on a syntax parse tree. The basic process of this method is as follows:

[0055] (1) Model the input data packet for the Web management end, analyze the service interface to obtain the parse tree for creating input data of test cases, and establish a corresponding seed pool for each data structure segment. The initial seeds can be obtained from test cases or provided in a configuration file. Then generate a parse tree as shown in Figure 2 . A seed pool corresponding to a data block is attached under each leaf node of the tree, and subsequent seed selection and generation will be processed based on this tree structure.

[0056] (2) Combine the established parse tree, select seeds from each seed pool, and mutate them using the mutation method corresponding to the data type. Finally, combine them together to generate new input data. The combination method avoids complex calculations and time consumption during data parsing, and ensures that the data generated each time conforms to the input format.

[0057] (3) Obtain code coverage through instrumentation. If a new path is generated, put the corresponding data block into the corresponding seed pool, and save the relevant information of the code block and path covered during execution. According to the feedback of the target test program, assign weights to the seeds.

[0058] (4) Monitor the log of the target program and the status code it returns to determine if an exception occurs. If an exception occurs, record it for subsequent exception report analysis.

[0059] (5) When the program under test has no exceptions for a long time or no new paths are generated, subsequent tests are guided based on conditional probability.

[0060] Figure 2 The parsing tree and data packet splicing are shown as follows. Each leaf node is a separate structure that mounts the corresponding seed pool. The seed pool stores some seeds corresponding to this data part. Each seed has a structure that records its own relevant information for subsequent seed weight calculation. Its relevant information includes: the weight of the seed, the code blocks that the seed can cover, the execution path, etc. The node data block and structure information in the leaf node generate test cases through the data packet splicing algorithm. The algorithm process is as follows:

[0061]

[0062]

[0063] During the process of generating data packets, seed mutation is required. Different mutation methods are used for different data types of data, making the seed mutation more targeted. When a new path is generated, the data block needs to be added to the corresponding seed pool. At this time, a structure corresponding to the seed needs to be created first, and the code blocks it can cover are corrected, and then it is added to the seed pool. The seed here records the code blocks it can cover for subsequent weight calculation. The main idea is to bind each data block to the code blocks in the program and clarify the relationship between the data part and the code blocks. Therefore, a seed addition algorithm is designed as follows:

[0064]

[0065] Taking the Web interface in JSON format as an example, assume the input data format of the interface is as follows:

[0066] json

[0067] {

[0068] "username":"admin",

[0069] "password":"123456",

[0070] "timestamp":1632000000

[0071] }

[0072] 1. Construction of the syntax parse tree

[0073] Perform syntax parsing on the above JSON data to generate a syntax parse tree, as Figure 2 shown. Among them, the root node is a JSON object, and the leaf nodes are key-value pairs

[0074] "username":"admin", "password":"123456", and "timestamp":

[0075] 1632000000.

[0076] 2. Establishment of the seed pool

[0077] Establish a corresponding seed pool for each leaf node:

[0078] o The seed pool corresponding to "username" contains several username seeds, such as

[0079] "admin", "test", "user1", etc.

[0080] o The seed pool corresponding to "password" contains several password seeds, such as

[0081] "123456", "password", "12345678", etc.

[0082] o The seed pool corresponding to "timestamp" contains several timestamp seeds, such as 1632000000, 1632000001, 1632000002, etc.

[0083] 3. Generation of test cases

[0084] Select seeds from each seed pool and perform mutation to generate new test cases.

[0085] For example:

[0086] o Mutate "username" to "admin'", mutate "password" to "123456",

[0087] Keep "timestamp" unchanged, and generate a test case:

[0088]

[0089] oMutate "timestamp" to 1632000000 + 1000 and generate test cases:

[0090]

[0091] 4. Monitoring of code coverage

[0092] Monitor the code coverage of the program under test through instrumentation technology. When the generated test cases trigger new code paths, add the corresponding seeds to the seed pool and update the weights of the seeds.

[0093] 5. Exception detection and reporting

[0094] Monitor the logs and returned status codes of the program under test. When an exception is detected, record the exception information and generate a report. For example, when a test case triggers an SQL injection vulnerability, record the test case and the corresponding exception information.

[0095] 6. Adjustment of test strategy

[0096] When no new code paths are triggered or no exceptions are detected in the program under test for a long time, adjust the test strategy according to the weights of the seeds and the code coverage information. For example, give priority to testing code paths with lower coverage.

[0097] It can be seen that a vulnerability mining method based on the syntax parse tree proposed by the present invention generates a syntax parse tree by performing syntax parsing on the input data, and performs targeted mutation on different data structures based on the parse tree, thereby improving the test efficiency and vulnerability mining ability. This method can be widely applied to the security testing of Web applications to help developers discover and fix security vulnerabilities in a timely manner.

[0098] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A vulnerability mining method based on grammar parse tree, characterized in that: The following steps are involved: Step 1: Construction of grammar parse tree Perform syntax analysis on the input data of the Web interface to generate a syntax parse tree, where each leaf node of the syntax parse tree corresponds to a data structure segment of the input data; Step 2: Establishment of seed pool A corresponding seed pool is established for each leaf node of the grammar parsing tree. The seed pool stores the initial seed of the data structure segment. The initial seed is generated by the test case or read from the configuration file. Step 3: Test case generation According to the structure of the grammar parse tree, seeds are selected from the seed pool of each leaf node, and the seeds are mutated according to the data type to generate new test cases; Variations include: Data type variation: add, delete, and modify numerical data, and perform fuzzy processing on character strings; Structural mutation: insert, delete, and sort arrays; Semantic mutation: mutation based on the semantics of the data, including incrementing or decrementing timestamps; Step 4: Code coverage monitoring The code coverage of the program under test is monitored through the instrumentation technology. When the generated test case triggers a new code path, the corresponding seed is added to the seed pool and the seed weight is updated; Step 5: Anomaly detection and reporting Monitor the logs and returned status codes of the program under test. When an exception is detected, record the exception information and generate a report. Step 6: Adjustment of testing strategy When no new code path is triggered or no exception is detected in the program under test for a preset long time, the test strategy is adjusted according to the seed weight and code coverage information.

2. The method according to claim 1, characterized in that Each leaf node is a separate structure, which mounts the corresponding seed pool, and the seed pool stores some seeds corresponding to the data part. Each seed has a structure that records its own related information for subsequent seed weight calculation. The related information includes: the weight of the seed, the code block that the seed can cover, and the execution path; in step 3, the node data block and structure information in the leaf node generate new test cases through the data packet splicing algorithm.

3. The method according to claim 2, characterized in that In step 3, the data packet splicing algorithm for generating new test cases is as follows: 1) Get the root node of the grammar parse tree; 2) Traverse the leaf nodes of the syntax tree in order, select seeds from the leaf node data structure, determine the data type, and go to one of steps 3)-5) according to the data type; 3) If the input data is required to carry the system time or contain CRC, the correct data is generated to replace the subsequent extended value and added to the data queue, and then go to step 6); 4) If the data delimiter does not change, add it directly to the data queue and go to step 6); 5) If it is a string, number, or logical value, mutate the corresponding data type, add it to the data queue, and go to step 6); 6) Select data from the data queue in sequence for data splicing to generate initial test cases; 7) Perform test case rationality check and output the final test case.

4. The method according to claim 3, characterized in that In step 4, first create a structure corresponding to the seed, correct the code block it can cover, and then add it to the seed pool. The seed here records the code block it can cover for subsequent weight calculations.

5. The method according to claim 4, characterized in that In step 4, the algorithm for adding seeds to the seed pool is as follows: 1') Traverse the syntax parse tree created by level-order traversal to determine whether the traversed node is a leaf node. If it is a leaf node, go to step 2'). If it is not a leaf node, continue level-order traversal; 2') If the traversed node is a leaf node, the code block coverage correction process is performed; the leaf node value is replaced with other random values ​​in the corresponding seed pool, and it is spliced ​​into a new test case and sent to the target test program to determine whether the new code path can still be covered. If it can, it means that the seed value of the data part is not necessarily related to the generation of the new path, and the seed value of the data part is not added to the seed pool; if it cannot, it means that the seed value is related to the generation of the new path, and the correction process is also a test of the target program, not a redundant operation, so there is no additional consumption; 3') Create a structure corresponding to the seed, fill in the value of the structure and the code block it can cover, and add the seed to the seed pool.

6. The method according to claim 1, characterized in that The data structure segment includes key-value pairs and arrays in JSON.

7. The method according to claim 1, characterized in that Adjusting the testing strategy includes prioritizing code paths whose coverage is below a preset threshold.

8. A system for implementing the method according to any one of claims 1 to 7.

9. Application of the method according to any one of claims 1 to 7 in the field of network security technology.

10. An application of the system as claimed in claim 8 in the field of network security technology.

Citation Information

Patent Citations

  • Seed scheduling weight distribution method for data packet splicing

    CN115048298A

  • Hybrid fuzzy test optimization method for mining security vulnerabilities

    CN117331826A