Directional fuzzy test method and device based on neural network model auxiliary input variation, equipment and storage medium
By using neural network models to assist input mutation in directional fuzz testing, the problem of lack of directionality of input mutation in the prior art is solved, and more efficient target code coverage and vulnerability triggering is achieved.
Patent Information
- Application Number
- CN202510151541.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-20
AI Technical Summary
The existing directional fuzz testing technology lacks directionality in the input mutation process, resulting in a large number of invalid mutations and failing to effectively reach the target code with specific control flow requirements.
A method based on neural network model assisted input mutation is adopted, and the key input fields are determined by static analysis and model training of the test program, and the fields are mutated based on the neural network model to generate effective mutation input.
Reduces the number of generated invalid mutant inputs, improves the efficiency of directional fuzzing, reduces system overhead costs, and significantly improves the coverage and vulnerability triggering capabilities of target codes.
Smart Images

Figure CN120179549A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a directed fuzz testing method, device, equipment and storage medium based on neural network model-assisted input mutation. Background Art
[0002] Fuzz testing technology is one of the most widely used vulnerability mining technologies. Fuzz testing can be divided into two categories according to its goals: coverage-guided fuzz testing and directed fuzz testing. The goal of coverage-guided fuzz testing is to cover as much program code as possible, while directed fuzz testing aims to reach specific target locations in the code. In recent years, directed grey-box fuzz testing has high practical value in tasks such as vulnerability reproduction, vulnerability verification, and patch testing.
[0003] Existing directed fuzz testing technologies can be divided into two categories according to strategies: 1) using the distance to the target code as feedback to schedule seed inputs; 2) optimizing input execution by performing target relevance analysis or target reachability analysis on program code. However, these strategies do not focus on the input mutation step.
[0004] In the prior art, a coverage-guided fuzz testing mutation scheme is usually adopted, and offsets and lengths are randomly selected during the mutation process. However, directed fuzz testing requires reaching target code with specific control flow requirements, which usually requires mutating specific input fields (i.e., key fields) to approach these target locations. However, the randomly selected offset and length mutation strategy adopted in the existing work lacks directionality and often results in a large number of invalid mutations. Therefore, there is great room for improvement in the input mutation of existing directed fuzz testing technologies. Summary of the Invention
[0005] The present invention provides a directed fuzz testing method, device, equipment and storage medium based on neural network model-assisted input mutation, which can reduce the generation of invalid mutated input fields and reduce the system overhead cost.
[0006] In a first aspect, the present invention provides a directed fuzz testing method based on neural network model-assisted input mutation, including the following steps: Determine the target code, and preprocess the program to be tested and the preset historical test inputs to determine the key input fields; Mutate the key input fields to obtain mutated key input fields; Based on the mutated key input fields and the preset fuzz testing engine, perform fuzz testing on the target code; Wherein, the preprocessing at least includes performing static analysis processing on the program to be tested, and performing model training processing based on the program to be tested and the historical test inputs.
[0007] Preferably, for the directed fuzz testing method based on neural network model-assisted input mutation provided by the present invention, the static analysis and processing of the program to be tested includes: Performing static analysis and processing on the program to be tested to obtain the program branches that affect the target code coverage; Constructing a branch encoder according to the program branches.
[0008] Preferably, for the directed fuzz testing method based on neural network model-assisted input mutation provided by the present invention, the model training and processing based on the program to be tested and the historical test inputs includes: Training a neural network model based on the historical test inputs and the encoding vectors; Wherein, the input of the neural network model is the historical test input, and the output of the neural network model is the encoding vector; The encoding vector is obtained by encoding the execution path corresponding to the historical test input using the branch encoder; the execution path is the path when the historical test input executes the test in the program to be tested.
[0009] Preferably, for the directed fuzz testing method based on neural network model-assisted input mutation provided by the present invention, after determining the trained neural network model, the method includes: Extracting the gradient information of the neural network model; Determining the key input fields from the historical test inputs based on the gradient information; The gradient information represents the gradient value composed of each execution path and the corresponding historical test input.
[0010] Preferably, for the directed fuzz testing method based on neural network model-assisted input mutation provided by the present invention, the determining the key input fields from the historical test inputs based on the gradient information includes: Based on the gradient information, determining the byte sequence of the historical test input with the largest gradient value in the historical test inputs as the target input field; Filtering out the noise fields from the target input field based on a preset gradient filtering strategy to determine the key input fields; Wherein, the gradient filtering strategy is to take the intersection of the input fields corresponding to the (n - 2)th, nth, and (n + 2)th dominant basic blocks when solving the key input fields of the nth dominant basic block to obtain the key input fields; the noise fields at least include fields with fixed values and keyword fields of adjacent dominant basic blocks.
[0011] Preferably, for the directed fuzz testing method based on neural network model-assisted input mutation provided by the present invention, the mutation processing of the critical input fields to obtain mutated critical input fields includes: Mutating the critical input fields based on a first mutation strategy to obtain the mutated critical input fields; or Mutating the critical input fields based on a second mutation strategy to obtain the mutated critical input fields; Wherein, the first mutation strategy is to sequentially mutate the critical input fields according to the field values within a preset field range; the second mutation strategy is to mutate the critical input fields according to preset dictionary values.
[0012] In a second aspect, the present invention provides a directed fuzz testing device based on neural network model-assisted input mutation, including: A preprocessing module, configured to determine target code, and perform preprocessing on the program to be tested and preset historical test inputs to determine critical input fields; A mutation module, configured to perform mutation processing on the critical input fields to obtain mutated critical input fields; A testing module, configured to perform fuzz testing on the target code based on the mutated critical input fields and a preset fuzz testing engine; Wherein, the preprocessing at least includes performing static analysis processing on the program to be tested, and performing model training processing based on the program to be tested and the historical test inputs.
[0013] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the directed fuzz testing method based on neural network model-assisted input mutation as described in any one of the above.
[0014] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the directed fuzz testing method based on neural network model-assisted input mutation as described in any one of the above.
[0015] In a fifth aspect, the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the directed fuzz testing method based on neural network model-assisted input mutation as described in any one of the above.
[0016] A method, device, equipment and storage medium for directed fuzz testing based on neural network model-assisted input mutation provided by the present invention determine target code, and based on preprocessing the program to be tested and preset historical test inputs, determine key input fields; perform mutation processing on the key input fields to obtain mutated key input fields; based on the mutated key input fields and a preset fuzz testing engine, perform fuzz testing on the target code; wherein, the preprocessing at least includes performing static analysis processing on the program to be tested, and performing model training processing based on the program to be tested and the historical test inputs. It can reduce the generation of invalid mutated input fields and reduce the system overhead cost. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 is one of the schematic flowcharts of the method for directed fuzz testing based on neural network model-assisted input mutation provided by the present invention.
[0019] Figure 2 is the schematic diagram of the node that dominates the basic block provided by the present invention.
[0020] Figure 3 is the second schematic diagram of the method for directed fuzz testing based on neural network model-assisted input mutation provided by the present invention.
[0021] Figure 4 is the schematic structural diagram of the device for directed fuzz testing based on neural network model-assisted input mutation provided by the present invention.
[0022] Figure 5 is the schematic structural diagram of the electronic equipment provided by the present invention. Detailed Embodiments
[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0024] In the related art, at least the following technical problems still exist: There are some mature input mutation schemes in the coverage-guided fuzzing scenario, mainly including correlation analysis techniques and symbolic execution techniques. However, these schemes face huge challenges when migrated to the directed fuzzing scenario.
[0025] Correlation analysis techniques, such as taint tracking, input-to-state relationships, mutation masking, etc., are widely used to optimize the input mutation in the coverage-guided fuzzing scenario. These techniques guide the fuzzing to mutate sensitive bytes by analyzing the correspondence between input bytes and constraints in the program, thereby breaking through the constraints, and are very effective in solving early input validity checks in the program such as magic bytes checks and checksums checks. However, correlation analysis is prone to false positives and false negatives when dealing with many other types of constraints, especially when the input fields corresponding to the constraints are variable-length fields. In the directed fuzzing scenario, the target code is often deep in the program and requires breaking through a large number of different types of constraints to reach. Therefore, the ability of correlation analysis techniques to solve constraints is difficult to meet the needs of directed fuzzing.
[0026] Symbolic / concolic execution is another type of input generation technique widely combined with fuzzing. This technique represents program constraints as symbolic expressions and uses a solver to analyze the constraints along the entire path, with a powerful ability to solve constraints. However, in real programs, there are often a large number of paths that can reach the goal of directed fuzzing, which makes symbolic execution encounter scalability problems such as path explosion when generating reachable inputs.
[0027] Existing mutation optimization techniques are difficult to meet the requirements of the directed fuzzing scenario in terms of both effectiveness and overhead. Therefore, it is necessary to develop a new mutation optimization technique for directed fuzzing. This technique supports the analysis of various constraints without relying on heavy path analysis, and has both lightweight and powerful analysis capabilities. Through this method, the effectiveness of the directed fuzzing technique can be effectively improved, with more powerful vulnerability discovery and verification capabilities, and can respond faster to emerging security threats.
[0028] The following combines Figures 1-5 Describe a directed fuzzing method, device, equipment and storage medium based on a neural network model-assisted input mutation of the present invention, which can reduce the generation of invalid mutated input fields and reduce the system overhead cost.
[0029] Figure 1It is one of the schematic flowcharts of a method for directed fuzz testing based on neural network model-assisted input mutation provided by the present invention. As Figure 1 shown, the method may include but is not limited to steps S100 to S300: S100, determine the target code, and based on preprocessing the program to be tested and the preset historical test inputs, determine the key input fields; S200, perform mutation processing on the key input fields to obtain mutated key input fields; S300, based on the mutated key input fields and the preset fuzz testing engine, perform fuzz testing on the target code; wherein, the preprocessing at least includes performing static analysis processing on the program to be tested, and performing model training processing based on the program to be tested and the historical test inputs.
[0030] In step S100 of some embodiments, determine the target code, and based on preprocessing the program to be tested and the preset historical test inputs, determine the key input fields.
[0031] It should be noted that the target code is the test target of directed fuzz testing, usually a specific line of code in the program to be tested. Directed fuzz testing aims to test this line of code faster and more. The target code can be the patched code in patch testing and the vulnerability trigger site in vulnerability reproduction according to different application scenarios.
[0032] The key input fields are the test input fields used to test the target code.
[0033] It can be understood that the preprocessing at least includes performing static analysis processing on the program to be tested, performing model training processing based on the program to be tested and the historical test inputs, and also includes extracting gradient information processing.
[0034] Further, performing static analysis processing on the program to be tested may specifically include but is not limited to: performing static analysis processing on the program to be tested to obtain the program branches that affect the coverage of the target code; constructing a branch encoder according to the program branches.
[0035] In some embodiments, the target dominance analysis technique is used to extract the dominant basic blocks of the target basic block in the control flow graph of the program to be tested as the target code, and at the same time record all the branches of these dominant basic blocks as the program branches. In this way, we narrow the huge search space of the program to be tested to the code strongly related to the target code.
[0036] It should be noted that target dominance analysis is a mature existing technology, which will analyze all the dominant basic blocks and / or necessary basic blocks of the target basic block in the program control flow graph.
[0037] The program branches that affect the target code coverage are the path branches composed of the dominating basic blocks and / or the essential basic blocks corresponding to the target basic blocks.
[0038] The branch encoder is used to encode the coverage status of the input execution path on the program branches into a multi-dimensional tensor. The branch encoder has at least the following technical advantages: 1) Accurately associate the input with the branch behavior, eliminate the interference of irrelevant branches through domination level sorting and sibling branch encoding, and focus on the target critical path.
[0039] 2) Support variable-length / variable-offset fields. The tensor encoding does not depend on the field position assumption and can adaptively identify the dynamically changing field boundaries.
[0040] 3) Lightweight design, only requiring a storage complexity of O(k) (k is the number of dominating basic blocks), significantly lower than the path explosion overhead of symbolic execution.
[0041] 1. Domination Basic Block Sorting and Tensor Dimension Mapping Input: The set of target-related dominating basic blocks extracted through static analysis (sorted by domination level, denoted as dom_bbs = [bb_1, bb_2,..., bb_k]).
[0042] Tensor Initialization: Create a one-dimensional tensor T of length k, with each dimension corresponding to a dominating basic block, arranged in the order of domination level (T[i] corresponds to bb_i).
[0043] Example 1: If the domination chain of the target code contains 5 dominating basic blocks (bb_1 to bb_5), then initialize the tensor T = <0,0,0,0,0>.
[0044] 2. Coverage Status Encoding Rules Rule 1 (Covered Dominating Basic Block): If the input execution path covers the dominating basic block bb_i, then set T[i] = 1.
[0045] Rule 2 (The First Uncovered Dominating Basic Block): Let bb_m be the first uncovered dominating basic block (i.e., bb_1 to bb_{m - 1} are covered, and bb_m is not covered): Classification of Predecessor Dominating Basic Blocks: Type A: bb_{m - 1} is directly connected to bb_m (through a single "domination edge").
[0046] Type B: There are non-dominating basic blocks between bb_{m - 1} and bb_m.
[0047] Processing of Type A: Extract all non-dominating edges of bb_{m-1} (i.e., branches that do not point to bb_m).
[0048] Generate an (n-1)-bit binary number (n is the total number of out-edges of bb_{m-1}), with each bit corresponding to the coverage status of a non-dominating edge (covered is 1, not covered is 0).
[0049] Convert the binary number to decimal and normalize it to [0,1]: Type B processing: Ignore the branch behavior of bb_{m-1} and set T[m] = 0.
[0050] Example 2: The dominating basic block bb_3 is the first uncovered block, and its predecessor bb_2 is of type A with 3 out-edges (including 1 dominating edge).
[0051] The non-dominating edge coverage status is 10 (binary), and the normalized value is 2 / 4 = 0.5. Therefore, T[3] = 0.5, and the encoding result is <1,1,0.5,0,0> 3. Sibling Branch Encoding Technique Problem background: Sibling branches of the same conditional statement (such as different cases in a switch statement) share the same key input fields, but traditional encoding cannot associate uncovered branches with key fields.
[0052] Solution: For the predecessor (type A) of the first uncovered dominating basic block, indirectly associate the key fields of sibling branches through non-dominating edge coverage status encoding.
[0053] Mathematical proof: Let the conditional statement variable be x, and its corresponding input field be f.
[0054] If the branch condition of bb_{m-1} is x = c (case c), then: Cover the specific value of field f corresponding to x = c; Cover the variations of field f that the sibling branches (such as x = a and x = b) also depend on.
[0055] By encoding the sibling branch status, the model can learn the global impact of field f on branch behavior.
[0056] Example 3: The program branch is x = 100 (uncovered), and its sibling branches x = 0 and x = 50 have been covered.
[0057] After encoding the status of x = 0 and x = 50, the model locates field f through the gradient to guide the variation of f to trigger x = 100.
[0058] It should be further noted that for the branch encoder, the first predecessor node that does not cover the dominated basic block: If it is of type A, generate a normalized value according to its non-dominated edge coverage status and write it into the tensor; If it is of type B, ignore its branch behavior and set the corresponding tensor dimension to 0.
[0059] Furthermore, the calculation method of the normalized value is as follows: Divide the binary value of the non-dominated edge coverage status by 2^{n - 1} (n is the total number of outgoing edges of the predecessor node).
[0060] Figure 2 It is a schematic diagram of the nodes of the dominated basic block provided by the present invention, a branch encoding example. For example, map the coverage situation of each dominated basic block to one dimension of the output vector. If a dominated basic block (dom_bbs[i]) is covered, it is considered that the branch behavior of the previous dominated basic block (dom_bbs[i - 1]) is appropriate, and the corresponding dimension of dom_bbs[i] is set to 1. Conversely, if dom_bbs[i] is not covered but dom_bbs[i - 1] is covered, according to the characteristics of the dominance tree, dom_bbs[i - 1] will be the last covered dominated basic block, and dom_bbs[i] is the first uncovered dominated basic block. Covering dom_bbs[i] will bring us closer to the target point. In this case, adjust the branch behavior of dom_bbs[i - 1] to cover dom_bbs[i].
[0061] It should be noted that further use a fraction between 0 and 1 to represent the branch behavior of the last covered dominated basic block (dom_bbs[i - 1]). Specifically, the dominated basic blocks are divided into two categories: type A and type B. A type A dominated basic block s refers to a dominated basic block s that is directly adjacent to the next dominated basic block and is connected by a single edge (referred to as a "dominant edge"), as shown in Figure 2 (a), (b), and (d) in. While a type B dominated basic block s refers to the situation where there are other non-dominated basic blocks s between it and the next dominated basic block, as shown in Figure 2 (c) in.
[0062] If the last covered dominating basic block in the execution path is of type A, the next dominating basic block can be covered by modifying its branching behavior once. Therefore, the branching behavior of the type-A dominating basic block s is highly correlated with reaching the target point. We mark all the output edges of the type-A dominating basic block s. For a type-A dominating basic block with n output edges, we use an (n - 1)-bit binary number to represent the coverage status of its n - 1 "non-dominating edges", where coverage is represented as 1 and non-coverage is represented as 0. Subsequently, we divide this binary number by 1 << n (2^n) and place the result in the output vector dimension corresponding to the next dominating basic block.
[0063] Conversely, if the last covered dominating basic block is of type B, modifying its branching behavior may not help in approaching the target point. For example, in the figure, modifying the branching behavior of the dominating basic block c does not directly affect the path "2 -> 5 -> 6 -> 8". To simplify the analysis, we ignore the branching behavior of the type-B dominating basic block s and simply set the corresponding dimension to 0.
[0064] In some embodiments of the present invention, the model training process based on the program to be tested and the historical test inputs includes: Training a neural network model based on the historical test inputs and the encoding vectors; wherein, the input of the neural network model is the historical test input, and the output of the neural network model is the encoding vector; The encoding vector is obtained by encoding the execution path corresponding to the historical test input using the branch encoder; the execution path is the path when the historical test input is executed in the program to be tested.
[0065] It should be noted that the model structure of the neural network model is designed as follows: Input layer: Receives an input byte sequence (fixed length of 1024 bytes, padded with zeros if insufficient), the historical test input.
[0066] Hidden layer: A 3-layer fully connected network, with 512 neurons in the hidden layer and the activation function being ReLU.
[0067] Output layer: Normalized by the Sigmoid function, outputs a branch encoding tensor (dimension equal to the number of dominating basic blocks), the encoding vector obtained by encoding the historical branches (execution path) corresponding to the historical test input.
[0068] The training parameters include at least but are not limited to the following: Optimizer: Adam (learning rate 0.001); Loss function: Mean Squared Error (MSE); Batch size: 32.
[0069] Example 4: Implement the model using PyTorch. After 500 rounds of training, the model can accurately predict the impact of the input byte sequence on branch coverage.
[0070] In some embodiments of the present invention, after determining the trained neural network model, the method includes: Extracting the gradient information of the trained neural network model; Determining the critical input field from the historical test inputs based on the gradient information; The gradient information represents the gradient value composed of each execution path and the corresponding historical test input.
[0071] Specifically, obtain the gradient information (gradient matrix) of the output with respect to the input bytes through backpropagation.
[0072] In some embodiments of the present invention, the determining the critical input field from the historical test inputs based on the gradient information includes: Based on the gradient information, determining the byte sequence of the historical test input with the largest gradient value in the historical test inputs as the target input field; Filtering out the noise fields from the target input field based on a preset gradient filtering strategy to determine the critical input field; Wherein, the gradient filtering strategy is to take the intersection of the input fields corresponding to the (n - 2)th, nth, and (n + 2)th dominated basic blocks when solving the critical input field of the nth dominated basic block to obtain the critical input field; the noise fields at least include fields with fixed values and key fields of adjacent dominated basic blocks.
[0073] Gradient calculation for keyword field positioning: Obtain the gradient matrix of the output with respect to the input bytes through backpropagation.
[0074] Noise filtering: For the nth dominated basic block, filter the intersection of its gradient value with the gradient values of the (n - 2)th, nth, and (n + 2)th dominated basic blocks, and exclude magic bytes (such as 0x7F in ELF files) and adjacent key bytes.
[0075] Field clustering: Use the DBSCAN algorithm (eps = 3, min_samples = 2) to aggregate adjacent bytes with high gradient values to identify critical input fields.
[0076] In step S200 of some embodiments, mutate the critical input field to obtain a mutated critical input field.
[0077] In some embodiments of the present invention, the mutating the critical input field to obtain a mutated critical input field includes: Mutate the key input field based on the first mutation strategy to obtain the mutated key input field; or Mutate the key input field based on the second mutation strategy to obtain the mutated key input field; wherein, the first mutation strategy is to mutate the key input field successively according to the field values within a preset field range; the second mutation strategy is to mutate the key input field according to a preset dictionary value.
[0078] The first mutation strategy is byte-by-byte mutation: traverse all values from 0 to 255 for the key input field. The preset field range is 0 to 255.
[0079] The second mutation strategy is dictionary replacement: replace the key input field with a preset value (such as the boundary value "0xFFFF").
[0080] In step S300 of some embodiments, perform fuzz testing on the target code based on the mutated key input field and a preset fuzz testing engine.
[0081] In Figure 3 In the lower left part, the fuzz testing engine (AFL) replacement mutation strategy is the method of the present invention. Perform boundary value replacement for the keyword field "offset" and successfully trigger the target code.
[0082] In some embodiments of the present invention, it at least further includes AFL fuzz testing engine framework adaptation, which specifically includes the following steps: Instrumentation modification: Record the out-edge coverage status of the dominating basic block in the instrumentation code of AFL.
[0083] Mutation module replacement: Replace the native mutate() function with the gradient-guided mutation logic of the present invention.
[0084] Real-time interaction: Model training and fuzz testing are executed in parallel, and keyword field information is passed through shared memory.
[0085] The directed fuzz testing method based on neural network model-assisted input mutation provided by the present invention has a target code coverage rate increased by about 70% and a vulnerability triggering time shortened to 80% compared with the traditional AFL.
[0086] Figure 3 It is the second schematic diagram of the directed fuzz testing method based on neural network model-assisted input mutation provided by the present invention, which may include but is not limited to three stages: static analysis, model training, and gradient mutation.
[0087] In the static analysis stage, step 1: Extract the interprocedural dominance chain Input: The LLVM intermediate representation (IR) bitcode of the program to be tested. Figure 3The program under test in it is the program to be tested in the embodiments of the present invention.
[0088] Call graph construction: Use the CallGraph analysis module of LLVM to extract the program call graph and determine the function call chain from the program entry to the target code.
[0089] Dominator tree analysis: For each function in the function call chain, extract the dominator tree through the DominatorTreeWrapperPass to identify the dominance chain from the function entry basic block to the call point basic block.
[0090] Interprocedural dominance chain synthesis: Connect the dominance chains of each function in the call order to form the interprocedural dominance chain of the target code. If there are multiple call points for the called function, take the intersection of the dominance chains of each call point.
[0091] Instrumentation: Insert instrumentation code at all out-edges of the dominant basic block to record the branch coverage status.
[0092] Embodiment 1.1: Taking Figure 2 the upper left part as an example, the target code is located in the function func_c, and the program entry is the main function. Through dominance chain analysis, extract the set of dominant basic blocks of the path main -> func_a -> func_b -> func_c, and insert instrumentation to track the out-edge coverage of each dominant basic block.
[0093] Branch encoding module Step 2: Branch coverage tensor generation Tensor dimension mapping: Sort the dominant basic blocks according to the dominance level and correspond to each dimension of the tensor in turn.
[0094] Coverage status encoding: If the input covers a dominant basic block, set its corresponding dimension to 1; If not covered, check whether its predecessor is a type A dominant basic block (directly connected to the next dominant basic block): Type A: Generate an (n - 1)-bit binary number (n is the number of out-edges), mark the coverage status of non-dominant edges (covered as 1, not covered as 0), and write it to the dimension corresponding to the next dominant basic block after normalization.
[0095] Type B (not directly connected): Ignore and set the corresponding dimension to 0.
[0096] Example: The path "2 -> 4" covers the dominant basic blocks 1 and 2, and the dominant basic block 2 is of type A, and its out-edges include dominant and non-dominant edges. The coverage status of the non-dominant edge is "10" (binary), normalized to 0.5, and encoded as <1, 1, 0.5, 0, 0> (as Figure 2 shown).
[0097] Dynamic dataset generation phase, step 3: Training set refinement Initial screening: Screen the inputs that cover the top 20% of the number of dominated basic blocks.
[0098] Similarity clustering: For the seed inputs that have undergone ≤ 2 mutations, calculate the Hamming distance of their byte distributions; use the K-means algorithm for clustering, and retain one representative input for each cluster.
[0099] Dynamic update: Trigger condition: Discover a new dominated basic block or more than 10 minutes have passed since the last training; retrain the model after updating the training set.
[0100] Example 2.1: Assume that the fuzz testing discovers a new dominated basic block "dom_5", immediately trigger the dataset update, remove the outlier inputs and add new samples, and then retrain the model.
[0101] The model training phase, gradient filtering, and mutation phase are as described in the above embodiments and will not be elaborated here.
[0102] As Figure 3 As shown in the upper right part, a neural network model is used to guide the input mutation. Our neural network model is a three-layer multi-layer perceptron (MLP) implemented using PyTorch. The ReLU function is used as the activation function for the hidden layer, and the sigmoid function is used for the output layer to normalize each output dimension to the range [0, 1]. The hidden layer of the model contains 512 neurons.
[0103] As Figure 3 As shown in the lower left part, in the embodiment of the present invention, a fuzz testing engine is used as the base to carry the input mutation module designed by the present invention. This fuzz testing engine needs to conform to the mainstream framework of AFL (American Fuzzy Lop). Replace the original random input mutation strategy with the strategy of this solution in the form of a patch in the framework source code.
[0104] The present invention provides a directed fuzz testing method, device, equipment, and storage medium based on a neural network model to assist input mutation. By determining the target code, and based on preprocessing the program to be tested and the preset historical test inputs, the key input fields are determined; the key input fields are mutated to obtain mutated key input fields; based on the mutated key input fields and the preset fuzz testing engine, the target code is fuzz tested; wherein, the preprocessing at least includes static analysis processing of the program to be tested, and model training processing based on the program to be tested and the historical test inputs. It can reduce the generation of invalid mutated input fields and reduce the system overhead cost.
[0105] The following describes the device for directed fuzz testing that aids input mutation based on a neural network model provided by the present invention. The device for directed fuzz testing that aids input mutation based on a neural network model described below can be correspondingly referred to in relation to the method for directed fuzz testing that aids input mutation based on a neural network model described above.
[0106] Figure 4 FIG. 4 is a schematic structural diagram of the device for directed fuzz testing that aids input mutation based on a neural network model provided by the present invention. A device for directed fuzz testing that aids input mutation based on a neural network model includes: A preprocessing module 410, configured to determine a target code, and based on preprocessing the program under test and preset historical test inputs, determine key input fields; A mutation module 420, configured to perform mutation processing on the key input fields to obtain mutated key input fields; A testing module 430, configured to perform fuzz testing on the target code based on the mutated key input fields and a preset fuzz testing engine; Wherein, the preprocessing at least includes performing static analysis processing on the program under test, and performing model training processing based on the program under test and the historical test inputs.
[0107] Preferably, the device for directed fuzz testing that aids input mutation based on a neural network model provided by the present invention is specifically configured to perform static analysis processing on the program under test to obtain program branches that affect the coverage of the target code; construct a branch encoder according to the program branches.
[0108] Preferably, the device for directed fuzz testing that aids input mutation based on a neural network model provided by the present invention is specifically configured to train a neural network model based on the historical test inputs and an encoding vector; Wherein, the input of the neural network model is the historical test input, and the output of the neural network model is an encoding vector; The encoding vector is obtained by encoding the execution path corresponding to the historical test input using the branch encoder; the execution path is the path when the historical test input executes the test in the program under test.
[0109] Preferably, the device for directed fuzz testing that aids input mutation based on a neural network model provided by the present invention is specifically configured to extract gradient information of the trained neural network model; Determine the key input fields from the historical test inputs based on the gradient information; The gradient information represents the gradient value formed by each execution path and the corresponding historical test input.
[0110] Preferably, the directed fuzz testing device based on neural network model-assisted input mutation provided by the present invention is specifically configured to, based on the gradient information, determine the byte sequence of the historical test input with the largest gradient value among the historical test inputs as the target input field; Based on a preset gradient filtering strategy, filter out the noise fields from the target input field to determine the critical input field; Wherein, the gradient filtering strategy is to take the intersection of the input fields corresponding to the (n-2)-th, n-th, and (n+2)-th dominating basic blocks when solving the critical input field of the n-th dominating basic block to obtain the critical input field; the noise fields at least include fields with fixed values and keyword fields of adjacent dominating basic blocks.
[0111] Preferably, the directed fuzz testing device based on neural network model-assisted input mutation provided by the present invention is specifically configured to perform mutation processing on the critical input field based on a first mutation strategy to obtain the mutated critical input field; or Perform mutation processing on the critical input field based on a second mutation strategy to obtain the mutated critical input field; Wherein, the first mutation strategy is to sequentially mutate the critical input field according to the field values within a preset field range; the second mutation strategy is to mutate the critical input field according to a preset dictionary value.
[0112] The present invention also provides a directed fuzz testing method, device, equipment, and storage medium based on neural network model-assisted input mutation. By determining the target code and preprocessing the program to be tested and preset historical test inputs, the critical input field is determined; mutation processing is performed on the critical input field to obtain the mutated critical input field; based on the mutated critical input field and a preset fuzz testing engine, fuzz testing is performed on the target code; wherein, the preprocessing at least includes performing static analysis processing on the program to be tested and performing model training processing based on the program to be tested and the historical test inputs. It can reduce the generation of invalid mutated input fields and reduce the system overhead cost.
[0113] Figure 5 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute a directed fuzz testing method based on neural network model-assisted input mutation. The method includes: determining target code, and based on preprocessing the program to be tested and preset historical test inputs, determining key input fields; performing mutation processing on the key input fields to obtain mutated key input fields; based on the mutated key input fields and a preset fuzz testing engine, performing fuzz testing on the target code; where the preprocessing at least includes performing static analysis processing on the program to be tested, and performing model training processing based on the program to be tested and the historical test inputs.
[0114] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, external hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0115] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a directed fuzz testing method based on neural network model-assisted input mutation provided by the above-mentioned various methods. The method includes: determining target code, and based on preprocessing the program to be tested and preset historical test inputs, determining key input fields; performing mutation processing on the key input fields to obtain mutated key input fields; based on the mutated key input fields and a preset fuzz testing engine, performing fuzz testing on the target code; where the preprocessing at least includes performing static analysis processing on the program to be tested, and performing model training processing based on the program to be tested and the historical test inputs.
[0116] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a directed fuzz testing method based on neural network model-assisted input mutation, and the method includes: determining target code, and performing preprocessing on the program to be tested and preset historical test inputs to determine key input fields; performing mutation processing on the key input fields to obtain mutated key input fields; based on the mutated key input fields and a preset fuzz testing engine, performing fuzz testing on the target code; wherein the preprocessing at least includes performing static analysis processing on the program to be tested, and performing model training processing based on the program to be tested and the historical test inputs.
[0117] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, also by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0119] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A directed fuzzy testing method based on neural network model assisted input variation, characterized in that: include: Determine the target code and pre-process the program to be tested and the preset historical test input to determine the key input fields; Performing mutation processing on the key input field to obtain a mutated key input field; Based on the mutated key input field and a preset fuzz testing engine, performing fuzz testing on the target code; The preprocessing at least includes static analysis of the program to be tested and model training based on the program to be tested and the historical test input.
2. The directional fuzzy testing method based on neural network model assisted input variation according to claim 1 is characterized in that: The static analysis of the program to be tested includes: Performing static analysis on the program to be tested to obtain program branches that affect the coverage of the target code; A branch encoder is constructed according to the program branches.
3. The directional fuzzy testing method based on neural network model assisted input variation according to claim 2 is characterized in that: The model training process based on the program to be tested and the historical test input includes: Training a neural network model based on the historical test input and the encoding vector; Wherein, the input of the neural network model is the historical test input, and the output of the neural network model is the encoding vector; The encoding vector is obtained by encoding the execution path corresponding to the historical test input using the branch encoder; the execution path is the path taken when the historical test input is tested in the program to be tested.
4. The directional fuzzy testing method based on neural network model assisted input variation according to claim 3 is characterized in that: After determining the trained neural network model, the method includes: Extracting gradient information of the neural network model; determining the key input fields from the historical test inputs based on the gradient information; The gradient information represents a gradient value composed of each of the execution paths and the corresponding historical test inputs.
5. The directional fuzzy testing method based on neural network model assisted input variation according to claim 4 is characterized in that: The determining the key input field from the historical test input based on the gradient information comprises: Based on the gradient information, determining a byte sequence of the historical test input having the largest gradient value from the historical test input as a target input field; Based on a preset gradient filtering strategy, filtering the noise field from the target input field to determine the key input field; Among them, the gradient filtering strategy is to take the intersection of the input fields corresponding to the n-2th, nth and n+2th dominating basic blocks when solving the key input field of the nth dominating basic block to obtain the key input field; the noise field at least includes a field with fixed values and the key field of the adjacent dominating basic blocks.
6. The directed fuzzy testing method based on neural network model assisted input variation according to any one of claims 1 to 5, characterized in that: The step of performing mutation processing on the key input field to obtain a mutated key input field includes: Perform mutation processing on the key input field based on a first mutation strategy to obtain the mutated key input field; or Performing mutation processing on the key input field based on a second mutation strategy to obtain the mutated key input field; Among them, the first mutation strategy is to mutate the key input fields in sequence according to the field values within the preset field range; the second mutation strategy is to mutate the key input fields according to the preset dictionary values.
7. A directional fuzzy testing device based on neural network model assisted input variation, characterized in that: include: A preprocessing module is used to determine the target code and preprocess the program to be tested and the preset historical test input to determine the key input fields; A mutation module, used for performing mutation processing on the key input field to obtain a mutated key input field; A testing module is used to perform fuzz testing on the target code based on the mutated key input field and a preset fuzz testing engine; wherein the preprocessing at least includes static analysis processing on the program to be tested, and model training processing based on the program to be tested and the historical test input.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, it implements the directional fuzzy testing method based on neural network model assisted input variation as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the directed fuzzy testing method based on neural network model-assisted input variation as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the directed fuzzy testing method based on neural network model-assisted input variation as described in any one of claims 1 to 6.