Format perception fuzzy test method and system based on taint analysis
Through the format-aware fuzz testing method based on stain analysis, the shortcomings of existing tools in input area dependence and structure recognition are solved, more efficient test case generation is achieved, and the effectiveness of fuzz testing is improved.
Patent Information
- Application Number
- CN202510772762.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-11
AI Technical Summary
The existing format-aware fuzz testing tools lack dependence on input area and structure recognition, resulting in the generated test cases that cannot meet program requirements and are inefficient.
Using a taint analysis method, we use instruction-level taint tracking and file format tree construction to identify the area type and dependencies of the input file to generate legitimate test cases.
It improves the efficiency and accuracy of fuzz testing, can effectively identify the input file area and association relationship, and generate test cases that are more in line with program requirements.
Smart Images

Figure CN120277683A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and particularly relates to a format-aware fuzz testing method and system based on taint analysis. Background Art
[0002] Vulnerability detection is an important issue in the field of software security. With the continuous development of software technology and the increasing complexity of related product functions, the number of vulnerabilities is increasing day by day. Many codes with security risks are inevitably exploited by malicious attackers, which not only brings great troubles to ordinary users, but also causes huge property losses and security threats to companies.
[0003] As a dynamic software testing method, fuzz testing has shown remarkable effects in exploring and revealing unknown security vulnerabilities. Format-aware fuzz testing understands the format of the input based on feedback information during the fuzz testing process, thereby more effectively changing the input and improving the fuzz testing efficiency. Usually, fuzz testing tools generate test cases according to preset strategies and conduct comprehensive testing on the target application program, and reveal potential security vulnerabilities by monitoring the exceptions during program operation. The key to improving the fuzz testing efficiency is to generate inputs in valid formats. A standard format input case usually contains a sequence of data fields with specific semantics: for example, buffer size, checksum, etc. The program parses these fields according to the functions it implements, and illegal input cases may be rejected in the early stage of program execution, making it difficult to improve the test coverage of the program.
[0004] Therefore, how to solve the problem that the existing format-aware fuzz testing tools mainly focus on identifying fields and lack the recognition of input area dependencies and structures, resulting in the generated test cases not meeting the program requirements, is a technical problem that those skilled in the art need to solve urgently. Summary of the Invention
[0005] To achieve the object of the present invention, the present application provides a format-aware fuzz testing method based on taint analysis, including: Step S1: Input the input sample and the seed queue into the target program for testing, and perform taint analysis at the instruction level on the target program to obtain a list of taint instructions; Step S2: Based on the list of taint instructions, divide the input sample into regions by bytes, identify the region types and dependency relationships according to the instruction control flow and data flow relationships, and construct a file format tree and a dependency relationship table; Step S3: Mutate the nodes of the file format tree according to the file format tree, obtain the dependency relationships of the mutated nodes, adjust the relevant fields according to the dependency relationships of the mutated nodes to generate legal test cases, and add the test cases to the seed queue.
[0006] In some of the specific embodiments, step S1 includes: Step S11: Perform instruction-level pollution tracking on the target program to obtain the execution trace of the input sample; Step S12: Use the input sample as the stain source, select bytes as the basic stain propagation unit with byte as the stain propagation granularity, and determine the individual offset of the contaminated value in the input sample; Step S13: Track the execution instructions of the target program, obtain the byte behavior characteristics of the input sample, and generate a stain instruction list; Step S14: Merge the redundant records of the stain instruction list and optimize the record format.
[0007] In some of the specific embodiments, in step S2, constructing the file format tree includes: Step S21: Sort and classify the stain instructions according to the initial byte and region size of the stain area based on the stain instruction list; Step S22: Traverse the instructions in the order of the initial byte of the stain area, specify the region boundary according to the termination byte of the stain instruction, and divide the input sample into file regions; Step S23: Initialize the file format tree; Step S24: Recursively perform steps S21 - S23 to add other file region nodes to the file format tree and initialize the current node. When the region interval of the node to be inserted belongs to the region interval of the current node, search for a node in the child node list of the current node such that the region interval of the node to be inserted belongs to the region of the child node. If no child node that meets the condition can be found, add the node to be inserted to the child node region list; if a child node that meets the condition is found, point the current node to the child node; Step S25: Recursively execute step S24 until the node to be inserted is added to the file format tree; Step S26: Continue to execute steps S24 and S25 until all region nodes are added to the file format tree.
[0008] In some of the specific embodiments, in step S2, constructing the dependency relationship table based on the dependency mutation method includes: Use the DFS method to traverse the file format tree to find the leaf nodes of the file format tree; Obtain the dependency relationship table elements of all nodes on the path from the leaf node to the root node and add them to the dependency queue of the leaf node; Perform exploratory mutation and destructive mutation according to the region dependency relationship rule set; Continue to traverse other nodes until all leaf nodes are traversed.
[0009] In some of the specific embodiments, step S3 further includes: Take out the first test case from the seed queue; input it into the target program for running or exception handling; determine whether an exception occurs. If an exception occurs, perform exception handling; otherwise, determine whether the current test case contains illegal characters or out-of-bounds access to structure members. If there are illegal characters or out-of-bounds access to structure members, directly discard the current test case.
[0010] To achieve the same inventive purpose, the present application also provides a format-aware fuzz testing system based on taint analysis, including: Taint instruction analysis module: used to input the input sample and the seed queue into the target program for testing, and perform taint analysis at the instruction level on the target program to obtain a list of taint instructions; Structure mutation module: used to divide the input sample into regions by bytes based on the list of taint instructions, identify the region types and dependency relationships according to the instruction control flow and data flow relationships, and construct a file format tree and a dependency relationship table; Test case generation module: used to mutate the nodes of the file format tree according to the file format tree, obtain the dependency relationships of the mutated nodes, adjust the relevant fields according to the dependency relationships of the mutated nodes to generate legal test cases, and add the test cases to the seed queue.
[0011] In some of the specific embodiments, the taint instruction analysis module is used to perform the following steps: Step S11: Perform contamination tracking at the instruction level on the target program to obtain the execution trace of the input sample; Step S12: Use the input sample as the taint source, bytes as the taint propagation granularity, select bytes as the basic taint propagation unit, and determine the individual offset of the contaminated value in the input sample; Step S13: Track the instructions executed by the target program, obtain the byte behavior characteristics of the input sample, and generate a list of taint instructions; Step S14: Merge the redundant records of the list of taint instructions and optimize the record format.
[0012] In some of the specific embodiments, in the structure mutation module, constructing the file format tree includes: Step S21: Sort and classify the taint instructions according to the initial byte and region size of the taint region based on the list of taint instructions; Step S22: Traverse the instructions in the order of the initial byte of the taint region, specify the region boundary according to the termination byte of the taint instruction, and divide the input sample into file regions; Step S23: Initialize the file format tree; Step S24: Recursively execute steps S21 - S23 to add other file region nodes to the file format tree and initialize the current node. When the region interval of the node to be inserted belongs to the region interval of the current node, search for a node in the list of child nodes of the current node such that the region interval of the node to be inserted belongs to the region of the child node. If no child node that meets the condition can be found, add the node to be inserted to the list of child node regions; if a child node that meets the condition is found, set the current node to point to the child node. Step S25: Recursively execute step S24 until the node to be inserted is added to the file format tree. Step S26: Continue to execute steps S24 and S25 until all region nodes are added to the file format tree.
[0013] In some specific embodiments, in the structure mutation module, constructing a dependency relationship table based on the dependency mutation method includes: Use the DFS method to traverse the file format tree to find the leaf nodes of the file format tree; Obtain the dependency relationship table elements of all nodes on the path from the leaf node to the root node and add them to the dependency queue of the leaf node; Perform exploratory mutation and destructive mutation according to the region dependency relationship rule set; Continue to traverse other nodes until all leaf nodes have been traversed.
[0014] In some specific embodiments, the test case generation module is further configured to: Take out the first test case from the seed queue; input it into the target program for running or exception handling; determine whether an exception occurs. If an exception occurs, perform exception handling; otherwise, determine whether the current test case contains illegal characters or out - of - bounds access to structure members. If there are illegal characters or out - of - bounds access to structure members, directly discard the current test case.
[0015] Advantages of the above - mentioned technical solutions: The present invention proposes a format - aware fuzz testing method based on taint analysis. Through the input file region boundary recognition and type recognition methods based on taint analysis, the input file regions and their association relationships can be effectively recognized. A method for saving the input structure based on a file format tree and a dependency relationship table is proposed, which can store seed file region information more efficiently and improve the mutation efficiency of fuzz testing. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0017] Figure 1 Schematic flowchart of a format-aware fuzz testing method based on taint analysis provided by an embodiment of the present invention; Figure 2 Schematic structural diagram of a format-aware fuzz testing system based on taint analysis provided by an embodiment of the present invention; Figure 3 Schematic framework diagram of a format-aware fuzz testing method based on taint analysis provided by an embodiment of the present invention; Figure 4 Schematic diagram of simplified taint analysis instruction operations for dynamic taint analysis and instruction analysis of a format-aware fuzz testing method based on taint analysis provided by an embodiment of the present invention; Figure 5 Schematic diagram of the method for recording a field dependency table of a format-aware fuzz testing method based on taint analysis provided by an embodiment of the present invention; Figure 6 Schematic diagram of a jpg file header format tree generated by a format-aware fuzz testing method based on taint analysis provided by an embodiment of the present invention; Figure 7 Schematic diagram of the dependency record of some nodes of a format-aware fuzz testing method based on taint analysis provided by an embodiment of the present invention. Detailed implementation manners
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, rather than all, embodiments of the present invention.
[0019] Examples of the embodiments are shown in the accompanying drawings, where the same or similar symbols represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.
[0020] Embodiment 1 An embodiment of the present invention provides a format-aware fuzz testing method based on taint analysis. Referring to Figure 1 、 Figure 3 shown, it includes: Step S1: Input the sample and the seed queue into the target program for testing, and perform taint analysis at the instruction level on the target program to obtain a list of tainted instructions; Step S2: Based on the list of tainted instructions, divide the input sample into regions by bytes, identify the region types and dependencies according to the instruction control flow and data flow relationships, and construct a file format tree and a dependency table; Step S3: Mutate the nodes of the file format tree according to the file format tree, obtain the dependencies of the mutated nodes, adjust the relevant fields according to the dependencies of the mutated nodes to generate legal test cases, and add the test cases to the seed queue.
[0021] Specifically, define that a file region is generated by a series of consecutive bytes of an input file, including the initial byte position, the region length, and the region type, expressed as ; when ignoring the region type, it can be abbreviated as ; define that the input format tree is a fine-grained representation method of the input file structure. The input format tree follows the general definition of a tree, and its nodes are composed of file regions and a list of sub-region pointers. The leaf nodes of the file format tree are expressed as ; define that the region dependency table is a record table representing the input region relationships. The region dependency table records the corresponding dependency fields or sets of valid values for each field. The description of the region dependency table is , where records the key value of the field, is the set of valid values, recording the corresponding dependency fields or sets of dependency values of field F.
[0022] In a specific embodiment of the present invention, Step S1 includes: Step S11: Perform pollution tracking at the instruction level on the target program to obtain the execution trace of the input sample; Step S12: Use the input sample as the taint source, bytes as the taint propagation granularity, select bytes as the basic taint propagation unit, and determine the single offset of the contaminated value in the input sample; Specifically, in order to obtain the byte behavior characteristics of the test sample, the taint tracking engine tracks the instructions executed by the target program, records the program jump and numerical operation instructions such as cmp, load, je, etc., and generates an execution trace list. The execution trace list is composed of taint tracking instructions and contains all the instructions involving taint variables during program execution.
[0023] The taint tracking instruction is defined as a quadruple , where: represents the address of the instruction, represents the opcode, that is, the operation type of the instruction, Represent two operands of the instruction, and each operand record is further defined as a binary tuple , where records the byte offset of the input file corresponding to the tainted operand, records the specific value of the constant variable.
[0024] Due to the existence of loop structures and single-byte comparison statements in the code, there are a large number of redundant statements in the trace instruction list. At the same time, many instructions with the same field data are executed at the same address, but multiple tainted trace instructions are generated, affecting the efficiency of field recognition.
[0025] Step S13: Trace the instructions executed by the target program, obtain the byte behavior characteristics of the input sample, and generate a tainted instruction list; Step S14: Merge the redundant records in the tainted instruction list and optimize the record format.
[0026] Specifically, according to the characteristics of the trace instruction list, redundant records are merged and the record format is optimized. The trace records are merged based on the following rules: (1) For consecutive instructions at the same address, merge the variable offset fields therein to generate one instruction.
[0027] (2) For instructions at adjacent addresses, if the operators are the same and the tainted fields of a certain operand are the same, merge the two instructions and merge the tainted fields of the different operands.
[0028] (3) Due to the repeated instructions generated by the program loop statements, only one is retained.
[0029] Based on the above rules, FieldsFuzz merges fine-grained operation instructions, streamlines the instruction analysis list, and reduces the computational overhead of region recognition.
[0030] In a specific embodiment of the present invention, in step S2, constructing the file format tree includes: Step S21: Sort and classify the tainted instructions according to the initial byte of the tainted region and the region size according to the tainted instruction list; Step S22: Traverse the instructions in the order of the initial byte of the tainted region, specify the region boundary according to the termination byte of the tainted instruction, and divide the input sample into file regions; Step S23: Initialize the file format tree; Step S24: Recursively perform steps S21 - S23 to add other file region nodes to the file format tree and initialize the current node. When the region interval of the node to be inserted belongs to the region interval of the current node, search for a node in the list of child nodes of the current node such that the region interval of the node to be inserted belongs to the child node region. If no child node satisfying the condition can be found, add the node to be inserted to the list of child node regions; if a child node satisfying the condition is found, set the current node to point to the child node. Step S25: Recursively execute step S24 until the node to be inserted is added to the file format tree. Step S26: Continue to execute steps S24 and S25 until all region nodes are added to the file format tree.
[0031] Specifically, differentiating the boundaries of different fields is the basis for dividing the input file regions. Due to the phenomenon of overlapping valid regions in the instructions, it is usually inappropriate to divide based on instruction - level fields. For example, an instruction to read bytes may parse multiple regions, resulting in incorrect merging of different fields or subdivision of fields of the same type. The present invention determines the input file region boundaries and the nested relationship of adjacent regions by creating an input format tree. The method for constructing the file format tree is described as follows: Step 1: Sort and classify the tainted instructions in the list of reduced instruction analyses according to the initial byte of the tainted region and the region size.
[0032] Step 2: Traverse the instructions in the order of the initial bytes, and divide the input sample into file regions according to the region boundaries specified by the termination bytes of the instructions. 。
[0033] Step 3: Initialize the file format tree. First, define the root node as 。Among them, , represents the length of the input seed.
[0034] Step 4: Use the recursive method to add other file region nodes to the file format tree in the order of Step 1. Initialize the current node as 。When the region interval of the node to be inserted belongs to the region interval of the current node, that is ,search for a node in the list of child nodes of the current node such that 。If no node satisfying the condition can be found, then add to ; if a node satisfying the condition is found, set the current node to point to 。
[0035] Step 5: Recursively execute step 4 until the node to be inserted is added to the file format tree.
[0036] Step 6: Continue to execute Step 4 and Step 5 until all region nodes are added to the file format tree.
[0037] The type of the file region directly affects the ability and efficiency of generating valid test cases during the mutation process in fuzz testing. Based on the file format standard and functions, the present invention determines that the main classifications of identifying valid file region types are shown in Table 1. There are differences in the instruction records and the tainted bytes of comparison instructions for different types of regions. In the region type identification stage, the method of the present invention initializes a region rule set, and the region rule set contains some common rule features. In the type identification stage, the method of the present invention enumerates the regions of all nodes in the file format tree, reads the list of tainted instructions in the region, extracts instruction features, and identifies the corresponding region type based on the tainted instruction feature rules corresponding to different regions.
[0038] Table 1 Region Identification Field Types and Rules
[0039] The purpose of dependency reasoning is to analyze the effective scope according to the list of tainted instructions in the region, so as to obtain the dependency relationship between different regions. Dependency reasoning focuses on the correlation between regions of input file length, checksum, and offset types and other regions. The present invention determines the dependency relationship between these special type regions and other data regions by extracting the list of tainted instructions in these regions and tracking the byte regions contaminated by the operands of cmp and load instructions. For example, the checksum field is compared with a certain variable, and the tainted byte region covered by this variable is usually the region verified by this checksum field. In order to efficiently implement the insertion and lookup operations of region relationships, the present invention constructs a region dependency table based on a hash table to save the region dependency relationship. The structure of the region dependency table is as Figure 4 、 Figure 5 shown. For different fields, the related fields and valid value sets of the field are recorded. For the Magic number and Enumeration fields, the valid value sets are saved in the dependency table. For the Length, Checksum, and Offset fields, the hash table records their action regions, and the data fields record the other fields related to them, which are saved in the form of a linked list. Based on the field dependency table, the dependency relationship between regions can be quickly added and queried.
[0040] In a specific embodiment of the present invention, in step S2, constructing a dependency relationship table based on the dependency mutation method includes: Traverse the file format tree using the DFS method to find the leaf nodes of the file format tree; Obtain the dependency table elements of all nodes on the path from the leaf node to the root node, and add them to the dependency queue of the leaf node; Perform exploratory mutation and destructive mutation according to the regional dependency relationship rule set; Continue to traverse other nodes until all leaf nodes are traversed.
[0041] Based on taint analysis, a file format tree and a dependency table are obtained. The present invention uses a file format mutation method to improve the efficiency of fuzz testing. The mutation method based on dependency relationships aims to achieve mutations at the regional level rather than being limited to mutations at the byte level. During the field mutation process, the dependent fields of the mutated field are modified according to the rules to ensure the validity of the mutated test cases. The rules of the mutation method based on dependency are as follows: Table 2 Field Mutation Rules of the Mutation Method Based on Dependency
[0042] In a specific embodiment of the present invention, step S3 further includes: Take out the first test case from the seed queue; input it into the target program for running or exception handling; determine whether an exception occurs. If an exception occurs, perform exception handling; otherwise, determine whether the current test case contains illegal characters or out-of-bounds access to structure members. If there are illegal characters or out-of-bounds access to structure members, directly discard the current test case.
[0043] To evaluate the format-aware fuzz testing method proposed by the present invention, this application implements this method based on the fuzz testing tool AFL and the taint analysis tool Pin, named FieldsFuzz.
[0044] Use the method proposed by the present invention to test the open-source software jhead, record the analysis results of the jpg file format and generate a visual file format tree, and record the start byte and end byte of the file format tree nodes as labels. Figure 6Shows the format tree for the APP0 field and the DQT dat[0] field of the seed file not_kitty.jpg. At the same time, use 010editor to parse the seed not_kitty.jpg provided by AFL and record its field division as shown in Table 3. By comparing and analyzing the generated results, it can be verified that FieldsFuzz largely restores the syntax structure corresponding to the test case. Although this tool exhibits high accuracy and consistency, there are certain deviations in the recognition in some specific areas. This is because the program does not fully follow the specifications during the file parsing process. FieldsFuzz uses the taint analysis method to identify the program processing process, and the identified file structure is consistent with the program parsing method. This application exports the dependency table generated by FieldsFuzz as a json file for recording, Figure 7 Shows partial Magic Number, Enumeration, and Length field dependency records. Through Figure 7 the recorded results shown, it is found that FieldsFuzz can effectively record the valid value sets or dependent fields of these three fields for the three fields, indicating that the field dependency table can effectively record and identify the field types and dependency records.
[0045] Table 3 JPG file header format record
[0046] To evaluate the performance of the method of the present invention in identifying the input file area, this application conducts a comparative analysis with other input structure inference methods (i.e., WEIZZ, ProFuzzerer, and NestFuzz). To correctly calculate the accuracy rate of area recognition, first preprocess the input file. For each input file, use the common format template from 010 Editor to parse the file and export all field records with boundary and type information, while manually marking the misidentified fields in the template and deleting some redundant fields. Subsequently, mark the area type according to the field function. Since the field area division granularity of each fuzzing tool is different, the accuracy rate of boundary analysis is defined as the number of areas correctly identified by the fuzzing tool divided by the total number of areas identified by the tool. For the area category, judge whether it is correct by comparing the area type defined by the tool with the actual function of the area. The precision rate is defined as the number of correctly identified area types divided by the total number of areas identified by the tool. Since WEIZZ only identifies the checksum field, the accuracy rate of area type recognition of WEIZZ is not counted. The accuracy rate of fuzzing attack to identify the input file area is shown in Table 4. The results show that the average accuracy of FieldsFuzz in area recognition accuracy and type recognition is 96.23% / 94.13%, which is higher than other comparison programs.
[0047] Accuracy Results of the Format-Aware Fuzzing Tool for Identifying File Regions in Table 4
[0048] Example Two An embodiment of the present invention provides a format-aware fuzzing system based on taint analysis. Referring to Figure 2 as shown, it includes: Taint Instruction Analysis Module 10: Used to test the input sample and the seed queue in the target program, and perform taint analysis at the instruction level on the target program to obtain a list of taint instructions; Structure Mutation Module 20: Used to divide the input sample into regions by bytes based on the list of taint instructions, identify the region types and dependency relationships according to the instruction control flow and data flow relationships, and construct a file format tree and a dependency relationship table; Test Case Generation Module 30: Used to mutate the nodes of the file format tree according to the file format tree, obtain the dependency relationships of the mutated nodes, adjust the relevant fields according to the dependency relationships of the mutated nodes to generate legal test cases, and add the test cases to the seed queue.
[0049] In a specific embodiment of the present invention, the taint instruction analysis module 10 is used to perform the following steps: Step S11: Perform pollution tracking at the instruction level on the target program to obtain the execution trace of the input sample; Step S12: Use the input sample as the taint source, bytes as the taint propagation granularity, select bytes as the basic taint propagation unit, and determine the single offset of the contaminated value in the input sample; Step S13: Track the execution instructions of the target program, obtain the byte behavior characteristics of the input sample, and generate a list of taint instructions; Step S14: Merge the redundant records in the list of taint instructions and optimize the record format.
[0050] In a specific embodiment of the present invention, in the structure mutation module 20, constructing the file format tree includes: Step S21: Sort and classify the taint instructions according to the initial byte and region size of the taint region based on the list of taint instructions; Step S22: Traverse the instructions in the order of the initial byte of the taint region, specify the region boundary according to the termination byte of the taint instruction, and divide the input sample into file regions; Step S23: Initialize the file format tree; Step S24: Recursively execute steps S21 - S23 to add other file area nodes to the file format tree and initialize the current node. When the area range of the node to be inserted belongs to the area range of the current node, search for a node in the child node list of the current node such that the area range of the node to be inserted belongs to the child node area. If no child node that meets the condition can be found, add the node to be inserted to the child node area list; if a child node that meets the condition is found, set the current node to point to the child node; Step S25: Recursively execute step S24 until the node to be inserted is added to the file format tree; Step S26: Continue to execute steps S24 and S25 until all area nodes are added to the file format tree.
[0051] In a specific embodiment of the present invention, in the structural variation module 20, constructing a dependency relationship table based on the dependency variation method includes: Use the DFS method to traverse the file format tree to find the leaf nodes of the file format tree; Obtain the dependency relationship table elements of all nodes on the path from the leaf node to the root node and add them to the dependency queue of the leaf node; Perform exploratory variation and destructive variation according to the area dependency relationship rule set; Continue to traverse other nodes until all leaf nodes have been traversed.
[0052] In a specific embodiment of the present invention, the test case generation module 30 is further configured to: Take out the first test case from the seed queue; input it into the target program for running or exception handling; determine whether an exception occurs. If an exception occurs, perform exception handling; otherwise, determine whether the current test case contains illegal characters or out - of - bounds access to structure members. If there are illegal characters or out - of - bounds access to structure members, directly discard the current test case.
[0053] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
[0054] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1Steps of the functions specified in one or more boxes. Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention. Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the said element.
[0055] The above has introduced the method and device provided by the present invention in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation of the present invention.
[0056] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", "one specific embodiment" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic description of the terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0057] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or equivalently replace some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A format-aware fuzz testing method based on taint analysis, characterized in that, Including: Step S1: Input the input sample and the seed queue into the target program for testing, perform taint analysis at the instruction level on the target program, and obtain a list of tainted instructions; Step S2: Based on the list of tainted instructions, divide the input sample into regions by bytes, identify the region types and dependencies according to the instruction control flow and data flow relationships, and construct a file format tree and a dependency table; Step S3: Mutate the nodes of the file format tree according to the file format tree, obtain the dependencies of the mutated nodes, adjust the relevant fields according to the dependencies of the mutated nodes to generate legal test cases, and add the test cases to the seed queue.
2. The format-aware fuzz testing method based on taint analysis according to claim 1, wherein Step S1 includes: Step S11: Perform instruction-level pollution tracking on the target program to obtain the execution trace of the input sample; Step S12: Use the input sample as the taint source, bytes as the taint propagation granularity, select bytes as the basic taint propagation unit, and determine the single offset of the contaminated value in the input sample; Step S13: Track the instructions executed by the target program, obtain the byte behavior characteristics of the input sample, and generate a list of tainted instructions; Step S14: Merge the redundant records of the list of tainted instructions and optimize the record format.
3. The format-aware fuzz testing method based on taint analysis according to claim 1, wherein, In step S2, constructing the file format tree includes: Step S21: Sort and classify the tainted instructions according to the initial byte and region size of the taint region according to the list of tainted instructions; Step S22: Traverse the instructions in the order of the initial byte of the taint region, specify the region boundary according to the termination byte of the tainted instruction, and divide the input sample into file regions; Step S23: Initialize the file format tree; Step S24: Recursively perform steps S21 - S23 to add other file region nodes to the file format tree and initialize the current node. When the region interval of the node to be inserted belongs to the region interval of the current node, find a node in the list of child nodes of the current node such that the region interval of the node to be inserted belongs to the region of the child node. If no child node that meets the condition can be found, add the node to be inserted to the list of child node regions; if a child node that meets the condition is found, point the current node to the child node; Step S25: Recursively execute step S24 until the node to be inserted is added to the file format tree; Step S26: Continue to execute steps S24 and S25 until all region nodes are added to the file format tree.
4. The format-aware fuzz testing method based on taint analysis according to claim 1, wherein In step S2, constructing the dependency table based on the dependency mutation method includes: Use the DFS method to traverse the file format tree to find the leaf nodes of the file format tree; Obtain the elements of the dependency table of all nodes on the path from the leaf node to the root node, and add them to the dependency queue of the leaf node; Perform exploratory mutation and destructive mutation according to the region dependency rule set; Continue to traverse other nodes until all leaf nodes are traversed.
5. The format-aware fuzz testing method based on taint analysis according to claim 1, wherein Step S3 also includes: Take out the first test case from the seed queue; input it into the target program for running or exception handling; determine whether an exception occurs. If an exception occurs, perform exception handling; otherwise, determine whether the current test case contains illegal characters or out-of-bounds access to structure members. If there are illegal characters or out-of-bounds access to structure members, directly discard the current test case.
6. A format-aware fuzz testing system based on taint analysis, characterized in that, Including: Taint instruction analysis module: used to test the input sample and the seed queue in the target program, perform taint analysis at the instruction level on the target program, and obtain a list of taint instructions; Structure mutation module: used to divide the input sample into regions by bytes based on the list of taint instructions, identify the region types and dependency relationships according to the instruction control flow and data flow relationships, and construct a file format tree and a dependency table; Test case generation module: used to mutate the nodes of the file format tree according to the file format tree, obtain the dependency relationships of the mutated nodes, adjust the relevant fields according to the dependency relationships of the mutated nodes to generate legal test cases, and add the test cases to the seed queue.
7. The format-aware fuzz testing system based on taint analysis according to claim 6, characterized in that The taint instruction analysis module is used to perform the following steps: Step S11: Perform contamination tracking at the instruction level on the target program to obtain the execution trace of the input sample; Step S12: Use the input sample as the taint source, bytes as the taint propagation granularity, select bytes as the basic taint propagation unit, and determine the single offset of the contaminated value in the input sample; Step S13: Track the instructions executed by the target program, obtain the byte behavior characteristics of the input sample, and generate a list of taint instructions; Step S14: Merge the redundant records in the list of taint instructions and optimize the record format.
8. The format-aware fuzz testing system based on taint analysis according to claim 6, characterized in that, In the structure mutation module, constructing the file format tree includes: Step S21: Sort and classify the taint instructions according to the initial byte and region size of the taint region based on the list of taint instructions; Step S22: Traverse the instructions in the order of the initial byte of the taint region, specify the region boundary according to the termination byte of the taint instruction, and divide the input sample into file regions; Step S23: Initialize the file format tree; Step S24: Recursively execute steps S21 - S23 to add other file region nodes to the file format tree and initialize the current node. When the region interval of the node to be inserted belongs to the region interval of the current node, find a node in the list of child nodes of the current node such that the region interval of the node to be inserted belongs to the child node region. If no child node that meets the condition can be found, add the node to be inserted to the list of child node regions; if a child node that meets the condition is found, point the current node to the child node; Step S25: Recursively execute step S24 until the node to be inserted is added to the file format tree; Step S26: Continue to execute steps S24 and S25 until all region nodes are added to the file format tree.
9. The format-aware fuzz testing system based on taint analysis according to claim 6, wherein In the structure mutation module, constructing the dependency table based on the dependency mutation method includes: Use the DFS method to traverse the file format tree to find the leaf nodes of the file format tree; Obtain the dependency table elements of all nodes on the path from the leaf node to the root node, and add them to the dependency queue of the leaf node; Perform exploratory mutation and destructive mutation according to the regional dependency rule set; Continue to traverse other nodes until all leaf nodes have been traversed.
10. The format-aware fuzz testing system based on taint analysis according to claim 6, characterized in that, The test case generation module is also used for: Take out the first test case in the seed queue; input it into the target program for running or exception handling; determine whether an exception occurs. If an exception occurs, perform exception handling; otherwise, determine whether the current test case contains illegal characters or out-of-bounds access to structure members. If there are illegal characters or out-of-bounds access to structure members, directly discard the current test case.
Citation Information
Patent Citations
Method for analyzing taint propagation path
CN103729295A
Industrial communication protocol reverse analysis method based on dynamic stain analysis
CN110213243A
Industrial control protocol grammar reverse analysis method under basic block granularity based on instrumentation
CN112905184A
Fuzzy testing method based on symbolic execution
CN115017516A
Method and device for extracting message format
US20150205963A1