Anti-fuzzy code identification method and device based on dynamic taint analysis
Through dynamic taint analysis, the function-level data flow diagram is constructed, and anti-fuzzy code is identified and removed, which solves the problem that redundant code is difficult to automatically remove in the existing Anti-fuzzing technology, and realizes automatic redundant code cleaning, reduces operation and maintenance costs and restores the performance of the fuzzer.
Patent Information
- Application Number
- CN202510200232.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-18
AI Technical Summary
In the existing Anti-fuzzing technology, the dependence between redundant code and the original code of the program is weak, and the modular characteristics of redundant code are obvious, which is difficult to resist the advanced program analysis technology, with high manual removal costs and increasing operation and maintenance workload.
Using a method based on dynamic taint analysis, a function-level data flow diagram is constructed, anti-fuzzy code is identified and removed, redundant code is cleaned through binary rewriting, and instruction-level instrumentation and data flow tracking is used for synchronous dynamic analysis framework of the execution engine and analysis engine, a tree-like data flow diagram is constructed, and redundant code is removed.
It realizes automatic removal of redundant code without changing the normal function of the program, restores the error search capability and coverage feedback mechanism of the fuzzer, and reduces operation and maintenance costs.
Smart Images

Figure CN120336162A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security technology, and particularly to an anti-fuzzy code recognition method and device based on dynamic taint analysis. Background Art
[0002] Fuzzing is currently the most efficient automated software defect detection technology. It not only provides convenience for developers to detect software defects, but also provides convenience for attackers to exploit vulnerabilities. Attackers continuously attempt to use Fuzzing technology to exploit vulnerabilities and then use the vulnerabilities to attack software systems. Since Fuzzing attacks play an important role in modern network security threats, traditional security defense measures have been unable to effectively cope with increasingly complex and hidden attack means.
[0003] To make up for the deficiencies in defense, anti-fuzzing technology targeting Fuzzing characteristics has emerged. Anti-fuzzing technology aims to develop advanced methods and strategies to detect and prevent the implementation of Fuzzing attacks. Previous anti-fuzzing research has shown that the key factor affecting Fuzzing efficiency is its coverage-based feedback mechanism. To protect the developer's advantage over attackers in error finding, anti-fuzzing technology compiles the target program into two versions: the original version and the fortified version. Developers retain the original version for thorough testing of the program and publicly release the fortified binary program. Anti-fuzzing technology inserts redundant code that does not affect the program function into the original code to mislead the coverage feedback mechanism of the fuzzer and make the Fuzzing process develop in the wrong direction. However, the dependency relationship between the redundant code added by anti-fuzzing and the original code of the program is weak, and the redundant code is highly modular and obvious in characteristics, making it easy to be recognized and difficult to resist the increasingly advanced program analysis technology. In daily operation and maintenance work, if redundant code added by the anti-fuzzing mechanism is to be removed manually, it will increase a huge amount of workload, seriously increasing operation and maintenance costs and human losses.
[0004] Therefore, how to analyze the fortified binary program without changing the normal function of the program and achieve the automatic removal of redundant code added by the anti-fuzzing mechanism has become a research topic. Summary of the Invention
[0005] An embodiment of the present invention provides an anti-fuzzing code recognition method and device based on dynamic taint analysis, which can analyze the fortified binary program without changing the normal function of the program, and achieve automatic removal of redundant codes added by the Anti-fuzzing mechanism.
[0006] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:
[0007] In the first aspect, the anti-fuzzing code recognition method based on dynamic taint analysis provided by the embodiments of the present invention includes:
[0008] S1. Input the binary program code fortified by the anti-fuzzing test mechanism into the binary dynamic analysis framework for taint analysis;
[0009] Among them, in this embodiment, instruction-level taint analysis is adopted, and the specific manifestation form of the analysis result is a data flow. What is recorded in the data flow is a series of instructions. For example: starting from the taint source (the register or memory marked as taint), the nodes passed through are all the instructions that the taint source propagates to during program execution.
[0010] S2. Use the taint analysis result to construct a data flow graph, and identify the anti-fuzzing code according to the data flow graph;
[0011] In this embodiment, the specific means of constructing the data flow graph using the taint analysis result is optimized. Since multiple taint analyses will obtain multiple data flows in practical applications, during the process of constructing the data flow graph, first promote the instruction-level data flow to the function level, which can be understood as traversing all the instructions of the program to see which function the instruction belongs to; at the same time, use the binary analysis tool IDApro to generate the function call graph of the target binary program; then, traverse all the data flows, and for the data flow relationship between every two data in the data flow, such as the data flow relationship from function A to function B, supplement the information of this edge on the function call graph; finally, obtain the function call graph containing the data flow relationship between functions, and use it as the data flow graph at the function level.
[0012] S3. According to the recognition result, rewrite the binary program code and remove the anti-fuzzing code.
[0013] Among them, the binary dynamic analysis framework includes: an execution engine and an analysis engine;
[0014] This framework includes two parts: an execution engine and an analysis engine. In this embodiment, taint analysis of the binary program is performed through the coordinated operation of these two engines. In the execution engine, instruction-level dynamic instrumentation is performed on the binary program code, and the instruction stream of the target program is dynamically intercepted during runtime; the process of dynamic instrumentation is shown as Figure 4As shown in the figure, the staking code is used to pause the original execution of the target program, collect the context information of the instructions to be executed, and synchronize it to the analysis engine.
[0015] Before each instruction is executed, the execution engine synchronizes the context information of the instruction to the analysis engine. The analysis engine that has completed the synchronization executes the instruction again based on the context information of the instruction and updates the taint analysis result in real time. As Figure 2 shown, the context information of the instruction includes: the operands of the instruction, the register information and memory information corresponding to the instruction, etc.
[0016] Specifically, in S2, it includes: traversing the data flow graph to extract the data flow starting from the main function, and regarding the data flow that terminates at the ordinary function as the redundant data flow to which the anti-obfuscation code belongs. The function code included in the redundant data flow is used as the anti-obfuscation code. In this embodiment, during the anti-obfuscation code recognition process, dynamic taint analysis is performed on the binary program after strengthening the anti-obfuscation test mechanism, and a data flow graph at the function level is constructed. Traverse the data flow graph, extract the data flow starting from the main function, regard the data flow that terminates at the ordinary function as the redundant data flow to which the anti-obfuscation code belongs, and regard the data flow that finally flows to the system call as the normal data flow to which the original function code of the program belongs. The function code included in the redundant data flow is the anti-obfuscation code. In practical applications, in the programming environment of C / C++ programs, the main function refers to the main function, which is a specific function and the entry of the program; except for the main function, other user-defined functions and API functions called from external libraries can be called ordinary functions.
[0017] Among them, S2 is specifically implemented as:
[0018] Identify the input of the target program and the API (Application Programming Interface) function calls related to the input of the target program. Among them, the target program refers to the binary program to be analyzed, which is a program strengthened by anti-obfuscation technology and includes anti-obfuscation code. The input of the target program includes: the command line parameters of the main function; mark the input of the target program as the taint source, change different inputs (including legal inputs and illegal inputs) and record the functions to which the taint propagates, and record the functions to which the taint propagates as the key functions of the target program. Among them, the input of the target program, the parameters and return values of the key functions can be marked as taint sources respectively, and the taint analysis process can be executed multiple times.
[0019] Then, mark the parameters and return values of the key functions as taint sources. Based on taint analysis technology, perform inter-procedural data flow tracking, construct a data flow graph at the function level. Among them, use the main function as the root node, and the key functions called by the main function as the second-layer nodes to establish a tree-shaped data flow graph, and repeat the nodes in the loop structure in the graph in the tree; Taint analysis refers to the general process of marking taints, tracking the taint propagation process, and analyzing the taint propagation path. "Inter-procedural" is a professional term, referring to the analysis process that spans multiple functions in the program. Traverse all paths starting from the root node in the data flow graph, and each path starting from the root node is regarded as a data flow. The data flow that finally flows to the system call is the normal data flow containing the original function of the program, while the data flow that terminates at an ordinary function is the redundant data flow containing the de-obfuscation code. For example: Take a simple program as an example, Figure 5 On the left is the source code of the program, whose function is to calculate the sum of a and b in the add function and output it in the main function; on the right are the main assembly instructions after compilation under the x86-64 architecture, including the function add and the function main. Among them, mark the variable a as a taint source. Take the execution of instruction 7 as an example to illustrate the real-time update process of the taint analysis result. After the analysis engine executes instruction 6, it waits for the execution engine to synchronize the information of the next instruction. At this time, the result of maintaining the taint analysis in the analysis engine is that the register EDX is contaminated. Before executing instruction 7 and after executing instruction 6, the execution engine inserts instrumentation code. The information synchronized by the instrumentation code to the analysis engine includes: the content of instruction 7 (the instruction to be executed), the value of register EDX is 1, and the value of register EAX is 2 (the context information of instruction 7). After receiving the instruction information, the analysis engine parses the semantics of the instruction and executes it. The semantics of this instruction is to add the value of register EDX to the value of register EAX, and the result is stored in EDX. Since this instruction uses the contaminated register EDX, mark this instruction as contaminated. This operation is to update the taint analysis result.
[0020] Take the execution of instruction 22 as an example to explain the processing of external functions (the same applies to system calls). The execution engine uses the tool QBDI. When simulating the execution of a binary program, the tool can jump to the link library normally, find the corresponding function and complete the call. Since QBDI is only a binary simulation execution tool and does not provide a specific data flow analysis tool, and the general instruction-level binary analysis tool cannot handle external function calls, a dynamic analysis framework that synchronizes the execution engine and the analysis engine is designed. After the execution engine synchronizes the information of instruction 22, it calls the printf function normally and prints the variable c to the standard output stream. After the analysis engine receives the information of instruction 22, it cannot find the specific instructions of the printf function and actually executes nothing. Next, the analysis engine continues to wait for the next instruction. The execution engine synchronizes the information of instruction 23. The context information at this time contains the impact of instruction 22 on the program after execution. Specifically, the standard output stream prints variable c. After receiving it, the analysis engine prints variable c to the standard output stream. At this point, the call of the external function printf is completed, and the states of the execution engine and the analysis engine are consistent, that is, the state of normal program execution.
[0021] Specifically, S3 includes: using IDApro to disassemble the hardened binary file (called old code or old file) to generate an assembly file containing all the instructions of the target program; modifying the assembly file and replacing the direct jump and indirect jump instructions of the anti-obfuscation function with no operation, where the anti-obfuscation function refers to the function added by the anti-obfuscation technology, and the anti-obfuscation code is viewed at the function level. In practical applications, the anti-obfuscation code is synonymous with the redundant code, and the anti-obfuscation function is synonymous with the redundant function. No operation refers to the NOP instruction. When the computer processor encounters the NOP instruction, it will not perform any actual operation and continue to execute subsequent instructions.
[0022] Generate an object file according to the modified assembly file, and then use the generated object file as a new code segment of the old binary file. At the same time, change the permission of the old code segment of the old binary file to be readable but not executable. Among them, the new assembly file can be compiled to generate an object file (.o file). Taking the object file as a new code segment of the old binary file and changing the permission of the old code segment to be readable but not executable to support potential data reading, the cleaned binary file is obtained. Among them, the old code refers to the code before modification, and the new code refers to the code after modification. Specifically in this embodiment, the old binary file refers to the target binary file before modification in this step. For user-defined normal functions, their instruction semantics are the same, but after being modified in this step, their instruction addresses are no longer the same, so an address mapping is created. The mapping here is specifically a table. For example, the instruction A at address 0x32 in the old code corresponds to instruction A' in the new code, with an address of 0x80, and the created address mapping is 0x32 -> 0x80.
[0023] During the process of modifying the assembly file, it includes: for direct jumps, symbolize the basic blocks and replace all direct jumps of the basic blocks with labels; for indirect jumps, create an address mapping between the old code and the new code, and map the old code address to the new address before the indirect jump. For example: modify this assembly file to obtain a new assembly file. The specific method is that for direct jumps, symbolize the basic blocks and replace all direct jumps of the basic blocks with labels. For example, use the label L_0x24 to refer to the basic block with a starting address of 0x24. For indirect jumps, create an address mapping between the old code and the new code, and map the old code address to the new address before each indirect jump. Finally, replace all instructions that are anti-fuzzing functions for direct and indirect jumps with no-operations to achieve the effect of removing anti-fuzzing code.
[0024] Furthermore, it also includes: evaluating the removal effect of the anti-fuzzing code and recording the evaluation results. For example: using the LAVA-M dataset and the utilities in Binutils as the target program, adopting AFL and QSym in QEMU mode as fuzzers, and the anti-fuzzing test methods being AntiFuzz and Fuzzification.
[0025] In a second aspect, the anti-fuzzing code recognition device based on dynamic taint analysis provided by the embodiments of the present invention includes:
[0026] An analysis module for inputting the binary program code fortified by the anti-fuzzing test mechanism into the binary dynamic analysis framework for taint analysis;
[0027] An identification module for constructing a data flow graph using the taint analysis results and identifying anti-fuzzing code according to the data flow graph;
[0028] A code filtering module, configured to rewrite the binary program code and remove the anti-fuzzing code according to the recognition result.
[0029] A testing module, configured to evaluate the removal effect of the anti-fuzzing code and record the evaluation result.
[0030] The anti-fuzzing code recognition method and device based on dynamic taint analysis provided by the embodiments of the present invention utilize a fully automatic, program-level binary simulation execution framework to analyze the fortified binary file, breaking through the limitation that existing binary simulation execution tools rely on hooking (hook) API calls. It can identify the input of the target program and determine the set of key functions in the program through the given test cases. Then, through taint analysis technology, inter-procedural data flow analysis is performed on the parameters and return values of the key functions, thereby constructing a program-level data flow tree. Based on in-depth analysis of the Anti-Fuzzing mechanism, the characteristics of the redundant code data flow can be determined, and these redundant codes can be cleaned up through binary rewriting. Thus, without changing the normal function of the program, the fortified binary program can be analyzed, and the redundant codes added by the Anti-fuzzing mechanism can be automatically removed. Description of the Drawings
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0032] Figure 1 It is a schematic flowchart of the method implemented by the present invention.
[0033] Figure 2 It is a schematic diagram of the binary dynamic analysis framework of the embodiments of the present invention.
[0034] Figure 3 It is a schematic diagram of the data flow of a specific example of the embodiments of the present invention.
[0035] Figure 4 、 5 It is a schematic diagram of a specific example provided by the embodiments of the present invention. Detailed Embodiments
[0036] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any and all combinations of one or more of the associated listed items. Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless defined as here.
[0037] The following further elaborates on this embodiment in conjunction with the accompanying drawings of the specification. As Figure 1 shown, this embodiment provides an anti-obfuscation code recognition method based on dynamic taint analysis, and the method includes the following steps:
[0038] Step 1: Design and implement a binary dynamic analysis framework, which is used to perform taint analysis on binary programs.
[0039] Step 2: Perform dynamic taint analysis on the binary program strengthened by the anti-obfuscation test mechanism, and construct a data flow graph at the function level. Traverse the data flow graph, and distinguish the redundant data flow containing anti-obfuscation code from the normal data flow containing the original function code of the program through the significant feature differences between the anti-obfuscation code and the original function code of the program.
[0040] Step 3: Rewrite the binary strengthened by the anti-obfuscation test mechanism to truly remove the anti-obfuscation code in the strengthened binary program.
[0041] Step 4: Experimentally verify and evaluate the effectiveness and safety of this embodiment.
[0042] In this embodiment, the following technical solutions are specifically adopted:
[0043] Step 1: Design and implement a binary dynamic analysis framework, as Figure 2 shown. This framework consists of two parts: an execution engine and an analysis engine. In this embodiment, it is used to perform taint analysis on binary programs. In the execution engine, dynamic instrumentation at the instruction level is performed on the binary program to dynamically intercept the instruction stream of the target program at runtime. While the instrumentation is being executed, the runtime environment of the program is maintained. When system calls and API calls occur, the target function can be found, parameters can be passed, the call can be executed, and the result can be returned. Before each instruction is executed, the instruction context information is synchronized to the analysis engine, and the analysis engine re-executes the instruction in the known context to update the analysis result in real time. Specifically, the execution engine directly uses the tool QBDI; the specific execution steps in the analysis engine use the tool Triton, and the taint propagation strategy can be customized. In this embodiment, synchronization between the execution engine and the analysis engine is performed.
[0044] Step 2: Perform dynamic taint analysis on the binary program after strengthening the anti-fuzz testing mechanism, and construct a data flow graph at the function level. Traverse the data flow graph, extract the data flow starting from the main function as the starting node, regard the data flow that terminates at ordinary functions as the redundant data flow belonging to the anti-fuzz code, and regard the data flow that finally flows to the system call as the normal data flow belonging to the original functional code of the program. The function code contained in the redundant data flow is the anti-fuzz code.
[0045] Step 2 of this embodiment includes the following steps:
[0046] Step 201: Identify the input of the target program, including the command-line parameters of the main function and the API (Application Programing Interface) function calls related to the input.
[0047] Step 202: Mark the input of the target program as the taint source, change different inputs (including legal inputs and illegal inputs), perform dynamic taint analysis on the target program, record the functions to which the taint propagates, and regard these functions as the key functions of the target program.
[0048] Step 203: Mark the parameters and return values of the key functions as taint sources respectively, and based on the taint analysis technology, implement interprocedural data flow tracking to construct a data flow graph at the function level. Using the main function as the root node and the key functions called by the main function as the second-layer nodes, display the data flow graph as a tree structure, and at the same time repeat the nodes in the circular structure in the graph in the tree.
[0049] Step 204: Traverse all the paths starting from the root node in the data flow graph. Each path is regarded as a data flow. All the data flows can be divided into three types. As Figure 3 shown, data flow type I is the redundant code introduced by the anti-fuzz testing mechanism, and data flows II and III are the original code of the program. The function of the program is mainly determined by the libraries and system calls it invokes. Therefore, the data flows belonging to the original code will ultimately flow to system calls. Some directly flow to system calls and the data flow terminates (data flow III). Some will first flow to some ordinary functions, and through the analysis of the propagation range of return values (data flows ③ and ④), the data flow will ultimately also terminate at system calls (data flow II). The redundant code introduced by the anti-fuzz testing mechanism does not contain any functions, and the data flow will ultimately terminate at ordinary functions.
[0050] Step 3: Rewrite the binary after strengthening the anti-fuzz testing mechanism to truly remove the anti-fuzz code in the strengthened binary program.
[0051] Step 3 of this embodiment includes the following steps:
[0052] Step 301: Use IDApro to disassemble the strengthened binary file (referred to as the old code or old file) to generate an assembly file containing all the instructions of the target program.
[0053] Step 302: Modify this assembly file to obtain a new assembly file. Specifically, for direct jumps, symbolize the basic blocks and replace all direct jumps of the basic blocks with labels. For example, use the label L_0x24 to refer to the basic block with the starting address of 0x24. For indirect jumps, create an address mapping between the old code and the new code, and map the old code address to the new address before each indirect jump. Finally, replace all the instructions with direct and indirect jumps to anti-fuzz functions with no-operations to achieve the effect of removing the anti-fuzz code.
[0054] Step 303: Compile the new assembly file to generate an object file (.o file).
[0055] Step 304: Use the object file as a new code segment of the old binary file, change the permission of the old code segment to be readable but not executable to support potential data reading, and then the cleaned binary file is obtained.
[0056] Step 4: Use the LAVA-M (Large Scale Automated Vulnerability Addition-Methodology) dataset and four utilities (objdump, nm, size, and strings) in Binutils as the target program to test the effectiveness and security of this embodiment.
[0057] Step 4 of this embodiment includes the following steps:
[0058] Step 401: Experimental setup: Use the LAVA-M (Large Scale Automated Vulnerability Addition - Methodology) dataset and four utilities (objdump, nm, size, and strings) in Binutils as the target programs; use two fuzzers, AFL (referred to as AFL-QEMU) in QEMU mode and QSym; select two state-of-the-art anti-fuzz testing methods, AntiFuzz and Fuzzification.
[0059] Step 402: Compile the four programs in LAVA-M into three versions, namely the original, AntiFuzz-protected, and Fuzzification-protected versions. For the fortified binary files protected by AntiFuzz and Fuzzification, use this embodiment for cleaning to obtain the cleaned binary files. Use the above five binary files as the target programs and conduct 24-hour fuzz testing with the two fuzzers, AFL-QEMU and QSym. Count the number of unique crashes finally discovered by the fuzzers to evaluate the effectiveness of this embodiment. The experimental results are the statistics of the number of unique crashes in Step 402 shown in Table 1. Both anti-fuzz testing mechanisms significantly reduce the ability of the fuzzer to discover false positives. For the fortified binary programs, the number of unique crashes that the fuzzer can discover is greatly reduced. After applying this embodiment to clean these two fortified binary programs, the number of unique crashes that the fuzzer can discover returns to near the original level, which demonstrates the effectiveness of this embodiment.
[0060] Table 1
[0061]
[0062] Step 403: Compile the four utilities (objdump, nm, size, and strings) in Binutils into three versions, namely the original, AntiFuzz-protected, and Fuzzification-protected versions. For the fortified binary files protected by AntiFuzz and Fuzzification, use this embodiment for cleaning to obtain the cleaned binary files. Use the above five binary files as the target programs and conduct 24-hour fuzz testing with AFL-QEMU. Count the branch coverage rate to evaluate the effectiveness of this embodiment.
[0063] The step 403 branch coverage statistics results shown in Table 2 of the experimental results indicate that AntiFuzz and Fuzzification did significantly reduce the branch coverage of the fuzzer, while this embodiment restored the branch coverage that the fuzzer could discover to a level close to that of the original program, with an average maximum of up to 96.25%, which demonstrates the effectiveness of this embodiment.
[0064] Table 2
[0065]
[0066] Step 404: For the four utilities (objdump, nm, size, and strings) in Binutils, use the DejaGnu test framework officially provided by GNU Binutils to test the binary files in five implementation methods, namely the original binary file, the binary file protected by AntiFuzz, the binary file protected by Fuzzification, the binary file with AntiFuzz cleared, and the binary file with Fuzzification cleared, and record whether compilation reports errors and whether the functional tests pass to evaluate the security of this embodiment. The experimental results show that all implementation methods passed compilation and passed the official functional tests, which demonstrates the security of this embodiment.
[0067] The anti-fuzzing code recognition method and device based on dynamic taint analysis provided by the embodiments of the present invention utilize a fully automatic, program-level binary simulation execution framework to analyze the fortified binary file, breaking through the limitation that existing binary simulation execution tools rely on hooking API calls. It can identify the input of the target program and determine the set of key functions in the program through the given test cases. Then, through taint analysis technology, interprocedural data flow analysis is performed on the parameters and return values of the key functions to construct a program-level data flow tree. Based on in-depth analysis of the Anti-Fuzzing mechanism, the characteristics of the redundant code data flow can be determined, and these redundant codes can be cleared through binary rewriting. Thus, without changing the normal function of the program, the fortified binary program can be analyzed, and the redundant codes added by the Anti-fuzzing mechanism can be automatically removed.
[0068] In practical applications, the present invention can effectively remove the redundant codes added to the program by existing anti-fuzzing testing technologies, and can restore the error-finding ability and coverage feedback mechanism of the fuzzer. And after rewriting the target program, the program still remains correct.
[0069] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the corresponding descriptions in the method embodiments. As mentioned above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An anti-obfuscation code recognition method based on dynamic taint analysis, characterized in that, Including: S1. Input the binary program code fortified by the anti-fuzz testing mechanism into the binary dynamic analysis framework for taint analysis; S2. Construct a data flow graph using the taint analysis results, and identify anti-fuzz code based on the data flow graph; S3. Rewrite the binary program code according to the identification result and remove the anti-fuzz code.
2. The method according to claim 1, wherein The binary dynamic analysis framework includes: an execution engine and an analysis engine; In the execution engine, perform instruction-level dynamic instrumentation on the binary program code, and dynamically intercept the instruction stream of the target program during runtime; Before each instruction is executed, the execution engine synchronizes the context information of the instruction to the analysis engine. The analysis engine that has completed the synchronization executes the instruction again based on the context information of the instruction, and updates the taint analysis result in real time.
3. The method according to claim 1, wherein In S2, it includes: Traverse the data flow graph to extract the data flow starting from the main function as the starting node, and regard the data flow that terminates at a normal function as the redundant data flow to which the anti-fuzz code belongs. The function code included in the redundant data flow is used as the anti-fuzz code.
4. The method according to claim 3, characterized in that, In S2, it includes: Identify the input of the target program and the API (Application Programming Interface) function calls related to the input of the target program. The input of the target program includes: the command-line parameters of the main function; Mark the input of the target program as a taint source, change different inputs and record the functions to which the taint propagates. The functions to which the taint propagates are recorded as the key functions of the target program; Then mark the parameters and return values of the key functions as taint sources, and construct a function-level data flow graph. Among them, with the main function as the root node and the key functions called by the main function as the second-layer nodes, a tree-like data flow graph is established; Traverse all paths starting from the root node in the data flow graph, and each path starting from the root node is used as a data flow.
5. The method according to claim 1, wherein In S3, it includes: Generate an assembly file containing all the instructions of the target program; Modify the assembly file, and replace the instructions of direct jumps and indirect jumps to anti-fuzz functions with no-operations; Generate an object file according to the modified assembly file, and then use the generated object file as a new code segment of the old binary file. At the same time, change the permission of the old code segment of the old binary file to readable but not executable.
6. The method according to claim 5, characterized in that, During the process of modifying the assembly file, it includes: For direct jumps, symbolize the basic blocks and replace all direct jumps of the basic blocks with labels; For indirect jumps, create an address mapping before the old code and the new code, and map the old code address to the new address before the indirect jump, where the old code includes the binary program code that has not been rewritten before executing S3.
7. The method according to claim 1, wherein It also includes: Evaluate the removal effect of the anti-fuzz code and record the evaluation result.
8. The method according to claim 7, wherein It also includes: Use the LAVA-M dataset and the utilities in Binutils as the target program, use AFL and QSym in QEMU mode as fuzzers, and the anti-fuzz testing methods are AntiFuzz and Fuzzification.
9. An anti-obfuscation code recognition device based on dynamic taint analysis, characterized in that Including: An analysis module, configured to input the binary program code strengthened by the anti-fuzz testing mechanism into the binary dynamic analysis framework for taint analysis; An identification module, configured to construct a data flow graph by using the taint analysis result and identify anti-fuzz code according to the data flow graph; A code filtering module, configured to rewrite the binary program code and remove the anti-fuzz code according to the identification result.
10. The method according to claim 9, wherein Further comprising: A testing module, configured to evaluate the removal effect of the anti-fuzz code and record the evaluation result.