Semantic hiding method based on data-driven code reuse
Through data-driven code reuse technology, search and encapsulate reusable Gadgets in the host program, and flatten the control flow, solving the problem of low fusion of steganographic code in software copyright protection, achieving efficient steganographic semantic hiding, and enhancing the ability to resist reverse analysis.
Patent Information
- Application Number
- CN202510383433.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-11
AI Technical Summary
In the existing software copyright protection technology, software obfuscation and watermarking methods are not robust enough, and the steganographic code and host program are low, which are easily identified and cracked by attackers, resulting in poor copyright protection effect.
The semantic hiding method based on data-driven code multiplexing is adopted to parse the hidden semantics through predefined semantic expressions, determine the reusable Gadget type, search and encapsulate code snippets in the host program, and perform control flow flattening processing, and compile it into a binary file to realize steganography.
It improves the anti-reverse analysis ability of software copyright protection, enhances the depth hiding of steganographic semantics, and is difficult to be located and identified by attackers, and improves the level of software copyright protection.
Smart Images

Figure CN120296705A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of program steganography, and particularly to a semantic hiding method based on data-driven code reuse. Background Art
[0002] Currently, software applications have become basic tools that humans rely on heavily in their lives. While providing convenience to humans, they are constantly subject to reverse attacks. For example, attackers can crack the program execution logic through reverse analysis of the software, and perform operations such as code cloning and repackaging. This not only seriously infringes on the copyright and economic interests of the original creators, but also poses a security risk of backdoor attacks. Therefore, the importance of software copyright protection is self-evident.
[0003] The main idea of software copyright protection is to embed copyright information into the software by means of program steganography. Common methods include software obfuscation technology and software watermark technology. Among them, software obfuscation technology embeds steganographic code into the program through semantic equivalent transformation of the program. The transformed code is difficult for attackers to understand, thus achieving software copyright protection. Common obfuscation technologies include layout obfuscation, control flow obfuscation, and data obfuscation. Layout obfuscation only changes the code structure and format, with weak robustness; control flow obfuscation and data obfuscation achieve semantic-level obfuscation, and the anti-reverse analysis ability is enhanced to some extent. However, in the actual software analysis process, it largely depends on the attacker not knowing which obfuscation method the host program uses. Assuming the attacker knows the specific obfuscation method and details, its protection effect will be reduced to varying degrees. Software watermark technology embeds watermark information into the host program and can extract this information when needed to prove software copyright. The watermark carrier can be data, code, or execution process characteristics. However, in the current implementation of most software watermark technologies, there is an obvious lack of integration between the embedded and extracted code and the host program code. This technical implementation method results in weak concealment of watermark information, which is easily located, identified, and damaged by attackers, thus reducing the actual effect of software copyright protection. Summary of the Invention
[0004] Based on this, it is necessary to provide a semantic hiding method based on data-driven code reuse for the above technical problems, which can improve the level of software copyright protection.
[0005] The present invention adopts the following technical solutions:
[0006] The present invention provides a semantic hiding method based on data-driven code reuse, including:
[0007] Semantically declare the semantic to be hidden through a predefined semantic expression to obtain a steganographic semantic;
[0008] Parse the steganographic semantics to determine the reusable Gadget types in the steganographic semantics; a Gadget is a code snippet for implementing basic operations.
[0009] Search for a set of reusable Gadgets in the host program according to the reusable Gadget types.
[0010] Obtain the semantic execution area of each code snippet in the set of reusable Gadgets in the host program.
[0011] Enclose the code of each semantic execution area in a control flow branch and perform control flow flattening on the enclosed code to obtain intermediate representation code.
[0012] Compile the intermediate representation code into a binary file to implement the steganography of the semantics to be hidden.
[0013] Optionally, the predefined semantic expressions include keyword design and expression design; keyword design includes declaring the meanings of multiple keywords; expression design includes defining basic operation types and corresponding formal descriptions; basic operation types include declaration, assignment, non-dereferencing operation, dereferencing operation, and output.
[0014] Optionally, searching for a set of reusable Gadgets in the host program according to the reusable Gadget types includes:
[0015] Perform static search on the intermediate code of the host program's compiler to determine the set of reusable Gadgets corresponding to the reusable Gadget types in the host program.
[0016] Optionally, the static search includes function traversal search; performing static search on the intermediate code of the host program's compiler to determine the set of reusable Gadgets corresponding to the reusable Gadget types in the host program includes:
[0017] Traverse all functions in the intermediate code of the compiler. If the current function is the target function or a sub-function of the target function, analyze each instruction in the current function; the target function is the function where the vulnerability is located.
[0018] If the instruction is a system call instruction, obtain the parameters of the instruction. If the parameters of the instruction are controllable parameters, then form a Gadget with the instruction and its context instructions.
[0019] If the instruction is not a system call instruction, obtain the source operand and destination operand of the instruction. If the source operand and destination operand of the instruction are global variables or controllable parameters, then form a Gadget with the instruction and its context instructions.
[0020] According to the reusable Gadget types in the steganographic semantics, match the corresponding type of Gadgets in the host program, and form a set of reusable Gadgets with all the matched Gadgets.
[0021] Optionally, the method further includes:
[0022] When there is no set of reusable Gadgets in the host program, obtain the compiler intermediate code of the host program and the Gadgets to be inserted;
[0023] If the type of the Gadget to be inserted is a global variable, directly insert the Gadget to be inserted into the compiler intermediate code;
[0024] If the type of the Gadget to be inserted is a local variable, an assignment type, or an arithmetic type, obtain the position information to be inserted, and insert the Gadget to be inserted into the compiler intermediate code according to the position information;
[0025] If the type of the Gadget to be inserted is a conditional type, obtain the position information to be inserted and the basic block information, and insert the Gadget to be inserted into the compiler intermediate code according to the position information and the basic block information.
[0026] Optionally, obtain the semantic execution area of each code segment in the set of reusable Gadgets in the host program, including:
[0027] In the intermediate code space of the host program, find the semantic safe area of each code segment in the set of reusable Gadgets; the semantic safe area represents the code area where the memory context is not affected when the corresponding code segment is executed in the host program;
[0028] According to the backbone instruction, obtain the semantic execution area from the semantic safe area; the backbone instruction is an instruction that will be executed only once in the module or function of the host program.
[0029] Optionally, find the semantic safe area of each code segment in the set of reusable Gadgets in the intermediate code space of the host program, including:
[0030] For each code segment, obtain the instruction set of the intermediate code space;
[0031] Obtain the instruction set that modifies the memory space corresponding to each variable in the code segment from the instruction set;
[0032] According to the predecessor instruction and the successor instruction of the code segment on the instruction set, determine the subset between the predecessor instruction and the successor instruction;
[0033] Take the intersection of the subsets obtained for each variable to get the semantic safe area of the code segment.
[0034] Optionally, the Gadget set includes multiple Gadgets; the method further includes:
[0035] Before obtaining the execution area of each code snippet in the reusable Gadget set from the host program, adopt the strategy of function inlining to fuse the functions where the Gadgets are located in the compiler intermediate code into one function, obtaining the fused compiler intermediate code.
[0036] Optionally, fusing the functions where the Gadgets are located in the compiler intermediate code into one function to obtain the fused compiler intermediate code includes:
[0037] Obtain the variable set of the target function in the compiler intermediate code;
[0038] Traverse all the sub-functions in the target function. For any sub-function, perform an initialization operation on the parameter list of the sub-function and add the local variables of the sub-function to the target function;
[0039] Obtain the start position and end position of the code of the sub-function, and fuse the sub-function into the target function according to the start position and end position of the code;
[0040] Determine whether the sub-function has a return value. If the sub-function has a return value, perform an operation to update the return value and update the external call relationship information of the sub-function;
[0041] After processing all the sub-functions, return the fused compiler intermediate code.
[0042] Optionally, the method further includes:
[0043] When performing control flow flattening on the encapsulated code, set branch predicates after each branch of the encapsulated code to achieve the stitching of steganographic semantics.
[0044] The present invention provides a semantic hiding device based on data-driven code reuse, including:
[0045] A declaration module, configured to perform semantic declaration on the semantics to be hidden through a predefined semantic expression to obtain steganographic semantics;
[0046] An analysis module, configured to analyze the steganographic semantics to determine the reusable Gadget types in the steganographic semantics; a Gadget is a code snippet for implementing basic operations;
[0047] A search module, configured to search for a reusable Gadget set in the host program according to the reusable Gadget types, and obtain the semantic execution area of each code snippet in the reusable Gadget set in the host program;
[0048] A processing module, configured to encapsulate the code of each semantic execution area in a control flow branch, and perform control flow flattening on the encapsulated code to obtain intermediate representation code;
[0049] A compilation module, configured to compile the intermediate representation code into a binary file to implement steganography of the semantic to be hidden.
[0050] The present invention provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-mentioned semantic hiding method based on data-driven code reuse.
[0051] The present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the above-mentioned semantic hiding method based on data-driven code reuse.
[0052] The above at least one technical solution adopted by the present invention can achieve the following beneficial effects:
[0053] In the present invention, first, the semantic to be hidden is expressed in a specific manner (pre-defined semantic expression) to obtain a steganographic semantic, and the steganographic semantic is parsed to find out the code fragments that implement basic operations, that is, Gadgets. These Gadgets are the basic components of the program. By finding reusable Gadgets, the steganographic semantic can be divided into multiple functional fragments for subsequent operations; searching for a set of reusable Gadgets in the host program is to find code fragments that already exist in the host program and can be reused. At the same time, obtaining the execution area of each code fragment means determining the position and context of these code fragments in the program execution flow; then, the code of each semantic execution area is encapsulated in a control flow branch to make it look more complex and disordered. The control flow flattening further confuses the control flow of the host program. On the premise of ensuring the equivalent transformation of the native semantics of the host program, the original clear structure is eliminated, and at the same time, the steganographic semantic is cleverly hidden therein. Finally, the intermediate representation code is compiled into a binary file. In this way, the processed code is converted into machine code, making it difficult to discover the hidden semantic by means of static analysis or dynamic analysis, etc., achieving deep hiding of the steganographic semantic, realizing effective resistance to dynamic reverse analysis, reducing the characteristics of the embedded code, making it difficult for attackers to locate and identify, and greatly improving the software copyright protection level. Description of the Drawings
[0054] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0055] Figure 1 Schematic flow chart of a semantic hiding method based on data-driven code reuse provided by the present invention;
[0056] Figure 2 Overall design flow chart of a semantic hiding method based on data-driven code reuse provided by the present invention;
[0057] Figure 3 Schematic diagram of function fusion, semantic security area, semantic execution area, control flow flattening, data-driven control flow flattening and branch predicate design provided by the present invention;
[0058] Figure 4 Control flow chart of the original program IR code provided by the present invention;
[0059] Figure 5 Control flow chart of the program after the algorithm is split by the semantic execution area provided by the present invention;
[0060] Figure 6 Structure diagram after flattening the control flow chart after splitting the semantic execution area provided by the present invention;
[0061] Figure 7 Structure diagram after flattening based on the original program control flow graph provided by the present invention;
[0062] Figure 8 Schematic diagram of a semantic hiding device based on data-driven code reuse provided by the present invention;
[0063] Figure 9 Schematic diagram of a computer device for implementing a semantic hiding method based on data-driven code reuse provided by the present invention. Detailed implementation manners
[0064] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0065] In view of problems such as the difficulty of reverse analysis of software depending on whether the obfuscation strategy is transparent and the low degree of integration of steganographic code with the host program, the present invention designs and implements DopSteg, which is a novel program steganography technique. The reverse analysis of it additionally adds the analysis of the steganographic code logic, thereby reducing the degree of dependence on whether the obfuscation strategy is transparent and enhancing the anti-dynamic analysis effect; at the same time, DopSteg uses the method of data-driven code reuse to implement program steganography, reusing as much code in the host program as possible, improving the code integration degree with the host program, and making the steganographic code more concealed. After analysis, DopSteg can achieve steganography with Turing-complete semantics. The present invention is the first attempt to apply the Data-Oriented Programming (DOP) technology in the field of software protection, expanding the application scope of this technology.
[0066] The present invention provides a semantic hiding method based on data-driven code reuse, which specifically includes:
[0067] For the first time, the DOP technology is introduced into the field of software protection, a complete set of program steganography models for data-oriented programming is proposed, and the effectiveness of this method is verified in cases.
[0068] A new strategy for flattening and dividing control flow branches is proposed. This strategy flattens branches based on the data execution area, which can ensure the consistency of data during steganographic semantic stitching.
[0069] A language for expressing hidden semantics (Hidden Semantic Representation Language, HSRL) is defined, which can accurately express complex hidden semantics.
[0070] Branch predicates are designed, which not only ensure the equivalent transformation of program semantics but also can correctly stitch hidden semantics through data-driven code reuse.
[0071] A function fusion algorithm is designed and implemented, effectively solving the problem of Gadgets stitching between functions.
[0072] First, elaborate on the related work before designing the method of the present invention. First, introduce the related work on using software obfuscation technology to implement program steganography, then introduce the related research work on software watermarking, an application technology of program steganography. Subsequently, introduce the principles of RopSteg and DOP technologies, which are most relevant to the method of the present invention. Finally, elaborate on the differences and advantages between the above technologies and the DopSteg technology.
[0073] Software Obfuscation: Software obfuscation techniques can be used to achieve program steganography. The idea is to hide the code to be steganographed while obfuscating the host program through obfuscation techniques. Specifically, it includes three obfuscation techniques. One is to make it difficult to disassemble machine code, and only a small part of the instructions can be disassembled. The second is to take advantage of the defect that x86 architecture instructions do not have length alignment. Through the overlapping instruction technique, the linear disassembler disassembles at the wrong starting point, resulting in the hiding of the real instructions. However, this method cannot escape recursive disassemblers and dynamic analysis. The third is to encrypt certain instructions to hide them without a key. Although encryption provides very strong protection, it always leaves doubts and attracts attention in program analysis, and the concealment is not strong. Using software obfuscation techniques to achieve program steganography cannot effectively counter dynamic reverse analysis.
[0074] Software Watermarking: Software watermarking technology extracts the watermark information pre-embedded in the carrier program through a specific watermark extraction algorithm to prove the ownership of software copyright. According to the form of watermark information display, it can be divided into data watermark and code watermark. Static data watermark embeds the watermark information into the static data area, and dynamic data watermark hides the watermark information in data structures such as the heap and linked list dynamically generated by the carrier program; static code watermark embeds the watermark information into a certain code segment, while dynamic code watermark encodes the state information of the program's dynamic execution as watermark information. Program steganography is very similar to software watermarking, but software watermarking usually hides a watermark information, such as copyright information. While program steganography focuses more on hiding additional code semantics.
[0075] RopSteg: RopSteg achieves program steganography based on the Return-Oriented Programming (ROP) technique. It constructs a ROP chain in the carrier program that is not used for attacks but implements the program steganography function, which is the first friendly application of the ROP technique. The ROP chain with steganography function constructed by RopSteg is invisible during static analysis. Only when the program runs, multiple ROP Gadgets scattered in different positions in the carrier program are dynamically concatenated to form hidden semantics, which is a typical code reuse technique. However, RopSteg has an obvious drawback, that is, when constructing the ROP hidden semantics chain, it violates the program execution control flow, and the behavioral characteristics during dynamic execution are relatively obvious, and it is easily detected by the ROP defense mechanism.
[0076] DOP: The DOP technique originated from the idea of non-control-flow data attacks. It is an attack launched at the data level without changing the dynamic execution control flow during the attack process, which can bypass mainstream defense mechanisms such as Data Execute Prevention (DEP), Canary, Address Space Layout Randomization (ASLR), and Control Flow Integrity (CFI), and the attack has strong concealment. The DOP technique was initially proposed as an attack technique. Its attack process is to input carefully designed Payload data (Payload can represent the key code or data to be steganographed in the program) to the program, reuse and reorganize the host program Gadgets to achieve an attack semantics different from the original semantic logic of the host program. If this attack semantics is understood as a hidden semantics, the DOP technique can achieve program steganography, which is also the original intention of this invention to introduce the DOP technique to achieve program steganography.
[0077] The DopSteg technique proposed in this invention can also be regarded as a new software obfuscation method because, in terms of technical implementation, it uses the classic control flow flattening method. However, the difference is that this invention newly designs a branch division strategy based on the data execution area, and most software obfuscation methods focus on the obfuscation of the code logic of the carrier program itself, while the method of this invention pays more attention to the hiding of steganographic semantics. This purpose is the same as that of software watermarking technology. The ultimate goal of software watermarking is to hide the watermark information in the carrier program so that it cannot be discovered by attackers. If the watermark information is in the form of code, it is consistent with the goal of this invention. However, most code watermarking algorithms are complex, and the implemented code needs to be inserted additionally, with a low coupling degree with the carrier program code. The RopStep technique significantly improves the coupling degree. It constructs a hidden semantics ROP chain through code reuse technology. Although it does not exist in static analysis, it violates the program's dynamic execution control flow and is easily detected by the ROP defense mechanism, with a large application limitation. The DopSteg technique proposed in this invention optimizes the branch division strategy, realizes program steganography of data-driven code reuse, has a high coupling degree with the carrier program code, simultaneously increases the instruction entropy by about 140% on average, enhances the anti-dynamic reverse analysis ability, and strictly follows the program's dynamic execution control flow, with a wider application range.
[0078] The following will detail the technical solutions provided by each embodiment of this invention with reference to the accompanying drawings.
[0079] Figure 1 It is a schematic flowchart of a semantic hiding method based on data-driven code reuse in this invention, specifically including the following steps:
[0080] S101. Semantically declare the semantic to be hidden through a predefined semantic expression to obtain a steganographic semantic.
[0081] Among them, the predefined semantic expression includes keyword design and expression design; keyword design includes declaring the meanings of multiple keywords; expression design includes defining basic operation types and corresponding formal descriptions; the basic operation types include declaration, assignment, non-dereference operation, dereference operation, and output.
[0082] The predefined semantic expression is the Hidden Semantic Representation Language (HSRL) for expressing the hidden semantic. The HSRL language is used to describe the steganographed semantic, and its keyword design is shown in Table 1.
[0083] Table 1
[0084] Keyword Meaning alloca Variable declaration i8 Declare 8-bit operation i16 Declare 16-bit operation i32 Declare 32-bit operation i64 Declare 64-bit operation assign Variable assignment add Addition operation sub Subtraction … Other non-dereferencing operations deref Dereferencing operation output Variable output
[0085] In Table 1, the alloca keyword declares a variable. i8, i16, i32, and i64 define basic variable types, representing 8-bit, 16-bit, 32-bit, and 64-bit types respectively. The assign keyword represents an assignment operation, keywords such as add and sub represent arithmetic operations, the deref keyword represents a variable dereference operation, and the output keyword represents an output operation. In the grammar definition of the HSRL language, var is used to represent all identifiers, and its naming rule is the same as that of the C language; con is used to represent all 16-bit numbers. Taking 64-bit integers as an example, its value range is from 0x0000000000000000 to 0xFFFFFFFFFFFFFFFF.
[0086] Expression design:
[0087] The HSRL language defines five basic operation types, namely declaration, assignment, non-dereference operation, dereference operation, and output. Next are their formal descriptions.
[0088] (1) Variable declaration. Define the number and size of independent and operable memory spaces required for the steganographic semantic, and it can be formally described using the Alloca statement, as shown in formula (1) specifically
[0089] Alloca := alloca i64 var | alloca i32 var | alloca i16 var | alloca i8 var (1)
[0090] (2) Variable assignment. Define which values in the steganographic semantics' operable memory spaces need to be modified by special means. It can be described using the Assign statement, specifically as shown in formula (2). The description is as follows:
[0091] Assign: = assign var (2)
[0092] (3) Variable operation. Define the arithmetic operations of steganographic semantic variables. Common operations include add, sub, mul, div, rem, xor, etc. Here, taking addition as an example, the HSRL language uses the Calculate statement for formal description, as shown in formula (3):
[0093] Calculate: = add var var | add var var con | add var con con (3)
[0094] (4) Variable dereference operation. Define the dereference behavior of pointers in steganographic semantics, and use the Deref statement for formal description, as shown in formula (4).
[0095] Deref: = deref var var (4)
[0096] (5) Variable output. Define which variables in steganographic semantics need to participate in user interaction. It is described using the Output statement, as shown in formula (5):
[0097] Output: = output var (5)
[0098] Combining the above formulas, the overall definition of the HSRL language can be formally described as follows, as shown in formulas (6) and (7):
[0099] Other: = Assign | Calculate | Deref | Output | Other (6)
[0100] HSRL: = Alloca HSRL | Alloca Other (7)
[0101] To define the HSRL language, variables must be declared first, and then statements are declared. The HSRL language can ensure the correct representation of steganographic semantics.
[0102] S102. Parse the steganographic semantics to determine the reusable Gadget types in the steganographic semantics; Gadget is a code snippet for implementing basic operations.
[0103] The steganographic semantics can be input into a pre-constructed parser to obtain the reusable Gadget types in the steganographic semantics; the process of constructing the parser includes: writing the pre-defined semantic expression logic into the basic parser to obtain the constructed parser.
[0104] S103. According to the reusable Gadget types, search for a set of reusable Gadgets in the host program, and obtain the semantic execution areas of each code segment in the set of reusable Gadgets in the host program.
[0105] Optionally, searching for a set of reusable Gadgets in the host program according to the reusable Gadget types includes: performing a static search on the compiler intermediate code of the host program to determine the set of reusable Gadgets corresponding to the reusable Gadget types in the host program.
[0106] The purpose of Gadget search is to find reusable code segments in the host program that can represent steganographic semantics. The Gadget search of the present invention is a static search on the intermediate representation (IR) code of the underlying virtual machine (Low Level Virtual Machine, LLVM) of the host program.
[0107] Optionally, the static search includes function traversal search; performing a static search on the compiler intermediate code of the host program to determine the set of reusable Gadgets corresponding to the reusable Gadget types in the host program includes: traversing all functions in the compiler intermediate code, and if the current function is the target function or a sub-function of the target function, analyzing each instruction in the current function; the target function is the function where the vulnerability is located; if the instruction is a system call instruction, obtaining the parameters of the instruction, and if the parameters of the instruction are controllable parameters, forming a Gadget by combining the instruction and its context instructions; if the instruction is not a system call instruction, obtaining the source operand and destination operand of the instruction, and if the instruction source operand and destination operand are global variables or controllable parameters, forming a Gadget by combining the instruction and its context instructions; matching the corresponding type of Gadget in the host program according to the reusable Gadget types in the steganographic semantics, and forming a set of reusable Gadgets by combining all the matched Gadgets.
[0108] Specifically, the search is divided into two levels: function traversal and Gadget recognition. Function traversal adopts a combination of module function traversal and call graph-based traversal; Gadget recognition takes the store instruction as the traceback point and traces backward. If the source operand and destination operand of the instruction are global variables or controllable parameters, then the critical arithmetic instruction and the context instructions together form a Gadget. Controllable parameters specifically refer to the parameters for calling sub-functions that can be controlled. Instruction context refers to all instructions related to a variable during the process from the loading of the variable to the storage or call of the variable, which consists of loading instructions, preprocessing instructions, critical arithmetic instructions, and storage or call instructions. A loading instruction is an instruction that fetches a value from memory into a register; a preprocessing instruction is an instruction that performs some logical processing operations on the variable value, specifically including four types: comparison, sign extension, index value extraction, and type conversion; critical arithmetic instructions refer to the operation instructions related to the Gadget definition, such as the addition instruction add, subtraction instruction sub, multiplication instruction mul, etc. Storage or call instructions refer to saving the variable generated by the critical arithmetic instruction or calling a function using the variable generated by the critical arithmetic instruction. The typical format of the instruction context is formally described as follows:
[0109] Temporary variable 1 = load <data type>, <data type> * <address name>;
[0110] Temporary variable 2 = logical processing on temporary variable 1 such as icmp, sext, bitcast, etc.;
[0111] Temporary variable 3 = critical operation related to temporary variable 2 such as add, sub, mul, etc.;
[0112] store <data type> temporary variable 3, <data type> * <address name>.
[0113] The input of the search algorithm is the intermediate M of the host program, the target function targetFunc, and the sensitive system call sysCall; its output is the set of Gadgets. Search idea: Process all functions func in M in a loop. If the current function curFunc is targetFunc or a child function childFunc of the target function, then each instruction in curFunc needs to be analyzed. If the instruction is a sysCall instruction, obtain the parameters of the instruction. If the parameters are controllable parameters, obtain the context instructions and add them to the Gadget. If the instruction is not a sysCall instruction, obtain the original operand src and the destination operand dst of the instruction. If both src and dst are controllable parameters, obtain the context instructions and add them to the Gadget. Finally, according to the reusable Gadget types in the steganography semantics, match the corresponding types of Gadgets in the host program, and form a set of reusable Gadgets with all the matched Gadgets and output it to complete the Gadget search.
[0114] In theory, DopSteg can implement steganography with Turing-complete semantics, but in practice, there may be a scenario where no suitable Gadget is found in the host program. The main reason is that steganography with Turing-complete semantics is implemented by code reuse, and the types and quantities of its Gadgets are highly dependent on the host program. When the host program does not have the required Gadgets, we adopt the method of Gadget insertion to solve it.
[0115] In one embodiment, it specifically includes: when there is no set of reusable Gadgets in the host program, obtain the Gadgets that need to be inserted into the host program; if the type of the Gadget to be inserted is a global variable, directly insert the Gadget to be inserted; if the type of the Gadget to be inserted is a local variable, an assignment type, or an arithmetic type, obtain the position information to be inserted, and according to the position information, insert the Gadget to be inserted into the specified position; if the type of the Gadget to be inserted is a conditional type, obtain the position information to be inserted and the basic block information, and according to the position information and the basic block information, insert the Gadget to be inserted into the specified position.
[0116] The input of the Gadget insertion algorithm is the IR code M of the source code of the host program to be inserted and the type of Gadget to be inserted; the output is the IR code after insertion. The idea of the insertion algorithm: First, it is necessary to read the IR code and the json file of the type information of the Gadget to be inserted, and then perform the insertion operation in a loop. Specifically, according to different types of Gadgets, different insertion strategies need to be adopted. If it is a global variable, it is inserted directly; if it is a local variable, an assignment-type Gadget, or an arithmetic-type Gadget, the location information needs to be obtained before insertion. If it is a conditional-type Gadget, the location and basic block information need to be obtained before insertion. The purpose of Gadget insertion is to prepare a complete Gadget for the follow-up.
[0117] In one embodiment, the Gadget set includes multiple Gadgets; this embodiment includes: before obtaining the execution area of each code segment in the reusable Gadget set from the host program, adopting the strategy of function inlining to fuse the functions where the Gadgets are located in the compiler intermediate code into one function, obtaining the fused compiler intermediate code.
[0118] The purpose of function fusion design is to solve the stitching problem between Gadgets existing in different functions. Ideally: all the required Gadgets exist inside the same function. In fact, most of the time, Gadgets exist in different functions. Facing the above problems, it is necessary to use the strategy of function inlining to fuse multiple functions where the Gadgets are located into a larger function, and then perform control flow flattening.
[0119] Specifically, fusing the functions where the Gadgets are located in the compiler intermediate code into one function to obtain the fused compiler intermediate code includes: obtaining the variable set of the target function in the compiler intermediate code; traversing all the sub-functions in the target function, for any one sub-function, initializing the parameter list of the sub-function and adding the local variables of the sub-function to the target function; obtaining the start position and end position of the code of the sub-function, and fusing the sub-function into the target function according to the start position and end position of the code; judging whether the sub-function has a return value, if the sub-function has a return value, then perform the operation of updating the return value and updating the external call relationship information of the sub-function; after processing all the sub-functions, return the fused compiler intermediate code.
[0120] The input of the fusion algorithm is the IR code M of the host program source code and the target function targetFunc; the output is the fused IR code M'. The algorithm idea is as follows: First, obtain the variable set VariableSet of the targetFunc function in M, loop through all the sub-functions in the targetFunc function, and obtain their parameter list ParaVarList. If there are parameters, first initialize the parameter list operation. Then, obtain the local variable list LocalVariableList of the current sub-function curFunc, add LocalVariableList to targetFunc, then create the position label FuncLabel of curFunc, and at the same time obtain the start and end positions of the curFunc code, and fuse curFunc into targetFunc. After that, determine whether curFunc has a return value. If there is a return value, perform the operation of updating the return value. Finally, update the external call relationship information of curFunc. After processing all the sub-functions, return the fused M' of the function. The DopSteg control flow flattening is carried out based on M', and the purpose of the function fusion design is to solve the stitching problem of Gadgets between functions.
[0121] In one embodiment, obtaining the semantic execution area of each code segment in the reusable Gadget set in the host program includes: searching for the semantic safe area of each code segment in the reusable Gadget set in the intermediate code space of the host program; the semantic safe area represents the code area where the memory context is not affected when the corresponding code segment is executed in the host program; obtaining the semantic execution area from the semantic safe area according to the backbone instruction; the backbone instruction is an instruction that will be executed only once in the module or function of the host program.
[0122] Optionally, searching for the semantic safe area of each code segment in the reusable Gadget set in the intermediate code space of the host program includes: for each code segment, obtaining the instruction set U of the intermediate code space; obtaining the instruction set U' that modifies the memory space corresponding to each variable vi in the code segment in the instruction set U vi ; according to the predecessor instruction and successor instruction of the code segment on the instruction set U vi to determine the subset between the predecessor instruction and the successor instruction; taking the intersection of the subsets obtained for each variable vi to obtain the semantic safe area of the code segment.
[0123] The DopSteg control flow flattening needs to consider the problem of context data consistency. To solve this problem, the present invention proposes the concepts of semantic safe area and semantic execution area, guides the division of branches, and designs the encapsulation of the semantic execution area. Under the combined action of these mechanisms, data-driven control flow flattening is realized.
[0124] Semantic safe area:
[0125] Definition 1: The semantic safe area refers to the code area where the steganographic semantic Gadget can ensure that the memory context is not affected when executed in the host program.
[0126] Here, a formal description of the determination of the semantic safe area is given. Let U be the set of instructions in the source code space of the host program, and the sorting order of the instructions is the physical sorting order of the source code of the host program. U(i, j) (i < j) is a proper subset of the instruction set U, representing the set of all instructions from the i-th instruction (excluding) to the j-th instruction (excluding) (automatically adjusted to the boundary when i and j cross the boundary). Let U v be the set of instructions in U that modify the memory space corresponding to the variable v, idx(U, inst) be the sorting of the instruction inst in the set U, pred(U, inst) be the direct predecessor instruction of the instruction inst in the set U, and succ(U, inst) be the direct successor instruction of the instruction inst in the set U. Combining the above definitions, if vi is the i-th operand of the Gadget g, a safe area S of g can be expressed using formula (8) g :
[0127]
[0128] Formula (8) expresses the method of obtaining the safe area S of g g . For each variable vi in g, find the predecessor instruction and successor instruction of g on U vi , and delimit the subset between these two instructions on U. This subset can ensure that the variable vi can be affected only by the semantic short instruction g in this subset. Take the intersection of the subsets obtained for each vi, and it can be ensured that each variable in g will be affected only by g in S g .
[0129] Semantic execution area:
[0130] Definition 2: The semantic execution area is a subset of the semantic safe area and can ensure that the Gadget will be executed only once in the execution area.
[0131] The semantic safe area S g ensures that each operand will not be affected by other instructions when the Gadget g is executed in this set, but in actual stitching execution, S g cannot be directly used as an execution area for g. Because when g is a semantic fragment in a branch structure (a loop structure can be regarded as a type of branch structure), and the branch judgment process is also in S gWhen it is [condition not specified], the number of executions of g becomes uncertain. It may not be executed, executed once, or executed multiple times. Therefore, it is necessary to make a more detailed division of S g to determine the execution area of g. We need to use the backbone instructions to help determine the execution area. The backbone instructions are defined as described in Definition 3.
[0132] Definition 3: The backbone instructions refer to the instructions that will be executed exactly once in certain modules or functions of the host program.
[0133] If the process of judging the backbone instruction is denoted as ifTrunk(inst), then the execution area P g can be calculated simply as follows:
[0134] P g = ifTrunk(g)? S g : {g} (9)
[0135] Formula (9) expresses the calculation method of obtaining the execution area P of g from the safe area S of g by combining the backbone instructions g Specifically, if g is a backbone instruction, then regardless of whether there is a branch structure in S g , g will be executed exactly once. More generally, in this case, any subset of S g that contains g can be used as the execution area P of g g ; if g is not a backbone instruction, then the number of executions of g in S g cannot be determined, so g can be extracted separately as the execution area P g , so that during the hidden semantic execution, P g will be directly executed, ensuring that g will be executed exactly once. g
[0136] S104, encapsulate the code of each semantic execution area in a control flow branch, and perform control flow flattening on the encapsulated code to obtain the intermediate representation code.
[0137] When performing semantic steganography based on data-driven code reuse, DopSteg first obtains the set of semantic Gadgets in the host program according to the steganographic semantic Payload written in the HSRL language, then calculates the execution area of each semantic Gadget, and finally encapsulates the code in each control flow branch.
[0138] When performing steganographic semantics, it is necessary to first control the branch selector as the semantic dispatcher Dispatcher of the semantic Gadget, and then sequentially execute the semantic Gadgets in each semantic execution area. Therefore, in each controllable branch, in addition to the semantic execution area, there should also be variables that can be controlled by the Dispatcher. Taking the control-flow flattened While-Swtich structure as an example, this control variable is the branch selection condition of the Switch. When executing the semantics of the original program, the branch condition ensures the correct execution of the original semantics; when executing the hidden semantics, the branch condition will be controlled by the Dispatcher and used to selectively stitch and execute the semantic Gadgets in each semantic execution area.
[0139] Through the joint cooperation of the above mechanisms such as the semantic security area, semantic execution area, and branch encapsulation, DopSteg can ensure that the steganographic semantic Gadget chain can be executed without memory context inconsistency problems when stitched and executed.
[0140] In one embodiment, when performing control-flow flattening on the encapsulated code, branch predicates are set after each branch of the encapsulated code to achieve the stitching of steganographic semantics.
[0141] Specifically, for the purpose of realizing data-driven code reuse, the present invention makes a special design for the branch predicate, rather than simply selecting branches through branch i. Instead, a variable J needs to be added. The specific form is as shown in formula (10).
[0142] i = (i + J) * J (10)
[0143] When DopSteg performs flattening, it is necessary to add a statement such as formula (10) at the end of each branch, which can not only realize the semantic equivalent transformation of the host program, but also achieve the stitching of steganographic semantics under special inputs. The specific implementation principle: Taking 32-bit integers as an example, when the program has a normal input, when the original control flow of the host program enters the While-Switch structure, the value of the branch selection variable i is 0x0, and the value of the variable J is 0x1. Formula (10) becomes i = i + 1. Therefore, the program will automatically execute basic block 1 -> basic block 2 -> basic block 3 in sequence, and this partial order relationship is consistent with that before flattening, realizing the equivalent transformation of semantics. If the steganographic semantics need to execute Gadget1 first and then Gadget2, then when the original semantics of the host program executes to basic block 1, by using stack overflow, make the variable i = 0x20 and the variable J = 0x0, then the next branch selection will choose basic block 3. In basic block 3, the value of i is modified to 0x0 again, so the control flow will return from basic block 3 to basic block 1. In basic block 1, make i = 0x01 and the J value remains unchanged, then basic block 2 can be executed. Similarly, after basic block 2 is executed, it will still return to basic block 1.
[0144] This method of constructing the Dispatcher structure by using the While-Switch structure in combination with stack overflow and J variables is a general construction method. When targeting other host programs, other construction methods can be selected. For example, when facing a host program of the message processing type, different types of messages can be sent to select different branches, and the semantic stitching function of the Dispatcher can also be implemented.
[0145] S105, Compile the intermediate representation code into a binary file to implement steganography of the semantic to be hidden.
[0146] The intermediate representation code is compiled into a binary file through the LLVM backend to implement steganography of the semantic to be hidden.
[0147] In an exemplary embodiment, the present invention provides a semantic hiding method based on data-driven code reuse. This method needs to solve the following problems: Problem 1: How to express the hidden semantic; Problem 2: How to find the reusable code of the host program, that is, the semantic Gadget; Problem 3: How to solve the inappropriateness of the semantic Gadget; Problem 4: How to ensure that the semantic Gadget can be correctly stitched.
[0148] For Problem 1, the present invention defines a language HSRL for expressing hidden semantics, and the HSRL language can realize the expression of hidden semantics. For Problem 2, the present invention defines the characteristics of the Gadget and designs a search algorithm. For Problem 3, the present invention designs a missing Gadget insertion algorithm. For Problem 4, the present invention gives the definitions of the semantic safe area and the semantic execution area, and conducts an encapsulation design on the semantic execution area. At the same time, a special design is made for the control flow flattening branch predicates. Under the combined action of these mechanisms, the host program after DopSteg processing can realize the correct stitching and execution of the hidden semantic Gadget chain, and finally complete program steganography.
[0149] Such as Figure 2As shown in the figure, the method includes the following steps: Step 1: A language HSRL is designed to express the semantic information to be hidden. After the semantic information to be hidden is expressed by HSRL, it is called the steganographic semantic Payload. Parse the Payload to obtain the corresponding Gadget type of the Payload; Step 2: Use a rule-based method to search for reusable Gadgets in the intermediate code of the host program. If the Gadget is not appropriate, the corresponding Gadget needs to be inserted to complete the missing Gadget; Step 3: Combine the semantic execution area to divide and encapsulate the host program P, and perform an equivalent transformation on P through control flow flattening to obtain the hidden semantic version of the program p'; Step 4: Compile the binary program for p' through the LLVM backend to implement the steganography of the semantic information.
[0150] To implement the method provided by the present invention, a semantic steganography prototype system DopSteg based on data-driven code reuse is designed. The system is divided into semantic steganography and execution verification as a whole. The semantic steganography process includes parsing the steganographic semantic Payload, searching for reusable semantic Gadgets in the host program, and under the guidance of data-driven, dividing and encapsulating the semantic execution area of the host program, and through control flow flattening processing, implementing an equivalent transformation of the host program. Finally, the transformed IR code (intermediate code) is compiled into a binary file through the LLVM backend to implement the steganography of the semantic information. The execution verification aims to use the trigger information provided by the steganography link to trigger the execution of the steganographic semantics and realize the verification of the steganographic semantics. The semantic steganography part is implemented in LLVM IR, and the execution verification process is implemented when executing the host program binary file. In the specific implementation process, the following four problems need to be solved:
[0151] Problem 1: How to implement the definition, recognition, and analysis of the steganographic semantic language HSRL;
[0152] Problem 2: How to extract the host program variable characteristics from the virtual registers and memory operations under the LLVM framework to implement Gadget search and insertion;
[0153] Problem 3: How to implement the extraction of the host main trunk, that is, how to implement the calculation of ifTrunk(inst) above to implement the encapsulation of the flattened branch;
[0154] Problem 4: How to implement the equivalent transformation of control flow flattening and the stitching of steganographic semantics.
[0155] For Problem 1, DopSteg uses the Antlr4 language analysis framework to implement the definition, recognition, and analysis of the HSRL language. Specifically, we use the g4 file to define the HSRL language. Table 2 gives a specific HSRL language definition code
[0156] Table 2
[0157]
[0158]
[0159] For problem two, the system uses a state machine structure to perform feature analysis on the host semantic variables according to the definition of the HSRL language. When performing specific analysis, Antlr4 is used to implement the HSRL language. The following selects several important types of instructions for explanation.
[0160] Alloca operation: By analyzing the llvm::AllocaInst instruction, the memory variables of the host program are obtained;
[0161] Calculate operation: Select llvm::AddInst as a representative, and analyze variable operations in combination with other instructions; Deref operation: Analyze llvm::LoadInst, llvm::StoreInst, and llvm::GetElementPtrInst; Output operation: Use llvm::CallInst to analyze the output operands of the printf function.
[0162] The state machine for analyzing the characteristics of host variables in the system includes eight states: Empty state, Alloca state, Load state, Calculate state, StructOffset state, SimpleDeref state, Call state, and Store state. There are three conditions for triggering state transitions. The first is the recognition of special LLVM instructions, the second is the analysis of the operands of LLVM instructions, and the third is the unconditional transfer after the completion of a complete memory operation.
[0163] For problem three, depth-first traversal is used to preprocess the loop structure in the host program, then breadth-first traversal is used to calculate the branch probabilities of the branch structure, and finally the extraction of the main instructions in the host program is realized. The essence of the loop structure preprocessing is to perform a loop removal operation on the loop structure in the host program. Specifically, it only needs to disconnect the first basic block of the loop body from its predecessor block to complete. The purpose of the loop structure preprocessing is to avoid the problem that the branch probability cumulative value is greater than 1 due to the calculation of the loop probability, thus affecting the probability value of the subsequent basic block.
[0164] To facilitate the calculation of branch probabilities, according to Definition 3, five inferences are given.
[0165] Inference 1: If an instruction will execute only once, then the basic block where it is located is the main basic block.
[0166] Corollary 2: If an instruction of a semantic Gadget falls into a trunk basic block, the result of judging the semantic Gadget using ifTrunk(g) is true.
[0167] Corollary 3: The initial execution probability of all basic blocks in the function to be analyzed is 0, but the basic block at the function entry is a trunk basic block, and its initial execution probability is 1.
[0168] Corollary 4: If a basic block B has P basic blocks that can be jumped to, then the execution probabilities of these basic blocks should increase by one over P of the execution probability of basic block B.
[0169] Corollary 5: If there are multiple basic blocks that can jump to basic block B, then the execution probability of basic block B is the sum of the execution probabilities of these basic blocks.
[0170] Specific steps: First, all trunk basic blocks can be analyzed based on the LLVM framework, and then it is judged whether the first instruction of the semantic Gadget falls into the trunk basic block. The truth value of this judgment is the truth value of the function ifTrunk(g). When encountering branch conditions that cause different execution probabilities for each branch, it can be judged whether it is a trunk instruction by calculating which basic blocks' execution probabilities are strictly equal to 1. When encountering a loop structure, first perform the loop removal operation, and then calculate the execution probability value.
[0171] For Question 4, the classic While-Switch structure in control flow flattening is used as an example for equivalent transformation, and the encapsulation of each semantic Gadget execution area is fully implemented. Based on function fusion, semantic safe areas, semantic execution areas, and the design of control flow flattening, the data-driven control flow flattening and branch predicate design are as Figure 3 shown.
[0172] Figure 3 On the left is the modification we made to the Case structure. In each basic block, the modification of the next-hop branch selection condition is not a direct assignment, but requires the participation of variable J. On the right is the local stack space structure. Taking the stack overflow control method in the host program as an example, when executing the Gadget semantic chain, not only the value of branch selection variable i needs to be modified, but also the value of variable J needs to be modified. How to obtain the predicate value to achieve steganographic semantic stitching, please refer to the relevant description of the branch predicate design.
[0173] The semantic hiding method based on data-driven code reuse proposed by the present invention is a general code semantic hiding method. However, in the specific implementation, the present invention selects to modify the branch predicate through an overflow vulnerability because it is easier to understand how the work of the present invention concatenates the hidden semantics in the normal logical code in this way. In fact, more other ways can be taken to complete the modification of the branch predicate. For example, for an interactive program, the modification can be carried out through its own interactive mode. As long as a way to achieve multiple modifications of the branch predicate can be found, the method of the present invention is applicable.
[0174] DopSteg is inspired by the DOP idea and realizes program steganography by the way of data-driven code reuse. Because DopSteg does not change the control flow, its applicable range is larger than that of RopStep. However, its implementation process adopts the control flow flattening method, which indeed changes the control flow. Here, not changing the control flow means not changing the control flow during the dynamic execution of the program, rather than the control flow during the static analysis of the host original program. The ultimate goal of the present invention is to be able to bypass CFI detection, which is also a prerequisite for being applicable to the actual environment and an advantage of the method of the present invention.
[0175] The essence of traditional code obfuscation to achieve semantic hiding is to increase the difficulty for software analysts to understand the code through the obfuscation method. This increased difficulty of understanding is more about the difficulty of understanding the code semantics of the original function of the host program. The method provided by the present invention is different. The method of the present invention focuses more on hiding the code semantics different from the original function into the host program through the way of data-driven code reuse. DopSteg can provide new ideas for software copyright protection.
[0176] The semantic hiding method based on data-driven code reuse proposed in the present invention integrates the idea of data programming. Through the semantic execution area, code partitioning is carried out on the intermediate representation of the host program. By using the special design of control flow flattening and branch predicates, an equivalent transformation of the host program is realized. At the same time, the reusable code fragments are encapsulated into the switch branch in the loop structure to hide the new code semantics therein. Theoretical analysis shows that DopSteg can have the ability to hide Turing-complete semantics, with stable additional overhead and practical value. Through experimental evaluation and comparative analysis, it shows that the DopSteg method is effective, with significant anti-homology detection and anti-reverse analysis effects, and a good balance is achieved between hiding complex semantics and space-time overhead. DopSteg can pay more attention to hiding new semantics different from the original function of the host program and can provide new ideas for software protection.
[0177] Among them, the semantic hiding method based on data-driven code reuse provided by the present invention is compared with conventional control flow flattening. The control flow flattening algorithm corresponding to the method of the present invention further splits each basic block according to the semantic execution area at the IR level in a fine-grained manner, making the modified program reorganization more flexible and convenient to organize into more complex stego semantics.
[0178] The main steps of the algorithm are divided into two stages: splitting and flattening:
[0179] 1. Splitting phase: The algorithm performs fine-grained splitting on each basic block in the IR. Specifically, it traverses the instructions of each basic block and decides whether to split based on the following conditions:
[0180] Store instruction: The function of the Store instruction is to store the value in the virtual register to the virtual stack, which usually corresponds to the real stack when the program is running. Since the store instruction often involves the assignment operation of the variables on the stack, choosing the store instruction as the splitting point can effectively ensure that the number of times the stack variables are modified in each basic block is at most once. This restriction reduces the risk of stack variables being frequently modified.
[0181] Call instruction: If the function called by the Call instruction has no return value, or its return value is not used, it usually means that the function is called to perform its side effects. In this case, using the Call instruction as a split point can significantly reduce the occurrence of multiple consecutive function calls in the program control flow, thereby reducing the possibility of program logic being hijacked by external parties.
[0182] 2. Flattening stage: This algorithm continues the basic idea of conventional control flow flattening and uses the switch instruction unique to LLVM IR to reorganize the control flow of the program in the form of a branch table. The specific implementation method is as follows:
[0183] Each basic block is considered as an independent case branch. A global state variable is used to transfer states between basic blocks. After each basic block executes its instructions, the value of the state variable is updated to point to the next basic block to be executed. This flattening method effectively disrupts the control flow order of the program and can effectively cover up the original semantics of the program.
[0184] An embodiment is provided below for specific explanation. For the native control flow graph of the program IR code, Figure 4The program, whose basic semantics is to determine whether the user input is "push" or "pop" and perform the corresponding operations, that is, the enqueue and dequeue semantics of the circular queue. It can be seen that the two semantics are implemented by one basic block respectively, and the remaining basic blocks only provide jump operations. Without changing the control flow, even after control flow flattening, the program can only execute the complete dequeue and enqueue semantics.
[0185] The CFG of the program after the splitting algorithm is as Figure 5 shown. It can be seen that the original 6 basic blocks are split into multiple basic blocks. After control flow flattening, the program can jump between basic blocks without changing the CFG, with higher flexibility in recombination. At the same time, the characteristic that each basic block modifies at most one stack variable ensures data integrity.
[0186] Comparison with the existing algorithm: The structure of the algorithm after splitting and flattening is as Figure 6 shown. It can be seen that its control flow becomes flattened compared with Figure 4 . At the same time, each basic block returns to the switch basic block after execution, making it difficult for static analysis to judge the execution path of the program, having a semantic hiding effect. Compared with the ordinary control flow flattening algorithm without splitting, the splitting of basic blocks is more meticulous, as Figure 7 shown. Therefore, it has a richer semantic expression ability. Splitting with the store instruction at the end ensures data integrity and consistency when restoring the hidden semantics.
[0187] For example, for Figure 7 the basic block numbered 18, which represents the push semantics of the source program, it is equivalent to the basic blocks 33405058667480 in this new flattening algorithm. For the ordinary control flow flattening algorithm, executing the basic block numbered 18 can only execute the complete push semantics, while for the DopSteg algorithm, it can jump to the scanf semantics of the 33 basic block, the printf semantics of the 80 basic block, etc., improving the flexibility of program Gadget reuse and recombination.
[0188] It should be noted that Figures 4 - 7 only shows the overall framework of the algorithm program, and the specific code representation content is not limited.
[0189] When applying the semantic hiding method based on data-driven code reuse provided by the present invention, it is not necessary to execute according to the Figure 1 sequence of each step shown. The specific execution sequence of each step can be determined as needed, and the present invention does not limit this.
[0190] The above is a semantic hiding method based on data-driven code reuse provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding semantic hiding device based on data-driven code reuse, as Figure 8 shown.
[0191] Figure 8 FIG. is a schematic diagram of a semantic hiding device based on data-driven code reuse provided by the present invention. The device 800 includes:
[0192] A declaration module 801, configured to perform semantic declaration on the semantic to be hidden through a predefined semantic expression to obtain a steganographic semantic;
[0193] An analysis module 802, configured to analyze the steganographic semantic to determine the reusable Gadget type in the steganographic semantic; Gadget is a code snippet for implementing basic operations;
[0194] A search module 803, configured to search for a set of reusable Gadgets in the host program according to the reusable Gadget type, and obtain the execution area of each code snippet in the set of reusable Gadgets in the host program;
[0195] A processing module 804, configured to encapsulate the code of each semantic execution area in a control flow branch, and perform control flow flattening processing on the encapsulated code to obtain intermediate representation code;
[0196] A compilation module 805, configured to compile the intermediate representation code into a binary file to implement the steganography of the semantic to be hidden.
[0197] For the specific limitations on the semantic hiding device based on data-driven code reuse, reference may be made to the limitations on the semantic hiding method based on data-driven code reuse in the above text, which will not be elaborated here. Each module in the above semantic hiding device based on data-driven code reuse can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0198] The present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program can be used to execute the above Figure 1 provided semantic hiding method based on data-driven code reuse.
[0199] The present invention also provides Figure 9 the structural schematic diagram of the computer device shown in FIG., as Figure 9As shown, at the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 semantic hiding method based on data-driven code reuse provided.
[0200] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0201] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present invention.
Claims
1. A semantic hiding method based on data-driven code reuse, characterized in that, Including: Semantically declare the semantic to be hidden through a predefined semantic expression to obtain a steganographic semantic; Parse the steganographic semantic to determine the reusable Gadget types in the steganographic semantic; A Gadget is a code snippet for implementing a basic operation; Search for a set of reusable Gadgets in the host program according to the reusable Gadget types; Obtain the semantic execution area of each code snippet in the set of reusable Gadgets in the host program; encapsulate the code of each semantic execution area in a control flow branch, and perform control flow flattening on the encapsulated code to obtain intermediate representation code; Compile the intermediate representation code into a binary file to implement the steganography of the semantic to be hidden.
2. The method according to claim 1, characterized in that, The predefined semantic expression includes keyword design and expression design; the keyword design includes declaring the meanings of multiple keywords; the expression design includes defining basic operation types and corresponding formal descriptions; the basic operation types include declaration, assignment, non-dereferencing operation, dereferencing operation, and output.
3. The method according to claim 1, characterized in that, The searching for a set of reusable Gadgets in the host program according to the reusable Gadget types includes: Perform static search on the intermediate code of the host program compiler to determine the set of reusable Gadgets corresponding to the reusable Gadget types in the host program.
4. The method according to claim 3, characterized in that, The static search includes function traversal search; the performing static search on the intermediate code of the host program compiler to determine the set of reusable Gadgets corresponding to the reusable Gadget types in the host program includes: Traverse all functions in the intermediate code of the compiler. If the current function is the target function or a sub-function of the target function, analyze each instruction in the current function; the target function is the function where the vulnerability is located; If the instruction is a system call instruction, obtain the parameters of the instruction. If the parameters of the instruction are controllable parameters, form a Gadget with the instruction and its context instructions; If the instruction is not a system call instruction, obtain the source operand and destination operand of the instruction. If the instruction source operand and destination operand are global variables or controllable parameters, form a Gadget with the instruction and its context instructions; Match Gadgets of the corresponding type in the host program according to the reusable Gadget types in the steganographic semantic, and form the set of reusable Gadgets with all the matched Gadgets.
5. The method according to claim 1, wherein The method further includes: When there is no set of reusable Gadgets in the host program, obtain the intermediate code of the host program compiler and the Gadgets to be inserted; If the type of the Gadget to be inserted is a global variable, directly insert the Gadget to be inserted into the intermediate code of the compiler; If the type of the Gadget to be inserted is a local variable, assignment type, or operation type, obtain the position information to be inserted, and insert the Gadget to be inserted into the intermediate code of the compiler according to the position information; If the type of the Gadget to be inserted is a conditional type, obtain the position information and basic block information to be inserted, and insert the Gadget to be inserted into the compiler intermediate code according to the position information and basic block information.
6. The method according to claim 1, wherein The obtaining of the semantic execution area of each code snippet in the reusable Gadget set in the host program includes: Search for the semantic safe area of each code snippet in the reusable Gadget set in the intermediate code space of the host program; the semantic safe area represents the code area where the memory context is not affected when the corresponding code snippet is executed in the host program. Obtain the semantic execution area from the semantic safe area according to the backbone instruction; the backbone instruction is an instruction that will be executed only once in the module or function of the host program.
7. The method according to claim 6, characterized in that, The searching for the semantic safe area of each code snippet in the reusable Gadget set in the intermediate code space of the host program includes: For each code snippet, obtain the instruction set of the intermediate code space. Obtain the instruction set that modifies the memory space corresponding to each variable in the code snippet from the instruction set. Determine the subset between the predecessor instruction and the successor instruction according to the predecessor instruction and the successor instruction of the code snippet on the instruction set. Take the intersection of the subsets obtained for each variable to obtain the semantic safe area of the code snippet.
8. The method according to claim 6, wherein The Gadget set includes multiple Gadgets; the method further includes: Before obtaining the execution area of each code snippet in the reusable Gadget set from the host program, adopt the strategy of function inlining to fuse the functions where the Gadgets are located in the compiler intermediate code into one function to obtain the fused compiler intermediate code.
9. The method according to claim 8, wherein The fusing of the functions where the Gadgets are located in the compiler intermediate code into one function to obtain the fused compiler intermediate code includes: Obtain the variable set of the target function in the compiler intermediate code. Traverse all the sub-functions in the target function. For any sub-function, perform an initialization operation on the parameter list of the sub-function and add the local variables of the sub-function to the target function. Obtain the start position and end position of the code of the sub-function, and fuse the sub-function into the target function according to the start position and end position of the code. Judge whether the sub-function has a return value. If the sub-function has a return value, perform an operation to update the return value and update the external call relationship information of the sub-function. After processing all the sub-functions, return the fused compiler intermediate code.
10. The method according to claim 1, wherein The method further includes: When performing control flow flattening on the encapsulated code, set branch predicates after each branch of the encapsulated code to achieve the stitching of the steganographic semantics.