Code obfuscation method and system based on function splitting
By splitting functions and dividing control flow graph dominator trees, a highly dispersed sub-function structure is generated, which solves the balance problem between running overhead and obfuscation strength in traditional code obfuscation methods, and achieves lower performance overhead and higher obfuscation concealment.
Patent Information
- Application Number
- CN202510938200.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-16
AI Technical Summary
Traditional code obfuscation methods have difficulty balancing obfuscation strength and runtime overhead, and cannot effectively resist deep learning-driven binary code similarity detection.
By constructing a code obfuscation method for function splitting, using control flow graphs and dominator trees to perform subgraph division, a highly dispersed sub-function structure is generated. Combining data dependencies and control flow repair, the original function is rewritten into sub-function code to enhance the obfuscation concealment.
Without significantly increasing code size and operating overhead, it reduces the recall rate of mainstream detection models, improves the concealment of code obfuscation, and resists code static analysis.
Smart Images

Figure CN120654216A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software security technology, and in particular to a code obfuscation method and system based on function splitting. Background Art
[0002] Code obfuscation refers to the act of converting the code of a computer program into a form that is functionally equivalent but difficult to read and understand. Code obfuscation technology plays a key role in software security protection, and is mainly used to protect source code from being easily copied or reverse engineered. Traditional code obfuscation methods increase the difficulty of reverse engineering by transforming structures, symbols, and data, but they are essentially "hidden security" and cannot withstand dynamic debugging and advanced analysis tools. However, binary code similarity detection technology in reverse engineering based on deep learning can quickly identify similarities between code and known vulnerabilities, forcing traditional obfuscation methods to increase the intensity of obfuscation to meet the challenge, but this also increases the program running overhead. Summary of the Invention
[0003] To this end, the present invention provides a code obfuscation method and system based on function splitting, which solves the problem that traditional obfuscation is difficult to balance obfuscation strength and running overhead. The function splitting strategy destroys the code structure that similarity detection relies on, while ensuring the code obfuscation strength and achieving lower code running performance overhead.
[0004] According to the design scheme provided by the present invention, on the one hand, a code obfuscation method based on function splitting is provided, comprising:
[0005] Constructing a control flow graph of the original function in the source code to be processed, and dividing the control flow graph into a plurality of subgraphs according to the dominance relationship between nodes in the control flow graph, wherein the control flow graph and the subgraphs are composed of code basic blocks;
[0006] Based on the subgraph basic block path and variables, a set of variables with data dependency relationships between functions corresponding to the subgraph is obtained, the subgraph entry nodes are uniformly identified according to the subgraph control flow, and the original function code is rewritten into sub-function codes corresponding to several subgraphs based on the set of variables with data dependency relationships between functions and the unified identification, so as to transform the control flow relationship within the original function into a direct and / or indirect function call relationship between several sub-functions.
[0007] As the code obfuscation method based on function splitting of the present invention, further, a control flow graph of the original function in the source code to be processed is constructed, including:
[0008] Obtain a basic block set based on the continuous branchless code in the original function of the source code to be processed, and obtain an edge set based on the control flow transfer relationship between the basic blocks;
[0009] A directed graph describing the control flow graph of the original function is constructed using a basic block set and an edge set.
[0010] As the code obfuscation method based on function splitting of the present invention, the control flow graph is further divided into several subgraphs according to the dominance relationship between nodes in the control flow graph, including:
[0011] Constructing a control flow graph dominance tree based on dominance relationships between nodes, wherein the dominance relationships are used to describe the situation where all paths from the control flow graph entry to the basic block pass through other basic blocks;
[0012] Based on hierarchical progression, the control flow graph dominator tree is traversed in post-order and cohesive regions are generated. The control flow graph is then divided into several subgraphs according to the cohesive regions.
[0013] As the code obfuscation method based on function splitting of the present invention, further, based on hierarchical progression, the control flow graph dominator tree is traversed in post-order and cohesive regions are generated, including:
[0014] Set the dominating subtree size threshold and subgraph balance ratio;
[0015] Traverse the control flow graph dominator tree nodes from the bottom up, recursively select nodes with dominance relationships to construct subgraphs, and determine whether the size of each node subtree is greater than the dominator subtree size threshold. If so, recursively divide the current node subtree. Otherwise, merge the current node subtree into a region and add it to the region set R.
[0016] Detect the natural loop set in the node subtree control flow graph and record the loop head node and loop body of each loop. If the region contains the loop head node, merge the corresponding loop body into the corresponding region of the node subtree;
[0017] Candidate regions are selected from the region set, and several subgraphs are generated according to the subgraph balance rate.
[0018] As the code obfuscation method based on function splitting of the present invention, further, according to the subgraph basic block path and variables, a variable set with data dependency relationship between the corresponding functions of the subgraph is obtained, including:
[0019] Compare the value sets of each variable in the subgraph at the program point before the head node and the program point after the tail node to obtain the program state difference, and obtain the cross-function transfer variable set based on the forward call data flow and reverse return data flow of the cross edges connecting each subgraph, where the program point before the head node and the program point after the tail node are two program points in the original function execution path obtained based on the cross edges of the subgraph;
[0020] According to the program state difference and the variable set transferred across functions, a variable set with data dependency between the corresponding functions of the subgraph is obtained. The variable set with data dependency is a variable set based on function calls and returns.
[0021] As the code obfuscation method based on function splitting of the present invention, further, the subgraph entry nodes are uniformly identified according to the subgraph control flow, including:
[0022] Assign a unique identifier to the entry node of each subgraph, and set a corresponding identifier at the subgraph exit node. The entry node of the subgraph is the target node of all edges entering the subgraph from other areas, and the subgraph exit node is the source node of all edges leaving the subgraph to other subgraphs.
[0023] As the code obfuscation method based on function splitting of the present invention, further, according to the variable set and unified identification of the data dependency relationship between functions, the original function code is rewritten into sub-function codes corresponding to several subgraphs, including:
[0024] Set the structure type and structure variables based on the variable set with data dependency relationship between the corresponding functions of the subgraph and the unified identification of the subgraph entry node;
[0025] In the original function execution path and at the head node and tail node of the subgraph intersection edge, code for structure variable assignment / removal operations and function call / return code are generated, and a switch basic block for selecting the target branch based on a unified identifier is set at the entry and call return of each subfunction.
[0026] As the code obfuscation method based on function splitting of the present invention, further, according to the variable set and unified identification of the data dependency relationship between functions, the original function code is rewritten into sub-function codes corresponding to several subgraphs, which also includes:
[0027] Traverse all basic blocks in the original function and merge all exit basic blocks; randomly generate a new basic block as the unified exit node, replace the return instruction with a jump instruction, so that the control flow graph has only one exit basic block;
[0028] And / or, convert the Phi nodes to memory-based operations and fix the escaped variables through memory operations.
[0029] As the code obfuscation method based on function splitting of the present invention, further, according to the variable set and unified identification of the data dependency relationship between functions, the original function code is rewritten into sub-function codes corresponding to several subgraphs, which also includes:
[0030] The rewritten sub-function is further split and rewritten, and the sub-function call is triggered by a triggering method, wherein the triggering method includes function pointer triggering, signal parameter triggering or exception triggering;
[0031] And / or, when dividing the control flow graph into several subgraphs, the subgraph cross-edge nodes with a dominating relationship in the subgraph partition area are further divided into a group, each group corresponds to an independent branch selection block, and multiple branch selection basic blocks are inserted into the corresponding sub-function of the subgraph to hide the upper and lower hierarchical features of the subgraph;
[0032] And / or, a subgraph is selected from the continued sub-function splitting and rewriting of each subgraph sub-function, the selected subgraphs are merged and a new function is constructed, so that the call graph between the original functions presents a mesh structure through multi-subfunction mixed splitting.
[0033] On the other hand, the present invention also provides a code obfuscation system based on function splitting, comprising: a division module and an obfuscation module, wherein:
[0034] A partitioning module is used to construct a control flow graph of the original function in the source code to be processed, and divide the control flow graph into several subgraphs according to the dominance relationship between the nodes in the control flow graph. The control flow graph and subgraphs are composed of code basic blocks;
[0035] The obfuscation module is used to obtain the set of variables with data dependency relationships between the functions corresponding to the subgraph based on the subgraph basic block path and variables, uniformly identify the subgraph entry nodes according to the subgraph control flow, and rewrite the original function code into sub-function codes corresponding to several subgraphs based on the set of variables with data dependency relationships between functions and the unified identification, so as to transform the control flow relationship within the original function into a direct and / or indirect function call relationship between several sub-functions.
[0036] Beneficial effects of the present invention:
[0037] The present invention splits the original function into multiple sub-functions, and through dominator tree-driven subgraph partitioning, loop-aware adjustment and random optimization, it can generate a highly dispersed sub-function structure without significantly increasing the code volume and runtime, destroying the function-level code structure that the existing similarity detection relies on, reducing the recall rate of mainstream detection models (such as Gemini, Asm2Vec, JTrans), and can improve the concealment of code obfuscation through enhanced strategies such as multi-function mixed splitting and cross-edge grouping, resist code static analysis, and meet the usage requirements of software code obfuscation. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is a schematic diagram of the code obfuscation process based on function splitting in the embodiment;
[0039] Figure 2This is a schematic diagram of the architecture of the code obfuscation algorithm FunSp based on function splitting in the embodiment;
[0040] Figure 3 This is an example of a dominator tree in the embodiment;
[0041] Figure 4 This is an example of data stream repair in the embodiment;
[0042] Figure 5 This is an example of control flow repair in the embodiment;
[0043] Figure 6 This is a schematic diagram of the function code rewriting process in the embodiment;
[0044] Figure 7 This is a schematic diagram of function code preprocessing in the embodiment;
[0045] Figure 8 This is an illustration of the mixed splitting of multiple functions in the embodiment. DETAILED DESCRIPTION
[0046] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention is further described in detail below with reference to the accompanying drawings and technical solutions.
[0047] In order to deal with the threat of artificial intelligence driven binary code similarity detection technology to software security, the embodiments of the present invention, such as Figure 1 As shown in the figure, a code obfuscation method based on function splitting is provided, which specifically includes the following contents:
[0048] S101 , constructing a control flow graph of the original function in the source code to be processed, and dividing the control flow graph into a plurality of subgraphs according to the dominance relationship between nodes in the control flow graph, wherein the control flow graph and the subgraphs are composed of code basic blocks.
[0049] S102. Obtain a set of variables with data dependency relationships between functions corresponding to the subgraph based on the subgraph basic block path and variables, uniformly identify the subgraph entry nodes based on the subgraph control flow, and rewrite the original function code into sub-function codes corresponding to several subgraphs based on the set of variables with data dependency relationships between functions and the unified identification, so as to transform the control flow relationship within the original function into a direct and / or indirect function call relationship between several sub-functions.
[0050] Set Function to a five-tuple, expressed as:
[0051] F=<Name,Params,ReturnType,Body>
[0052] Among them, Name is the unique identifier of the function, Params is the parameter list, ReturnType is the return value type, BodyWie is the set of basic blocks, and Body(F) = {Entry, b1, b2, b3, ..., Exit}.
[0053] The control flow graph G of a function is used to represent the control flow information of the function. The nodes of the control flow graph are basic blocks, and the edges represent the relationships between basic blocks. A basic block is a continuous and branchless sequence of instructions.
[0054] The original function to be split is recorded as:
[0055] f unc_o=<′f unc_o′,params_o,ret_o,{Entry_o,b1,..b i ,...b j ,Exit_o}>
[0056] Select some basic blocks and the edges between them from the control flow graph of function func_o to construct a new function func_b, and the remaining basic blocks and edges construct function func_a, denoted as
[0057] func_a=<′f unc_a′,params_o,ret_o,{Entry_a,b1,...,b i ,Exit_a}>
[0058] func_b=<′f unc_b′,params_b,ret_b,{Entry_b,b i+1 ,...,b j ,Exit_b}>
[0059] Functions func_a and func_b jointly implement the functionality of the original function.
[0060] The algorithm FunSp architecture for splitting the original function func_o into func_a and func_b is as follows Figure 2 As shown, first, the control flow graph G of the function func_o is divided into two subgraphs G a , G b , respectively contain some basic blocks of the original function; use data flow analysis technology to analyze the subgraph G a , G b Perform reach definition analysis and active variable analysis on the basic block of G to obtain the variable set with data dependency between sub-functions; perform control flow analysis on the subgraph G a , G bThe entry node is uniformly identified, which is used to repair the control flow between sub-functions. Based on the results of data flow and control flow analysis, the code of function func_o is rewritten. If the output sub-function is re-input, more sub-functions can be constructed. Program analysis and rewriting are two important steps in compiler optimization. Function splitting aims to obfuscate the program, not optimize it. However, the process is similar to optimization: the first three steps involve program analysis, and the fourth step involves program rewriting.
[0061] The control flow graph of the original function in the source code to be processed is constructed and can be designed to include:
[0062] Obtain a basic block set based on the continuous branchless code in the original function of the source code to be processed, and obtain an edge set based on the control flow transfer relationship between the basic blocks;
[0063] A directed graph describing the control flow graph of the original function is constructed using a basic block set and an edge set.
[0064] Selecting some basic blocks from function func_o to construct function func_b . The basic block selection problem can be mapped to the subgraph partitioning problem of the control flow graph. A function's control flow graph represents the function's control flow information. Nodes in the control flow graph are basic blocks, and edges represent relationships between basic blocks. Let the control flow graph of function func_o be a directed graph G = (V, E), where V is the set of basic blocks, and each basic block v∈V represents a continuous, branchless section of code. It is a set of edges that represent the transfer relationship of control flow. If there is an edge (u,v)∈E, then there is an execution path from basic block u to v in the program.
[0065] Specifically, the control flow graph is divided into several subgraphs according to the dominance relationship between the nodes in the control flow graph, which can be designed to include:
[0066] Constructing a control flow graph dominance tree based on dominance relationships between nodes, wherein the dominance relationships are used to describe the situation where all paths from the control flow graph entry to the basic block pass through other basic blocks;
[0067] Based on hierarchical progression, the control flow graph dominator tree is traversed in post-order and cohesive regions are generated. The control flow graph is then divided into several subgraphs according to the cohesive regions.
[0068] Divide graph G into two subgraphs, denoted as subgraph G a and G b , satisfying: V = V a ∪V b ,and Among them, V a and V b They are subgraphs Ga and G b A set of basic blocks. Each subgraph is defined as:
[0069] G a =(V a ,E a ), where E a ={(u,v)∈E|u,v∈V a}
[0070] G b =(V b ,E b ), where E b ={(u,v)∈E|u,v∈V b}
[0071] A cross edge is an edge that connects two different subgraphs, namely:
[0072] E cross ={(u,v)∈E|(u∈V a ∧v∈V b )∨(u∈V b ∧v∈V a )}
[0073] The opposite of the intersection edge is the inner edge.
[0074] In the process of subgraph partitioning, subgraph G b The vertex set V b is func o The set of basic blocks that need to be transferred, G a The vertex set V a It is the set of remaining basic blocks in func_o. The entry basic block and exit basic block of function func_o cannot be divided into subgraph G. b In addition, some basic blocks can be randomly selected for subgraph division. Random selection reflects the flexibility of the method, but it is not a good choice because different subgraph division methods will bring different runtime and space overheads to the program.
[0075] In the FunSp algorithm, after the function func_o is split, the sub-function func_a needs to call func_b to implement the functionality of the original function. The calling frequency is related to two factors: the first factor is the number of cross edges between subgraphs. Both the inner edges and the cross edges are located on the execution path of the function func_o. The nodes (basic blocks) of the inner edges are inside the function, while the nodes of the cross edges are distributed in the two subgraphs. The execution path corresponding to the cross edges needs to be repaired through function calls after the function is split. Therefore, the more cross edges there are, the higher the calling frequency between sub-functions; the second factor is whether the cross edge is located in a loop. The loop may be executed multiple times. If it is located in a loop, the calling frequency is increased.
[0076] After the function is split, the head node and tail node of the cross edge will be located in func_a and func_b respectively. Variables need to be passed between the two functions. The variables themselves require storage space, and the instructions for assigning and getting values of the variables also require space. In addition, the assignment and retrieval instructions increase the running overhead. Therefore, an algorithm is needed to reduce the subgraph G a ,G b The number of crossing edges between them.
[0077] In the analysis of control flow graph, subgraph partitioning based on dominator tree is an effective strategy. For any two basic blocks b in the control flow graph, i with b j , if from the starting entrance Entry to b j All paths through b i , then b i Dominate b j , at this time b i That is b j Except for Entry, each basic block b j Its direct dominance point idom(b j ) is called a dominator tree. Nodes with dominance relationships are selected to construct subgraphs. The edges between these nodes become internal edges within the subgraph, thus reducing the number of cross-edges between subgraphs. To this end, in this embodiment, subgraph partitioning can be performed based on the dominator tree, hoping to effectively reduce the number of cross-edges between subgraphs.
[0078] Specifically, based on hierarchical progression, the control flow graph dominator tree is traversed in post-order and cohesive regions are generated, which can be designed to include:
[0079] Set the dominating subtree size threshold and subgraph balance ratio;
[0080] Traverse the control flow graph dominator tree nodes from the bottom up, recursively select nodes with dominance relationships to construct subgraphs, and determine whether the size of each node subtree is greater than the dominator subtree size threshold. If so, recursively divide the current node subtree. Otherwise, merge the current node subtree into a region and add it to the region set R.
[0081] Detect the natural loop set in the node subtree control flow graph and record the loop head node and loop body of each loop. If the region contains the loop head node, merge the corresponding loop body into the corresponding region of the node subtree;
[0082] Candidate regions are selected from the region set, and several subgraphs are generated according to the subgraph balance rate.
[0083] like Figure 3 As shown in Figure 2, a dominating subtree is split as a whole. However, the binary code similarity detection model can recover this dominating subtree through function inlining, which leads to a decrease in the adversarial effect. In contrast, the complex splitting method selects basic blocks (such as b4, b7, b 10 ,b 13 ). Based on the above analysis, the subgraph partitioning algorithm can adopt a hierarchical and progressive design, which can be summarized into three stages: dominator-driven preliminary partitioning (Dominator-DrivenPartition), loop-aware adjustment (Loop-Aware Adjustment) and randomized refinement (RandomizedRefinement). The algorithm contains two optional parameters: the dominator subtree size threshold T and the subgraph balance rate α. High cohesion areas are generated by post-order traversal of the dominator tree, and the subgraph ratio is dynamically balanced during the random allocation process to ensure that the subgraph partition meets the ratio requirement of α. The specific steps of the subgraph partitioning algorithm can be summarized as follows:
[0084] 1. Construct a dominator tree: Use the Lengauer-Tarjan algorithm to generate a dominator tree (Dominator Tree) based on the control flow graph.
[0085] 2. Post-order traversal of the dominator tree: traverse the dominator tree nodes in post-order (bottom-up).
[0086] 3. Recursively divide the subtree: For each node v, calculate the size of its subtree |Subtree(v)|:
[0087] –If |Subtree(v)|>T (T is the preset threshold), recursively partition the subtree.
[0088] –If |Subtree(v)|≤T, merge Subtree(v) into a region and add it to the region set R.
[0089] 4. Detect natural loops: Identify the natural loop set L in the control flow graph and record the head node of each loop (Loop
[0090] Header) and loop body (Loop Body).
[0091] 5. Merge loop area: traverse the set of areas R that are initially divided. If an area R i Contains the loop header node,
[0092] Then merge the entire loop body into R i In the process, ensure the integrity of the loop structure.
[0093] 6. Randomly generate candidate partitions: Based on the current partition R, randomly select candidate members from it and balance them according to the subgraph
[0094] The rate α generates the subgraph G a and G b , ensure |G a | / (|G a |+|G b |)≈α.
[0095] By setting the dominating subtree threshold T and the subgraph balancing rate α, the algorithm can strike a balance between program overhead and obfuscation effect: the threshold T controls the granularity of the subgraph. A smaller T value will generate more fine-grained subgraphs, but may increase runtime overhead. The balancing rate α adjusts the subgraph G a and G b The algorithm maintains the size ratio of the subgraphs to avoid overly skewed subgraph divisions. Through a hierarchical and progressive design, the algorithm retains the structural characteristics of the dominator tree while enhancing its ability to combat binary code similarity detection through randomization, while ensuring that the program's performance overhead is controllable.
[0096] Function splitting will destroy the data dependency within the original function, making it difficult to ensure the semantic consistency of the program. The execution of instructions can change the state of the program. Each execution of an instruction will convert an input state into an output state. This input state is associated with the program point before the statement, and the output state is associated with the program point after the statement. The program point is denoted as p. In a basic block, the program point after an instruction is the same as the program point before the next instruction, such as Figure 4 As shown in Figure 1, if a basic block b1 has an edge pointing to b2, then the program point before the first statement of b2 may be immediately after the program point after the last statement of b1. The sequence of program points constitutes the execution path of the program.
[0097] Divide the control flow graph G of function func_o into two subgraphs, subgraph G a Basic block set V aIn the function func_a, the subgraph G b Basic block set V b In the function func_b, such as Figure 4 As shown, the cross edge e cross Connected to subgraph G a , G b , its head node and tail node are in two subgraphs respectively, and the program point before the head node is denoted as p head , the program point after the tail node is recorded as p tail These two program points exist in an execution path of the function func_o. After the function is split, the execution path corresponding to the cross edge will be disconnected. This will cause the data flow between the two subgraphs to be interrupted, so at program point p head 、p tail Data flow repair is required. intra The corresponding execution path is not broken, so no repair is required.
[0098] The data flow between subgraphs is broken and needs to be repaired by passing variables when calling functions. In the process of function splitting, it is necessary to accurately identify the set of variables that need to be passed between sub-functions. To do this, it is necessary to compare p head 、p tail According to the definition of program state in compiler theory, the program state can be expressed as the set of values of all variables at a specific program point. Figure 4 As shown in the figure, the values of these variables are distributed in different storage areas such as the code segment, data segment, heap area, stack area and system memory. Functions func_a and func_b are executed in the same process, each with its own independent stack frame, which is used to store local variables, return addresses and other information during function execution. In addition, functions func_a and func_b share other system resources and memory areas. Therefore, program point p head 、p tail The difference in program state comes from func a and func b local variables.
[0099] To this end, in an embodiment of the present case, the program state difference is obtained by comparing the value sets of each variable in the subgraph at the program point before the head node and the program point after the tail node, and the cross-function transfer variable set is obtained based on the forward call data flow and the reverse return data flow of the cross edges connecting each subgraph. The program point before the head node and the program point after the tail node are two program points in the original function execution path obtained based on the cross edges of the subgraph; based on the program state difference and the cross-function transfer variable set, the variable set with data dependency between the corresponding functions of the subgraph is obtained, and the variable set with data dependency is a variable set based on function calls and returns.
[0100] To ensure data consistency between sub-functions, the connected subgraph G a and G b The cross edge e cross , respectively at its head node program point p head and the tail node program point p tail Perform data flow analysis. Specifically, in p tail Perform the arrival definition analysis to obtain all possible variable definition sets Def that can reach this point p tail, while in p head Perform live variable analysis to determine the set of variables required for subsequent execution p head. The intersection of these two sets Var e cross=Def p tail∩
[0101] Live p head, which is the cross edge e cross The corresponding variable set is passed as a parameter during the sub-function call.
[0102] Based on the characteristics of the control flow direction, the cross edges are divided into two types of key data flow paths:
[0103] One is the forward call data flow, from the subgraph G a Pointing to G b The cross edge set is denoted as E a→b , indicating that from func a to func b The function call process, the set of variables that need to be passed can be expressed as:
[0104]
[0105] The second is to return the data flow in reverse, from subgraph G b Pointing to subgraph G a The cross edge set is denoted as E b→a , indicating that from func b arrive
[0106] func a The return process needs to pass the following variable set:
[0107]
[0108] The set of variables called and returned between sub-functions can provide a basis for rewriting function codes.
[0109] In the control flow graph G of a function, the edge represents the transfer relationship of the control flow, and its head node (basic block) and tail node are the predecessor and successor relationships. In the process of function splitting, the inner edge E of the subgraph intra The control flow transfer relationship represented by is preserved, and the cross edges E between subgraphs cross The control flow transfer relationship represented by it needs to be restored. Therefore, in this embodiment, this problem is solved by uniformly marking the subgraph entry nodes.
[0110] Specifically, the subgraph entry nodes are uniformly identified according to the subgraph control flow, which can be designed to include:
[0111] Assign a unique identifier to the entry node of each subgraph, and set a corresponding identifier at the subgraph exit node. The entry node of the subgraph is the target node of all edges entering the subgraph from other areas, and the subgraph exit node is the source node of all edges leaving the subgraph to other subgraphs.
[0112] Entry Nodes: The entry node of a subgraph is the target node of all edges entering the subgraph from other regions. a , its entry node set is:
[0113]
[0114] Exit Nodes: The exit nodes of a subgraph are the source nodes of all edges that leave the subgraph to other subgraphs. a , its exit node set is:
[0115]
[0116] Subgraph G a , G b There may be multiple entry nodes. Figure 5 As shown, subgraph G b There are four entry nodes b4, b7, b 10 ,b 13 , and the function func b There can be only one entry point. Similarly, the function func a Calling function func b After returning, only one program point can be reached, and the subgraph G a There are three entry nodes b8, b9, b 14 . Both of the above situations require control flow repair.
[0117] Pair graph G a , G b The entry node Entry(G a)、Entry(G b ) assigns a unique identifier flagId. Then at the exit node Exit(G a )、Exit(G b ) Set flagId. Use flagId as a parameter when calling a function and as the return value when the function returns. At the function entry point, you can set a switch basic block to select the correct branch node based on the identifier data. When returning from a call, you can also set a switch basic block to complete the control flow repair between functions.
[0118] Specifically, based on the variable set and unified identifier of the data dependency relationship between functions, the original function code is rewritten into sub-function codes corresponding to several subgraphs, which can be designed to include:
[0119] Set the structure type and structure variables based on the variable set with data dependency relationship between the corresponding functions of the subgraph and the unified identification of the subgraph entry node;
[0120] In the original function execution path and at the head node and tail node of the subgraph intersection edge, code for structure variable assignment / removal operations and function call / return code are generated, and a switch basic block for selecting the target branch based on a unified identifier is set at the entry and call return of each subfunction.
[0121] Divide the control flow graph G of function func_o into subgraphs G a and G b Afterwards, the basic blocks in the subgraph are located in functions func_a and func_b, respectively. Variables used in the basic blocks may be defined in function func_a but used in function func_b. Therefore, variables not defined in the subfunctions must be explicitly declared. Furthermore, at runtime, to ensure that functions func_a and func_b maintain consistent program states at the beginning and end of the intersection edge, variable assignment and retrieval operations must be performed during function calls. This is accomplished through code rewriting.
[0122] Based on the results of data flow analysis and control flow analysis, code rewriting needs to generate the following code in functions func_a and func_b: variable declaration, assignment, value retrieval, control flow identifier setting, and switch branch selection. The algorithm mainly includes two steps: First, define the structure type, which contains the following members: Var a→b 、Var b→a 、flagId. Using this structure type, you can a and func bThe structure variable is defined in the following code; the second is at the intersection of the head and tail program points p head ,p tail Generates code for assigning and getting values of structure variables and code for calling or returning functions.
[0123] Among them, the specific implementation of code generation can be summarized as:
[0124] At program point p tail Generate code at this location, which performs the following operations:
[0125] 1. Set flagId to identify the target basic block of the current branch.
[0126] 2. Var a→b / Var b→a Stored in a structure to save the program context.
[0127] 3. Create a call (or return) instruction and use the structure pointer as a parameter
[0128] Transfer. At program point p head Generate code at this location. These codes complete the following
[0129] operate:
[0130] 4. Load the variable value Var in the structure a→b / Var b→a , restore the program context.
[0131] 5. Use flagId to select the correct control flow branch.
[0132] Final function func a and func b The rewritten logical structure is as follows Figure 6 As shown, the code rewriting steps can be summarized as a five-tuple:
[0133]
[0134] in, Indicates the code generation steps from the state s in the original function to the state s′ in the target function. Each element corresponds to a code generation operation. set Generate inter-function control flow identifier assignment instructions, var set Generates instructions for assigning values to structures, call / ret generates instructions for calling or returning functions, var use Generate instructions to parse the data in the structure and restore it to the corresponding variables, flag use Generates instructions to read the value of the variable flag and generate a jump.
[0135] Code generation operations can be implemented in many ways. For example, call / ret represents the calling method between function a and function b, which can be direct, indirect, or based on exceptions. set ,var use The shared structure can be set as a global variable or temporary space can be allocated in the heap / stack.
[0136] The FunSp algorithm can be implemented at the intermediate language level based on LLVM. The clang compiler is used to compile the C language source code into an LLVM intermediate language file (.bc file). The LLVM intermediate language adopts the static single assignment (SSA) form and has the following characteristics: there are Phi nodes for merging variables from different basic blocks, which increases the association between basic blocks; based on the dominance tree, the scope of the variable is determined by its dominance relationship, and escape variables (that is, variables that cross basic blocks) may cause symbol dominance exceptions.
[0137] The above characteristics enhance the correlation between basic blocks, which is not conducive to subgraph partitioning on the control flow graph of the function, so it is necessary to preprocess the intermediate language file generated by the compiler.
[0138] like Figure 7 As shown in the figure, there may be multiple exit basic blocks (including ret instructions) in the control flow graph of the function. If the exit block is transferred, special processing is required. To this end, all exit basic blocks can be merged, all basic blocks in the function are traversed, and blocks containing 'ret' instructions are collected; then a new basic block is generated as a unified exit node, and the original return instruction is replaced with a jump instruction, so that the control flow graph of the function has only one exit block. The Phi instruction is a core component of the SSA (static single assignment) form, and its value depends on the predecessor basic block. This dependency means that when moving a basic block containing a Phi instruction, its predecessor block must be transferred together, which limits the flexibility of basic block selection. In order to support more flexible area selection, the Phi instruction needs to be converted into a memory-based operation. As shown in the figure, Figure 7 As shown in (b) in the figure. Escaped variables are variables whose scope exceeds the function or code block in which they are defined. Such variables may complicate data flow analysis during function splitting, increase memory access overhead, and introduce unnecessary data dependencies. To simplify the data flow processing process, escaped variables need to be repaired, as shown in Figure 7 For (c) in the figure, the repair method can refer to the phi instruction processing method.
[0139] To enhance the code obfuscation effect, you can set the following enhancement strategy:
[0140] Taking the split function as input, further split it into more sub-functions. This function splitting method will generate a large number of "single-call" functions. Therefore, direct function calls can be replaced with indirect calls, using function pointers or triggering function calls through signals (Signal) or exceptions (Exception), making it impossible for static analysis to resolve the specific call target.
[0141] In the FunSp algorithm, func o Split into func a 、func b After that, func a Need to call func b Complete the original function, in the algorithm framework, func a Only call func b Once. Figure 6 As shown, due to func a The control flow points to the same basic block containing the call instruction, resulting in func a The control flow graph (CFG) of a function exhibits a hierarchical structure. To address this, we can employ a cross-edge grouping method to conceal this characteristic. Specifically, we first construct a dominance tree for the original function. We then group the cross-edge nodes in the region that have a dominance relationship into groups, with each group corresponding to an independent branch selection block. Finally, we insert multiple branch selection basic blocks into the function funca to avoid the topological characteristics caused by a single dominating node.
[0142] like Figure 8 As shown, in multiple functions func m func n ...select a subgraph G′ in each m ,G′ n ,..., then merge these subgraphs to construct a function func mn The implementation process is similar to FunSp, except that the subgraph G′ needs to be m ,G′ n ...perform data flow analysis and control flow analysis as a whole. Based on the splitting of a single function, the calling relationships between sub-functions are simple. Mixed splitting of multiple functions will make the call graph between functions present a mesh structure, increasing the complexity of the program.
[0143] Furthermore, based on the above method, an embodiment of the present invention also provides a code obfuscation system based on function splitting, comprising: a division module and an obfuscation module, wherein:
[0144] A partitioning module is used to construct a control flow graph of the original function in the source code to be processed, and divide the control flow graph into several subgraphs according to the dominance relationship between the nodes in the control flow graph. The control flow graph and subgraphs are composed of code basic blocks;
[0145] The obfuscation module is used to obtain the set of variables with data dependency relationships between the functions corresponding to the subgraph based on the subgraph basic block path and variables, uniformly identify the subgraph entry nodes according to the subgraph control flow, and rewrite the original function code into sub-function codes corresponding to several subgraphs based on the set of variables with data dependency relationships between functions and the unified identification, so as to transform the control flow relationship within the original function into a direct and / or indirect function call relationship between several sub-functions.
[0146] To verify the effectiveness of this solution, the following experimental data is used to further explain it:
[0147] The FunSp algorithm, a proposed obfuscation algorithm, was used in obfuscation experiments on specific code examples. The experimental results show that FunSp significantly reduces the recall rate of mainstream detection models (Recall@1 down to 10%). Compared to traditional obfuscation tools such as O-LLVM, FunSp achieves lower performance overhead (12.7%) and a smaller code expansion ratio (2.06x) at the same level of feature perturbation.
[0148] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0149] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0150] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.
[0151] Those skilled in the art will appreciate that all or part of the steps in the above method can be performed by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disk. Alternatively, all or part of the steps in the above embodiment can be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or software functional modules. The present invention is not limited to any specific combination of hardware and software.
[0152] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A code obfuscation method based on function splitting, characterized in that: Include: Constructing a control flow graph of the original function in the source code to be processed, and dividing the control flow graph into a plurality of subgraphs according to the dominance relationship between nodes in the control flow graph, wherein the control flow graph and the subgraphs are composed of code basic blocks; Based on the subgraph basic block path and variables, a set of variables with data dependency relationships between functions corresponding to the subgraph is obtained, the subgraph entry nodes are uniformly identified according to the subgraph control flow, and the original function code is rewritten into sub-function codes corresponding to several subgraphs based on the set of variables with data dependency relationships between functions and the unified identification, so as to transform the control flow relationship within the original function into a direct and / or indirect function call relationship between several sub-functions.
2. The code obfuscation method based on function splitting according to claim 1 is characterized in that: Construct the control flow graph of the original function in the source code to be processed, including: Obtain a basic block set based on the continuous branchless code in the original function of the source code to be processed, and obtain an edge set based on the control flow transfer relationship between the basic blocks; A directed graph describing the control flow graph of the original function is constructed using a basic block set and an edge set.
3. The code obfuscation method based on function splitting according to claim 1 or 2, characterized in that: The control flow graph is divided into several subgraphs according to the dominance relationship between the nodes in the control flow graph, including: Constructing a control flow graph dominance tree based on dominance relationships between nodes, wherein the dominance relationships are used to describe the situation where all paths from the control flow graph entry to the basic block pass through other basic blocks; Based on hierarchical progression, the control flow graph dominator tree is traversed in post-order and cohesive regions are generated. The control flow graph is then divided into several subgraphs according to the cohesive regions.
4. The code obfuscation method based on function splitting according to claim 3 is characterized in that: Perform post-order traversal of the control flow graph dominator tree based on hierarchical progression and generate cohesive regions, including: Set the dominating subtree size threshold and subgraph balance ratio; Traverse the control flow graph dominator tree nodes from the bottom up, recursively select nodes with dominance relationships to construct subgraphs, and determine whether the size of each node subtree is greater than the dominator subtree size threshold. If so, recursively divide the current node subtree. Otherwise, merge the current node subtree into a region and add it to the region set R. Detect the natural loop set in the node subtree control flow graph and record the loop head node and loop body of each loop. If the region contains the loop head node, merge the corresponding loop body into the corresponding region of the node subtree; Candidate regions are selected from the region set, and several subgraphs are generated according to the subgraph balance rate.
5. The code obfuscation method based on function splitting according to claim 1 is characterized in that: According to the subgraph basic block path and variables, the variable set with data dependency relationship between the corresponding functions of the subgraph is obtained, including: Compare the value sets of each variable in the subgraph at the program point before the head node and the program point after the tail node to obtain the program state difference, and obtain the cross-function transfer variable set based on the forward call data flow and reverse return data flow of the cross edges connecting each subgraph, where the program point before the head node and the program point after the tail node are two program points in the original function execution path obtained based on the cross edges of the subgraph; According to the program state difference and the variable set transferred across functions, a variable set with data dependency between the corresponding functions of the subgraph is obtained. The variable set with data dependency is a variable set based on function calls and returns.
6. The code obfuscation method based on function splitting according to claim 1 is characterized in that: Uniformly identify subgraph entry nodes based on subgraph control flow, including: Assign a unique identifier to the entry node of each subgraph, and set a corresponding identifier at the subgraph exit node. The entry node of the subgraph is the target node of all edges entering the subgraph from other areas, and the subgraph exit node is the source node of all edges leaving the subgraph to other subgraphs.
7. The code obfuscation method based on function splitting according to claim 1 is characterized in that: Based on the variable set and unified identification of data dependency relationships between functions, the original function code is rewritten into sub-function codes corresponding to several subgraphs, including: Set the structure type and structure variables based on the variable set with data dependency relationship between the corresponding functions of the subgraph and the unified identification of the subgraph entry node; In the original function execution path and at the head node and tail node of the subgraph intersection edge, code for structure variable assignment / removal operations and function call / return code are generated, and a switch basic block for selecting the target branch based on a unified identifier is set at the entry and call return of each subfunction.
8. The code obfuscation method based on function splitting according to claim 1, characterized in that: Based on the variable set and unified identification of data dependencies between functions, the original function code is rewritten into sub-function codes corresponding to several subgraphs, which also includes: Traverse all basic blocks in the original function and merge all exit basic blocks; randomly generate a new basic block as the unified exit node, replace the return instruction with a jump instruction, so that the control flow graph has only one exit basic block; And / or, convert the Phi nodes to memory-based operations and fix the escaped variables through memory operations.
9. The code obfuscation method based on function splitting according to claim 1, characterized in that: Based on the variable set and unified identification of data dependencies between functions, the original function code is rewritten into sub-function codes corresponding to several subgraphs, which also includes: The rewritten sub-function is further split and rewritten, and the sub-function call is triggered by a triggering method, wherein the triggering method includes function pointer triggering, signal parameter triggering or exception triggering; And / or, when dividing the control flow graph into several subgraphs, the subgraph cross-edge nodes with a dominating relationship in the subgraph partition area are further divided into a group, each group corresponds to an independent branch selection block, and multiple branch selection basic blocks are inserted into the corresponding sub-function of the subgraph to hide the upper and lower hierarchical features of the subgraph; And / or, a subgraph is selected from the continued sub-function splitting and rewriting of each subgraph sub-function, the selected subgraphs are merged and a new function is constructed, so that the call graph between the original functions presents a mesh structure through multi-subfunction mixed splitting.
10. A code obfuscation system based on function splitting, characterized in that: Contains: partition module and obfuscation module, among which, A partitioning module is used to construct a control flow graph of the original function in the source code to be processed, and divide the control flow graph into several subgraphs according to the dominance relationship between the nodes in the control flow graph. The control flow graph and subgraphs are composed of code basic blocks; The obfuscation module is used to obtain the set of variables with data dependency relationships between the functions corresponding to the subgraph based on the subgraph basic block path and variables, uniformly identify the subgraph entry nodes according to the subgraph control flow, and rewrite the original function code into sub-function codes corresponding to several subgraphs based on the set of variables with data dependency relationships between functions and the unified identification, so as to transform the control flow relationship within the original function into a direct and / or indirect function call relationship between several sub-functions.