Dynamic control flow confusion method based on Gilbreath conjecture
Through the dynamic control flow obfuscation method based on the Gilbreath conjecture, opaque predicates and flattened control flow are constructed, which solves the problem of weakening effectiveness of the existing technology when facing modern reverse analysis tools, and achieves stronger analysis resistance and code protection effects.
Patent Information
- Application Number
- CN202510361171.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
Smart Images

Figure CN120217329A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software protection, and specifically to a dynamic control flow obfuscation method based on Gilbreath's conjecture. Background Art
[0002] With the rapid development of information technology, software security and intellectual property protection have become important issues that need to be solved urgently; as an effective protection means, code obfuscation technology has been widely used in commercial software, digital rights management (DRM) and other fields by complicating the syntax structure, control flow and data flow of programs to improve the unreadability and unpredictability of the code, so as to resist reverse analysis and malicious tampering.
[0003] However, traditional code obfuscation technologies mainly rely on simple transformations at the static level, such as variable renaming, inserting invalid code, etc.; these methods had a certain interference effect on manual reverse analysis and basic static analysis tools in the early stage, but their effectiveness has been greatly weakened in the face of modern efficient automated reverse analysis tools.
[0004] Current mainstream obfuscation methods expose various defects when facing advanced reverse engineering technologies, including: the control flow structure is easy to recover, false branches can be easily identified, the difficulty of path exploration is insufficient, etc.; these problems make the core algorithms and sensitive data still be efficiently parsed and restored under complex analysis tools, and cannot provide sufficient security and obfuscation strength.
[0005] Therefore, there is an urgent need for a more complex and highly unpredictable obfuscation method, which combines advanced strategies such as dynamic assignment, random path selection, opaque predicate construction, etc., to enhance the complexity of the control flow and data flow of the program, ensure the concealment of the program logic and the non-derivability of the execution path, so as to achieve stronger anti-analysis ability and code protection effect. Summary of the Invention
[0006] The purpose of the present invention is to propose a dynamic control flow obfuscation method based on Gilbreath's conjecture for problems such as the program being easily parsed and restored under reverse analysis means such as static analysis, dynamic debugging and symbolic execution, low obfuscation strength of the code and weak complexity of the execution path.
[0007] To achieve the above purpose, the present invention adopts the following technical solutions: a dynamic control flow obfuscation method based on Gilbreath's conjecture, including the following steps:
[0008] Step S1: Based on Gilbreath's conjecture, combined with dynamic factors of environmental variables, construct an opaque predicate that satisfies the two-state logic characteristic;
[0009] Step S2: Analyze the control flow structure of the source program, construct interference code with the same programming language, the same syntax rules, and compatible with the source program code logic as the source program, and split the two to obtain basic blocks;
[0010] Step S3: Flatten the control flow, insert the designed opaque predicates into the control flow, and design a dynamic branch assignment mechanism.
[0011] Specifically, the specific process of Step S1 is as follows:
[0012] Step S11: Gilbreath's conjecture: For any sequence of prime numbers arranged in ascending order, calculate the differences between adjacent numbers in turn, and repeat this process. The first term of each row is always 1; Use the dynamic prime number generation algorithm PGEN(N) to generate a sequence of prime numbers P not exceeding N, and set P = {p1, p2, p3,..., p n} arranged in ascending order, where N is obtained from user input, and p i represents the i-th prime number;
[0013] Step S12: Based on the obtained sequence of prime numbers P, define the difference sequence of the k-th layer as D k , where D0 = P0; The first-layer difference sequence D1 is obtained by subtracting adjacent elements of the prime number sequence pairwise:
[0014] D1 = {p2 - p1, p3 - p2,..., p n - p n-1};
[0015] Step S13: Continue to perform the same operation on D1 to calculate the difference sequences of higher levels; The k-layer difference sequence is defined as: D k [i] = |D k-1 [i + 1] - D k-1 [i]|, i = 1, 2,..., n - k; This process continues until the sequence only has one element left;
[0016] Step S14: After the iterative calculation is completed, verify the last element and combine it with the environment binding strategy to ensure that the program can only run in the specified environment. Specifically, first calculate the hash value H(env_key) based on the environmental characteristics, which is jointly determined by the hardware fingerprint and the system configuration information, that is, H(env_key) = Hash(hardware fingerprint || system configuration), and thus an opaque predicate can be constructed When the program is executed, it will check whether the last element of the final difference sequence is equal to 1, and further bind it to H(env_key) to ensure the uniqueness of the program running environment.
[0017] Specifically, the specific process of Step S2 is as follows:
[0018] Step S21: Traverse the source program using the abstract syntax tree structure to generate the abstract control flow graph of the source program, and identify and record the branch nodes, loop nodes, function call nodes, and jump instruction nodes in the source program;
[0019] Step S22: Cut the source program with key nodes such as the program entry, branch, loop, function call, jump, and exit as boundaries to form multiple correct basic blocks, and record the number of generated correct basic blocks as n;
[0020] Step S23: Construct interference code that meets the conditions and perform the same cutting operation as S21 and S22 to obtain multiple interference basic blocks.
[0021] Specifically, the specific process of step S3 is as follows:
[0022] Step S31: Define the types of basic blocks, mark the own attributes (AT) of different basic blocks, including two types. The attribute of the normal basic block is N, and the attribute of the interference basic block is D. At the same time, label all basic blocks;
[0023] Step S32: Define the basic block quadruple, mark the path information of different basic blocks, which is composed of the attribute (AT), its own basic block (BS), the predecessor basic block (BP), and the successor basic block (BN); when a basic block has no predecessor or successor, set the corresponding label to empty. If there is a predecessor or successor, record the label of the corresponding basic block;
[0024] Step S33: Create a dispatcher with a switch structure, flatten the ordered control flow and store it in an unordered manner, and implement the combination of normal basic blocks and interference basic blocks through the Gilbreath_OP constructed in S1;
[0025] Step S34: When the program executes, the execution path NP of the basic block is determined by the quadruple defined in step S32. After the current basic block is executed, the execution path NP will determine the next basic block to execute according to the attribute of the current basic block and the label of the successor basic block;
[0026] Step S35: Dynamically assign NP, define a F rand-switch Random jump function; according to the total number of current basic blocks n obtained in step S22, determine the next target basic block by generating a random number nB = rand() mod n;
[0027] Step S36: The verification rules for the generated random number are as follows: 1): The basic block label corresponding to nB must be consistent with the successor basic block label of the current basic block. 2) The attribute (AT) of the basic block represented by nB must be consistent with the attribute (AT) of the current basic block. If the verification rules are not satisfied, a random number is regenerated until a match is successful and then the corresponding basic block is executed;
[0028] Step S37: This process is looped until there is no successor basic block for the current basic block, at which point the program execution path ends and all basic blocks have been executed.
[0029] Compared with the prior art, the beneficial effects of the present invention are:
[0030] By constructing opaque predicates based on the Gilbreath conjecture and environmental variables, the execution path of the program will depend on dynamic factors, increasing the randomness and unpredictability of the path, enhancing the resistance to analysis methods such as static analysis and symbolic execution, preventing analysis tools from accurately restoring the program logic, and improving the security of the program;
[0031] By constructing interference code compatible with the source program and splitting it from the source program, multiple interference paths can be effectively generated, making the control flow highly confusing. Through the interference code constructed based on the source program, although it has the normal execution function, its control flow and logical structure are cleverly designed to generate a large number of pseudo-target paths in a manner similar to the real program path, thus greatly enhancing the misleading nature of the program and misleading analysis tools and attackers to distinguish the real path, preventing the accurate restoration of the program logic;
[0032] By introducing control flow flattening and random assignment mechanisms, the randomness of the program execution path is greatly increased, avoiding the restoration of program logic by static analysis tools through pattern matching and path tracking. It improves the program's defense capabilities against reverse engineering and dynamic debugging, effectively prevents common control flow recovery techniques, and increases the time and resources required for reverse analysis. Description of the Drawings
[0033] Figure 1 It is a flowchart of the flattening control flow obfuscation method based on the Gilbreath conjecture in this embodiment;
[0034] Figure 2 It is the control flowchart of step S21 in this embodiment;
[0035] Figure 3 It is the basic block split in step S22 of this embodiment;
[0036] Figure 4 It is the basic block split by the interference code constructed in step S23 of this embodiment;
[0037] Figure 5 The quadruple of the correct code constructed for step S32 of this embodiment case;
[0038] Figure 6 The quadruple of the interference code constructed for step S32 of this embodiment case;
[0039] Figure 7 Create a distributor for step S33 of this embodiment case to implement a flattened schematic diagram;
[0040] Figure 8 For step S35 of this embodiment case, implement NP from F rand_switch The flattened schematic diagram after the random jump function;
[0041] Figure 9 The comparison chart of the execution time of the embodiment (test - GCFF) of the present invention, the original program (test), and the ollvm obfuscation scheme (test - ollvm);
[0042] Figure 10 The comparison chart of the number of basic blocks of the embodiment (test - GCFF) of the present invention, the original program (test), and the ollvm obfuscation scheme (test - ollvm):;
[0043] Figure 11 The comparison chart of the program sizes of the embodiment (test - GCFF) of the present invention, the original program (test), and the ollvm obfuscation scheme (test - ollvm). Detailed implementation manners
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. For the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] Next, the embodiments of the present invention will be further described in detail in conjunction with the accompanying drawings.
[0046] Step S1: Based on the Gilbreath conjecture, combined with the dynamic factors of environmental variables, construct an opaque predicate that satisfies the two - state logic characteristics;
[0047] In this step, it can be divided into the following 4 steps for implementation:
[0048] Step S11: Generate a prime number sequence P not exceeding N using the dynamic prime number generation algorithm PGEN(N): Assume the parameter N = 20 input by the user, and use the dynamic prime number generation algorithm PGEN(N) to obtain the prime number sequence not exceeding N = 20: P = {2, 3, 5, 7, 11, 13, 17, 19};
[0049] Step S12: Based on the obtained prime number sequence P, calculate the first-layer difference sequence D1. The calculation method of this sequence is to subtract adjacent elements of P pairwise; the first-layer difference sequence is:
[0050] D1 = {3 - 2, 5 - 3, 7 - 5, 11 - 7, 13 - 11, 17 - 13, 19 - 17} = {1, 2, 2, 4, 2, 4, 2};
[0051] Step S13: On the basis of D1, continue to iteratively calculate higher-level difference sequences until the final difference sequence is calculated; the k-layer difference sequence is defined as:
[0052] D k [i] = |D k-1 [i + 1] - D k-1 [i]|, i = 1, 2,..., n - k; the obtained second-layer difference sequence: D2 = {1, 0, 2, 2, 2, 2}, the third layer: D3 = {1, 2, 0, 0, 0}, until the last-layer difference sequence is: D7 = {1};
[0053] Step S14: After the iterative calculation is completed, verify the last element and combine the environment binding strategy to ensure that the program can only run in the specified environment. Specifically, first calculate the hash value H(env_key) based on the environmental characteristics, which is jointly determined by the hardware fingerprint and system configuration information, that is, H(env_key) = Hash(hardware fingerprint || system configuration), and thus an opaque predicate can be constructed When the program is executed, it will check whether the last element of the final difference sequence is equal to 1 and further bind it with H(env_key) to ensure the uniqueness of the program running environment.
[0054] Step S2: Analyze the control flow structure of the source program, construct interference code with the same programming language, the same syntax rules, and compatible with the source program code logic as the source program, and split the two to obtain basic blocks;
[0055] In this step, it can be carried out in the following 3 steps:
[0056] Step S21: Traverse the source program using the abstract syntax tree structure to generate the abstract control flow graph of the source program, and identify and record the branch nodes, loop nodes, function call nodes, and jump instruction nodes in the source program, specifically as follows Figure 2 shown;
[0057] Step S22: Cut the source program at the key nodes such as the program entry, branch, loop, function call, jump, and exit to form multiple correct basic blocks, as follows Figure 3 shown;
[0058] Step S23: Construct interference code that meets the conditions, and perform the same cutting operation in the manner of S21 and S22 to obtain multiple interference basic blocks, as follows Figure 4 shown.
[0059] Step S3: Flatten the control flow, insert the designed opaque predicates into the control flow, and design a dynamic branch assignment mechanism.
[0060] In this step, it can be divided into the following 7 steps:
[0061] Step S31: Classify all basic blocks, define the types of basic blocks, and mark the own attributes of different basic blocks, including two types. The attribute of the normal basic block is N, and the attribute of the interference basic block is D. At the same time, label all basic blocks for subsequent control flow management;
[0062] Step S32: Define the basic block quadruple to record the path information of all basic blocks. This quadruple consists of the attribute (AT), its own basic block (BS), the predecessor basic block (BP), and the successor basic block (BN); when a basic block has no predecessor or successor, set the corresponding label to empty. If there is a predecessor or successor, record the label of the corresponding basic block, as follows Figure 5 , Figure 6 shown;
[0063] Step S33: Create a dispatcher with a switch structure to flatten the ordered control flow and store it in an unordered manner, and implement the combination of normal basic blocks and interference basic blocks through the opaque predicate Gilbreath_OP constructed in S1, as follows Figure 7 shown;
[0064] Step S34: When the program is executed, the execution path NP of the basic block is determined by the quadruple defined in Step S32. After the current basic block is executed, the execution path NP will determine the next basic block to be executed according to the attribute of the current basic block and the label of the successor basic block;
[0065] Step S35: Dynamically assign NP, define an F rand-switchRandom jump function; According to step S21, obtain the total number of current basic blocks n, and determine the next target basic block by generating a random number nB = rand() mod n, as Figure 8 shown;
[0066] Step S36: Check the generated random number, and the check rules are as follows: 1): The basic block label corresponding to nB must be the same as the successor basic block label of the current basic block. 2) The attribute (AT) of the basic block represented by nB must be the same as the attribute (AT) of the current basic block. If the check rules are not met, generate a random number again until the match is successful and then execute the corresponding basic block;
[0067] Step S37: This process is looped until there is no successor basic block for the current basic block, at which point the program execution path ends and all basic blocks have been executed.
[0068] From Figure 9 , Figure 10 and Figure 11 it can be seen that with the enhancement of the obfuscation strategy, this embodiment (test-GCFF) significantly exceeds OLLVM in terms of program volume and the number of basic blocks; the test-GCFF scheme increases the program volume by 78.6% compared to the original program, far exceeding 35.7% of test-ollvm. At the same time, the number of basic blocks increases from 4 to 43, demonstrating a stronger ability to complicate the code structure. The increase in volume and the number of basic blocks mainly stems from the fact that this embodiment adopts a more complex code organization method. By introducing a large number of interfering basic blocks, the control flow graph (CFG) shows a high degree of non-linearity, increasing the difficulty for analysis tools to identify valid paths; in addition, combined with dynamic assignment calculation and opaque predicates, the uncertainty of the path is further enhanced, making it difficult to restore the execution logic; although the obfuscation strategy improves the code complexity, this embodiment still maintains a reasonable computational overhead, avoiding excessive performance loss, and ensuring the efficient operation of the program while enhancing security. The experimental results show that this embodiment has significant advantages in balancing code complexity and execution efficiency.
[0069] The above is only the preferred embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the scope disclosed by the present invention, according to the technical solution and concept of the present invention, makes equivalent substitutions or changes, all belong to the protection scope of the present invention.
Claims
1. A dynamic control flow obfuscation method based on Gilbreath conjecture, characterized in that: The following steps are involved: Step S1: Based on Gilbreath conjecture and combined with the dynamic factors of environmental variables, an opaque predicate that satisfies the binary logic characteristics is constructed; Step S2: parsing the control flow structure of the source program, constructing interference code with the same programming language and grammatical rules as the source program and logically compatible with the source program code, and dividing the two to obtain basic blocks; Step S3: Flatten the control flow, insert the designed opaque predicate into the control flow, and design a dynamic branch assignment mechanism.
2. A dynamic control flow obfuscation method based on Gilbreath conjecture according to claim 1, characterized in that: The specific process of step S1 is as follows: Step S11: Gilbreath conjecture: For any prime number sequence arranged from small to large, calculate the difference between adjacent numbers in turn, repeat this process, and the first item of each row is always 1; use the dynamic prime number generation algorithm PGEN(N) to generate a prime number sequence P not exceeding N, let P = {p1, p2, p3, ..., p n } is arranged in ascending order, where N is obtained by user input, p i represents the i-th prime number; Step S12: Based on the obtained prime number sequence P, define the difference sequence of the kth layer as D k , where D0 = P0; the first-level difference sequence D1 is obtained by subtracting adjacent elements of the prime number sequence: D1={p2-p1,p3-p2,…,p n -p n-1 }; Step S13: Continue to perform the same operation on D1 to calculate a higher level difference sequence; the k-level difference sequence is defined as: D k [i]=|D k-1 [i+1]-D k-1 [i]|, i = 1, 2, ..., nk; the process continues until there is only one element left in the sequence; Step S14: After the iterative calculation is completed, the last element is verified and combined with the environment binding strategy to ensure that the program can only run in the specified environment. Specifically, the hash value H(env_key) based on the environment characteristics is first calculated. The value is determined by the hardware fingerprint and the system configuration information, that is, H(env_key) = Hash(hardware fingerprint||system configuration), thereby constructing an opaque predicate When the program is executed, it checks whether the last element of the final difference sequence is equal to 1, and further binds it to H (env_key) to ensure the uniqueness of the program running environment.
3. A dynamic control flow obfuscation method based on Gilbreath conjecture according to claim 1, characterized in that: The specific steps of step S2 are as follows: Step S21: traverse the source program using the abstract syntax tree structure to generate an abstract control flow graph of the source program, and identify and record branch nodes, loop nodes, function call nodes, and jump instruction nodes in the source program; Step S22: Cut the source program based on key nodes such as program entry, branch, loop, function call, jump and exit as boundaries to form multiple correct basic blocks, and record the number of generated correct basic blocks as n; Step S23: construct interference codes that meet the conditions and perform the same cutting operations as S21 and S22 to obtain multiple interference basic blocks.
4. The method for dynamic control flow obfuscation based on Gilbreath conjecture according to claim 1, characterized in that: The specific steps of step S3 are as follows: Step S31: define the type of basic blocks, mark the own attributes (AT) of different basic blocks, including two types, the attribute of normal basic blocks is N, the attribute of interference basic blocks is D, and label all basic blocks at the same time; Step S32: define a basic block quadruple, mark the path information of different basic blocks, and consist of the attribute (AT), the basic block itself (BS), the predecessor basic block (BP) and the successor basic block (BN); when the basic block has no predecessor or successor, set the corresponding label to null, if there is a predecessor or successor, record the label of the corresponding basic block; Step S33: Create a switch structure distributor to flatten the ordered control flow into an unordered storage, and realize the combination of normal basic blocks and interference basic blocks through Gilbreath_OP constructed in S1; Step S34: When the program is executed, the execution path NP of the basic block is determined by the four-tuple defined in step S32. After the current basic block is executed, the execution path NP will determine the next basic block to be executed according to the attributes of the current basic block and the label of the subsequent basic block; Step S35: Dynamically assign NP and define an F rand-switch Random jump function; according to step S22, the total number of current basic blocks n is obtained, and the next target basic block is determined by generating a random number nB=rand()modn; Step S36: The generated random number is verified according to the following rules: 1) the basic block number corresponding to nB must be consistent with the basic block number following the current basic block; 2) the attribute (AT) of the basic block represented by nB must be consistent with the attribute (AT) of the current basic block. If the verification rules are not met, a random number is regenerated until the corresponding basic block is executed after a match is successful. Step S37: This process is repeated until the current basic block has no successor basic block. At this time, the program execution path has ended and all basic blocks have been executed.