An llvm-based cross-function state entanglement algorithm

CN122818364APending Publication Date: 2026-09-25SHIJIAZHUANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610986486.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

1.防御孤岛问题:由于混淆局限于单函数边界,整个程序的防御呈现出孤岛状态

Benefits of technology

1.打破防御孤岛,实现全局防护:通过全局状态变量将多个原本独立的函数在执行语义上进行强关联,攻击者无法单独分析某一个函数,必须理清整个程序的状态演变过程,显著提升整体逆向难度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122818364A_ABST
    Figure CN122818364A_ABST
Patent Text Reader

Abstract

The application discloses a cross-function state entanglement algorithm based on LLVM, creates a global state variable in an LLVM intermediate representation layer, and inserts a memory loading instruction at the entrance of each target function to obtain a current global state value. Each target instruction in the function is traversed, an entanglement factor is generated according to the current state, the original instruction is replaced by an instruction converted by mixed Boolean arithmetic (MBA), a semantically equivalent but highly nonlinear transformation is realized. Then, the global state value is updated according to the current state and the instruction calculation result, and the global variable is written back. Different functions form a state dependence chain through reading and writing operations on the same global state variable, so that any local tampering causes an avalanche effect to cause all subsequent execution output error results instead of direct crash, and a silent failure mechanism is realized. The application breaks the defense island of traditional single-function confusion, and significantly improves the anti-symbol execution capability and anti-tampering capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of software security technology, and more specifically to a cross-function state entanglement algorithm based on LLVM. Background Technology

[0002] Traditional code obfuscation techniques, such as control flow flattening and spoof control flow, mostly strictly limit the scope of their transformations to a single function. These techniques increase the difficulty of analyzing the logic inside a function by using mechanisms such as function dispatchers and opaque predicates.

[0003] Objective disadvantages: 1. The problem of isolated defenses: Because obfuscation is limited to single function boundaries, the defense of the entire program appears isolated. Attackers can use automated reverse engineering tools such as symbolic execution and pollution analysis to independently disassemble and crack each function without understanding the complete program context.

[0004] 2. Weak resistance to automated analysis: For advanced analysis engines such as symbolic execution, the solution space for path constraints generated by single-function obfuscation is relatively small. Solvers can complete the solution of a single obfuscated function logic in a short time, typically within minutes, leading to the failure of protection.

[0005] 3. Lack of global linkage: The execution logic of different functions is independent of each other, and no causal dependency is formed. After an attacker successfully cracks a function, it will not affect the execution of other parts of the program, and a global protective barrier cannot be formed.

[0006] 4. Vulnerable to silent tampering: Traditional hardening programs often crash directly when illegally tampered with. This explicit failure mode is easily located by attackers and used as a breakthrough point for reverse engineering.

[0007] Therefore, how to provide a cross-function state entanglement algorithm based on LLVM to solve the problems of defense silos, weak resistance to automated analysis, and lack of global linkage in existing code obfuscation techniques, and significantly improve the software's ability to resist reverse engineering and automated attacks, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0008] This invention proposes a cross-function state entanglement algorithm based on LLVM, which expands the code defense dimension from a single function to the entire program. Specifically, it includes the following technical solutions: Create a global state variable that is visible to the entire program in the LLVM intermediate presentation layer, and set the initial value of the global state variable; Iterate through each target function in the program to be protected, insert a memory load instruction at the entry point of the target function, and load the current global state value from the global state variable; Iterate through each target instruction within the target function and perform the following operations on each target instruction: Generate an entanglement factor based on the current global state value; The target instruction is replaced with an instruction after a hybrid Boolean arithmetic transformation, wherein the hybrid Boolean arithmetic transformation uses the entanglement factor to perform an identity transformation on the original result of the target instruction; Calculate the next global state value based on the current global state value and the calculation result of the target instruction; Insert a memory storage instruction to write the next global state value back to the global state variable; In this context, instructions in different functions form cross-function state dependency chains through read and write operations on the global state variables.

[0009] Preferably, the global state variable is created through the LLVM's Module::getOrInsertGlobal interface, is of type 32-bit integer, and its link type is set to external link attribute through the GlobalValue::setLinkage method to prevent the compiler optimizer from optimizing it away.

[0010] Preferably, the initial value of the global state variable is set to a hexadecimal feature code.

[0011] Preferably, the entanglement factor is generated using the following nonlinear function: Z = (S_cur × 0x41C64E6D + X) ⊕ Salt Among them, the multiplication operator 0x41C64E6D fits the Hull-Dobell theorem, S_cur is the current global state variable, X is the state update constant, Salt is a random salt value, and ⊕ represents the XOR operation.

[0012] Preferably, the hybrid Boolean arithmetic transformation is selected from at least one of the following identity strategies: .

[0013] Preferably, the calculation of the next global state value uses the following nonlinear function: S_next = (S_cur ⊕ Result) × X + Y Where S_next is the updated global state variable, S_cur is the current global state variable, Result is the calculation result of the target instruction, ⊕ is the XOR operator, and X and Y are state update constants.

[0014] Preferably, when the target instruction is replaced with an instruction after mixed Boolean arithmetic transformation, an isolation-based proxy approach is adopted, using an instruction pointer set to track all newly generated intermediate computing nodes, and when performing usage redirection, the instructions inside the obfuscated logic can continue to reference the original nodes, while the obfuscated result is only applied to subsequent logic outside the isolation zone.

[0015] Preferably, prior to the hybrid Boolean arithmetic transformation, the following linearly decoupled scheduling order is followed: Inject opaque predicates to split basic blocks; Implement control flow flattening to reshape the function topology; Perform hybrid Boolean arithmetic transformations based on global state entanglement.

[0016] Preferably, the cross-function state dependency chain causes any modification to local logic of the program to trigger an avalanche effect in the global state evolution, resulting in all subsequent execution flows outputting incorrect results instead of directly crashing, thus achieving a silent failure mechanism.

[0017] Preferably, the target instruction is a binary arithmetic instruction or a binary logic instruction, including addition instructions, subtraction instructions, multiplication instructions, XOR instructions, OR instructions, and AND instructions.

[0018] Compared with existing technologies, the cross-function state entanglement algorithm based on LLVM disclosed in this invention has the following beneficial effects: 1. Break down defense silos and achieve global protection: By using global state variables, multiple originally independent functions are strongly correlated in terms of execution semantics. Attackers cannot analyze a single function in isolation and must understand the state evolution process of the entire program, which significantly increases the overall difficulty of reverse engineering.

[0019] 2. Effectively resist symbolic execution attacks: By introducing nonlinear logic of global state evolution and MBA transformation into the path constraints, the constraint space faced by the symbolic execution engine expands exponentially, causing path explosion, thus making it impossible to solve the core logic within the effective time.

[0020] 3. Enhanced resistance to tampering: The silent failure mechanism ensures that after an attacker or malicious software modifies the code, the program will not exit abnormally but will continue to run and output incorrect results. This covert failure method greatly increases the difficulty for attackers to determine whether the tampering was successful, and can be used to build high-strength software watermarks or integrity verification.

[0021] 4. Optimized balance between security and performance overhead: Experimental data shows that with an average increase of 142.55% in the number of effective instructions, the program execution time overhead only increases by an average of 118.51%, and the average growth rate of binary file size is only 7.15%, proving that this patent significantly improves security while keeping the impact on performance controllable.

[0022] 5. High compilation stability: A memory synchronization-driven SSA (Static Single Assignment) control repair mechanism is designed, which replaces virtual register transfer with Load / Store operations, effectively avoiding compiler validity check errors that may be caused by complex control flow transformations. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the cross-function entanglement principle of the present invention; Figure 2 This is a flowchart illustrating the state entanglement and MBA transition of a single instruction in this invention; Figure 3 The images show the normal result (left) and the modified result (right) of the RC4 function in this invention. Detailed Implementation

[0024] The technical solutions of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described technical solutions are only a part of the present invention, and not all of the technical solutions and measures. All other technical solutions obtained by those skilled in the art based on the implementation of the technical solutions of the present invention without creative effort are within the scope of protection of the present invention.

[0025] Example 1

[0026] This embodiment provides a cross-function state entanglement algorithm based on LLVM. The algorithm processes C / C++ source code and is implemented in the intermediate representation (IR) layer of the LLVM compiler. The entire algorithm is integrated into the compilation process as a custom LLVM Pass.

[0027] Step 1: Environment Initialization and Global State Machine Construction First, initialization is performed in the `runOnModule` method of the LLVM Pass. In this embodiment, to eliminate the variable dominance boundary failure problem that may be caused by transformations such as control flow flattening, the first step of the linear decoupling scheduling mechanism—environment initialization—is executed. This step maps all virtual register variables in SSA (Static Single Assignment) form back to the memory stack space and clears their lifetime markers.

[0028] Subsequently, a global state machine visible to the entire program is constructed.

[0029] By calling the Module::getOrInsertGlobal interface, a new global variable, denoted as G_State, is created within the module scope of LLVM IR.

[0030] The variable's type is set to `IntegerType::getInt32Ty(Context)`, a 32-bit integer, to match common processor word lengths. To prevent the compiler optimizer from incorrectly optimizing this variable away during dead code elimination, its linker type is set to `GlobalValue::ExternalLinkage` using the `GlobalValue::setLinkage` method. This causes the optimizer to recognize that the variable might be modified by unknown external symbols at runtime, thus forcing the preservation of all memory read and write operations targeting this variable. Finally, the initial value of this global variable is set to a specific hexadecimal signature, such as 0x5CDC, as the starting point for program logic transitions.

[0031] Step 2: Traversal of the objective function and state anchoring Next, Pass iterates through all functions in the program to be protected. For each target function that needs to be obfuscated, state read logic is inserted at its entry point. Specifically, an IRBuilder (instruction builder) is used to create a memory load instruction LoadInst before the first valid instruction of the function. This loads the current global state value, denoted as S_cur, from the memory address of the global variable G_State. This value will serve as the dynamic seed for all instruction obfuscation operations within the function.

[0032] Step 3, Obfuscation of target instructions This embodiment uses a binary arithmetic instruction within a function (e.g., the add instruction) as an example to describe in detail the obfuscation process for a single instruction. For logical operations (such as xor, or) or other arithmetic instructions, the processing method is similar; only the corresponding MBA identity needs to be adjusted according to the operand type of the instruction. Step S1: Load the current state and obtain the original result For the target instruction I (an add instruction), S_cur is loaded at the function entry point. At the same time, the original computation result Result of instruction I, i.e., the sum of the two operands T and Z, is preserved.

[0033] Step S2: Generate dynamic entanglement factor Based on the current global state S_cur and a preset constant, an entanglement factor Z is generated. In this embodiment, the nonlinear function shown in equation (1) is used for calculation: Z = (S_cur × 0x41C64E6D + X) ⊕ Salt Equation (1) Wherein, the multiplication operator 0x41C64E6D conforms to the Hull-Dobell theorem, ensuring that the global state machine is in a 32-bit integer ring. This allows for full-cycle coverage. Analyzing it from the perspective of bit entropy, the operator has balanced Hamming weights, generating a significant avalanche effect during rolling updates. Combined with an implicit module overflow mechanism, it transforms the originally linear state evolution into nonlinear lattice structure changes in a high-dimensional space, achieving deep interference with the symbolic execution engine at the logic layer. S_cur is the current global state variable, X is the state update constant, and Salt is a random salt value. ⊕ represents the XOR operation. At the LLVM IR level, multiplication, addition, and XOR instructions are created sequentially to calculate Z.

[0034] Step S3: Apply Mixed Boolean Arithmetic (MBA) Transformation This embodiment pre-configures several MBA identity strategies, as shown in Table 1 below. During runtime, one of these strategies is randomly selected to replace the original add instruction.

[0035] Table 1 Hybrid Boolean Identity Strategy

[0036] Taking the "additive recombination transformation" as an example, the obfuscation process is as follows: If we choose Result=(Result⊕Z)+2·(Result∧Z), Then generate Value*Term1=Builder.CreateXor(Result, Z), Value*Term2=Builder.CreateAnd(Result, Z), Value*Term3=Builder.CreateMul(Term2, ConstantInt::get(Int32Ty, 2)), Value * MBA_Result = Builder.CreateAdd(Term1, Term3). This MBA_Result is semantically completely equivalent to the original Result = T + Z, but its instruction representation becomes highly non-linear. Since the entanglement factor Z depends on the rolling updated global state S_cur, the same source code instructions will present completely different instruction sequences in different execution contexts or different obfuscation instances, thus effectively combating pattern matching-based deobfuscation tools.

[0037] Step S4: Nonlinear Global State Update After the original instruction I is replaced, the global state needs to be updated based on the calculation result of the instruction (i.e., MBA_Result). This embodiment uses the nonlinear function shown in Equation (2): S_next = (S_cur ⊕ Result) × X + Y Equation (2) In this equation, the CFSE algorithm abandons the traditional static state dependency and instead carries out the design of a rolling model based on the instruction sequence for updating. Before each target instruction is executed, the current global state must be loaded from memory, and the calculation result and the entanglement factor generated based on the current state are deeply integrated. After the instruction is executed, the global state S_next is updated using a nonlinear function, as shown in Equation (2). This design ensures that any local logical tampering will generate an avalanche effect through the state chain, causing all subsequent execution flows to enter a silent failure state. In this equation, S_next is the updated global state variable. S_cur is the current global state variable. Result is the calculation result of the target instruction. ⊕ is the XOR operator. X and Y are state update constants.

[0038] Step S5: Status Write-back and Instruction Replacement The Builder.CreateStore method creates a memory storage instruction that writes the calculated S_next back to the memory address of the global state variable G_State. Finally, the ReplaceInstWithInst interface is called to replace the original add instruction I with the generated obfuscated instruction tree (i.e., the MBA_Result node).

[0039] Step 4: Command Isolation and SSA Dominion Restoration To prevent unexpected recursive references within the obfuscated logic and avoid unnecessary compilation overhead during the aforementioned replacement process, this embodiment employs a proxy method based on an isolation zone. Specifically, SmallPtrSet is used.<Instruction*, 32> A data structure is used to track all newly generated intermediate computation nodes. When using the replaceUsesWithIf interface for usage redirection, a filtering logic ensures that instructions inside the obfuscated logic can continue to reference the original Result node, but the obfuscated final result MBA_Result is only applied to subsequent logic outside the isolation zone. This guarantees the integrity of the SSA graph during the transformation process.

[0040] Furthermore, since this method does not use virtual registers to pass state, but instead relies on memory Load / Store operations, it fundamentally avoids the SSA dominance constraints of LLVM. Specifically, when passing MBA_Result across basic blocks, it exists as a memory value, which can be obtained anywhere that needs it via the Load instruction without satisfying the dominance rules of virtual registers, thus ensuring compilation stability under complex control flow structures.

[0041] Step 5: Linear Decoupling Scheduling To further improve obfuscation quality and prevent conflicts between different transformation logics, this embodiment follows a linear decoupling scheduling order from local to global: Opaque predicate injection: First, inject opaque predicates that are always true or always false into the function to split the basic block and increase the branch complexity.

[0042] Control flow flattening: Next, the function topology is reshaped by introducing a dispatcher, which turns the original basic blocks into successors of the dispatcher, increasing the difficulty of static analysis.

[0043] Instruction Encryption (MBA Transformation): Finally, the MBA transformation based on global state entanglement described in Part III above is performed to replace arithmetic and logic instructions with a highly nonlinear logic forest.

[0044] Step Six: Silent Failure Mechanism and Effect Verification Through the steps described above, this method constructs a causal dependency chain that runs throughout the entire program. To verify its effectiveness, the classic RC4 encryption algorithm was used as the test object, and dynamic debugging was performed in IDA Pro.

[0045] In the debugger, when the program reaches a certain point, manually modify the value of the global state variable G_State. For example, if its correctly evolved intermediate value is 0x6F85A82, change it to 0x6F85A83. Then let the program continue execution. Experimental results show that the program does not crash immediately, but the output of subsequent encryption or decryption processes is completely incorrect (e.g., ...). Figure 3 As shown, the left side represents the correct ciphertext, and the right side represents the tampered ciphertext. This demonstrates that any local logical tampering (in this case, state value tampering) will trigger an avalanche effect through the state chain, causing all subsequent execution flows to enter a silent failure state, outputting incorrect results instead of directly crashing. This hidden failure mode greatly increases the difficulty of reverse engineering. In summary, the LLVM-based cross-function state entanglement method provided in this embodiment, by establishing a globally shared state machine, an instruction-driven causal logic chain, and an anti-decryption MBA transformation, successfully elevates the code protection dimension from the traditional single function to the entire program scope. It can effectively resist the cracking by automated analysis tools such as symbolic execution and possesses excellent performance and compilation stability.

[0046] The experimental environment will be described below using both hardware and software components. The hardware environment is shown in Table 2 below.

[0047] Table 2 Hardware Environment Configuration

[0048] The system layer uses Linux Ubuntu 24.04, along with the latest LLVM 22 component library, and utilizes the enhanced New PassManager (NPM interface) to enable efficient scheduling of custom passes. The selected software environment is shown in Table 3.

[0049] Table 3 Software Environment Configuration

[0050] Experiments were conducted, including cross-function state entanglement and confusion experiments, including silent failure tests and anti-symbolic execution tests based on angr.

[0051] Experiment on cross-function state entanglement algorithm: To verify the performance before and after obfuscation, the running time, file size, basic blocks and effective instructions of the four modules were compared before and after obfuscation. The dataset is angr_ctf, and the comparison data is shown in Table 4.

[0052] Table 4 Comparison of data before and after confusion

[0053] First, we will verify the silent failure mechanism. This experiment uses the classic RC4 algorithm as the test function. The experiment involves dynamic debugging in IDA Pro and changing global state values ​​to observe whether the program outputs completely incorrect encryption and decryption results without immediately crashing.

[0054] In the experiment, the global state variable The initial value was 0x5CDC. Dynamic debugging revealed that the first calculation was 0x6F85A82. The value was changed to 0x6F85A83, and the result was observed after running the program.

[0055] The normal result of the RC4 function is the left side, and the modified result is the right side, as shown below. Figure 3 As shown.

[0056] Secondly, a comparative experiment on anti-symbolic execution was conducted based on angr. The automated analysis engine angr was used to perform path detection on the obfuscated samples. In single-function protection mode, symbolic execution tools can usually solve logical constraints within minutes. However, after enabling CFSE protection, because path constraints are diffused into the global state evolution across processes, the constraint space faced by the solver expands exponentially, inducing a severe path explosion phenomenon. The test results are shown in Table 5 below.

[0057] Table 5. Experimental results of anti-symbolic execution based on angr.

[0058] Cross-function entanglement defense verification: In security testing, the silent failure mechanism was confirmed. Experiments showed that after tampering with the global state value, the program did not immediately crash, but instead output incorrect results, demonstrating strong stealth capabilities.

[0059] Comparative experiments on anti-symbolic execution based on angr show that the CFSE algorithm causes path explosion in the symbolic execution engine when multiple functions are entangled, making the test samples unsolvable and verifying its resistance to automated analysis.

[0060] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A cross-function state entanglement algorithm based on LLVM, characterized in that, Includes the following steps: Create a global state variable that is visible to the entire program in the LLVM intermediate presentation layer, and set the initial value of the global state variable; Iterate through each target function in the program to be protected, insert a memory load instruction at the entry point of the target function, and load the current global state value from the global state variable; Iterate through each target instruction within the target function and perform the following operations on each target instruction: Generate an entanglement factor based on the current global state value; The target instruction is replaced with an instruction after a hybrid Boolean arithmetic transformation, wherein the hybrid Boolean arithmetic transformation uses the entanglement factor to perform an identity transformation on the original result of the target instruction; Calculate the next global state value based on the current global state value and the calculation result of the target instruction; Insert a memory storage instruction to write the next global state value back to the global state variable; In this context, instructions in different functions form cross-function state dependency chains through read and write operations on the global state variables.

2. The cross-function state entanglement algorithm based on LLVM according to claim 1, characterized in that, The global state variable is created through LLVM's Module::getOrInsertGlobal interface. Its type is a 32-bit integer, and its link type is set to external link attribute through the GlobalValue::setLinkage method to prevent the compiler optimizer from optimizing it away.

3. The cross-function state entanglement algorithm based on LLVM according to claim 1, characterized in that, The initial value of the global state variable is set to a hexadecimal signature.

4. The cross-function state entanglement algorithm based on LLVM according to claim 1, characterized in that, The entanglement factor is generated using the following nonlinear function: Z = (S_cur × 0x41C64E6D + X) ⊕ Salt Among them, the multiplication operator 0x41C64E6D fits the Hull-Dobell theorem, S_cur is the current global state variable, X is the state update constant, Salt is a random salt value, and ⊕ represents the XOR operation.

5. The cross-function state entanglement algorithm based on LLVM according to claim 1, characterized in that, The hybrid Boolean transformation is selected from at least one of the following identity strategies: 。 6. The cross-function state entanglement algorithm based on LLVM according to claim 1, characterized in that, The calculation of the next global state value uses the following nonlinear function: S_next = (S_cur ⊕ Result) × X + Y Where S_next is the updated global state variable, S_cur is the current global state variable, Result is the calculation result of the target instruction, ⊕ is the XOR operator, and X and Y are state update constants.

7. The cross-function state entanglement algorithm based on LLVM according to claim 1, characterized in that, When replacing the target instruction with the instruction after mixed Boolean arithmetic transformation, an isolation-based proxy approach is adopted. The instruction pointer set is used to track all newly generated intermediate computing nodes. When redirecting the purpose, the instructions inside the obfuscated logic can continue to reference the original nodes, while the obfuscated result is applied only to the subsequent logic outside the isolation zone.

8. The cross-function state entanglement algorithm based on LLVM according to claim 1, characterized in that, Prior to performing the hybrid Boolean arithmetic transformation, the following linearly decoupled scheduling order is followed: Inject opaque predicates to split basic blocks; Implement control flow flattening to reshape the function topology; Perform hybrid Boolean arithmetic transformations based on global state entanglement.

9. The cross-function state entanglement algorithm based on LLVM according to claim 1, characterized in that, The cross-function state dependency chain enables any modification to local logic of the program to trigger an avalanche effect in the global state evolution, causing all subsequent execution flows to output incorrect results instead of directly crashing, thus achieving a silent failure mechanism.

10. The LLVM-based cross-function state entanglement algorithm according to claim 1, characterized in that, The target instruction is a binary arithmetic instruction or a binary logic instruction, including addition instructions, subtraction instructions, multiplication instructions, XOR instructions, OR instructions, and AND instructions.