Code virtualization obfuscation method based on reinforcement learning
Through the code virtualization obfuscation method based on reinforcement learning, the reinforcing learning agent selects obfuscation method and iterates through evaluation feedback, the increase in performance overhead caused by the improvement of code protection intensity in the existing technology is solved, and efficient code protection is achieved.
Patent Information
- Application Number
- CN202510347432.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-24
AI Technical Summary
While the existing code protection methods increase protection strength, they lead to increased performance overhead and reduce the practicality of the code.
Using a code virtualization obfuscation method based on reinforcement learning, the reinforcement learning agent selects obfuscation methods in the behavioral space, code obfuscation of key code files, and through NCD algorithm and performance overhead evaluation, feedback rewards guide the agent iterative selection until the balance between obfuscation intensity and performance overhead is met.
It achieves the maintenance of good performance and overhead performance while meeting higher obfuscation intensity, and improves the practicality of the code obfuscation method.
Smart Images

Figure CN120197157A_ABST
Abstract
Description
Technical Field
[0001] The present invention discloses a code virtualization obfuscation method based on reinforcement learning, which relates to the technical field of code protection. Background Art
[0002] With the rapid development of the software field, the code security in software has been increasingly emphasized. For a company, a large amount of manpower, material resources and financial resources are consumed in software R & D. If it is easily cracked, huge losses will be incurred. Therefore, code protection for software is of great importance.
[0003] At present, there are various ways of code protection, such as code obfuscation and virtualization obfuscation. However, most of the current code protection methods only focus on the protection intensity and obfuscation effect. But with the improvement of the protection intensity, the corresponding performance overhead will also increase accordingly, and more additional processing logics need to be added to support the security requirements. Therefore, the practicability of code obfuscation is reduced, which is not conducive to improving the code practicability. Summary of the Invention
[0004] In view of the problems of the prior art, the present invention provides a code virtualization obfuscation method based on reinforcement learning, which balances the obfuscation ability of the obfuscation method and the performance overhead limitation caused, and meets a relatively high obfuscation intensity and good performance overhead performance.
[0005] The specific solution proposed by the present invention is as follows:
[0006] The present invention provides a code virtualization obfuscation method based on reinforcement learning, including:
[0007] Step 1: For the key code files in the business application scenario, use the reinforcement learning agent to select an obfuscation method in the behavior space to obfuscate the key code files.
[0008] Step 2: Evaluate the obfuscation method selected by the reinforcement learning agent from two aspects of the obfuscation effect and the performance overhead performance. Among them, the NCD algorithm is used to evaluate the obfuscation effect, and the performance overhead performance of the obfuscation method is evaluated by directly comparing the running time and volume size of the executable program before and after code obfuscation, and the evaluation result is obtained.
[0009] Step 3: Provide the evaluation result as a feedback reward to the reinforcement learning agent to guide the reinforcement learning agent to select an obfuscation method in the next iteration. The process of loop selecting the obfuscation method and evaluating the obfuscation method is repeated until an obfuscation method that meets the obfuscation intensity and performance overhead performance in the business application scenario is obtained, and the obfuscation method is applied to obfuscate the key code files.
[0010] Further, in step 1 of the method for code virtualization obfuscation based on reinforcement learning, the key code file is preprocessed, converted into an intermediate code file, and the reinforcement learning agent is used to select an obfuscation method in the behavior space to obfuscate the intermediate code file, generating an obfuscated intermediate code file. Then, the obfuscated executable program file is generated through a compiler, and the unobfuscated executable program file of the intermediate code file is generated by the compiler as a benchmark file for evaluating the performance overhead.
[0011] Further, the set included in the behavior space in step 1 of the method for code virtualization obfuscation based on reinforcement learning consists of a virtualization obfuscation method based on LLVM IR and compiler optimization options. The virtualization obfuscation method based on LLVM IR includes a virtualization obfuscation method at the basic block level, an instruction-level virtualization obfuscation method, and combinations between virtualization obfuscation methods.
[0012] Further, in step 2 of the method for code virtualization obfuscation based on reinforcement learning, the NCD algorithm is used to evaluate the obfuscation effect, including: calculating the binary code similarity between the unobfuscated executable program file and the obfuscated executable program file using the NCD algorithm, and judging the difference between the executable programs before and after obfuscation according to the similarity. The smaller the difference, the lower the obfuscation intensity.
[0013] The present invention also provides a code virtualization obfuscation device based on reinforcement learning, including a code processing module, an evaluation module, an iteration module, and an application module.
[0014] The code processing module targets the key code file in the business application scenario, and uses the reinforcement learning agent to select an obfuscation method in the behavior space to obfuscate the key code file.
[0015] The evaluation module evaluates the obfuscation method selected by the reinforcement learning agent from two aspects of the obfuscation effect and the performance overhead performance. Among them, the NCD algorithm is used to evaluate the obfuscation effect, and the performance overhead performance of the obfuscation method is evaluated by directly comparing the running time and volume size of the executable programs before and after code obfuscation, and the evaluation result is obtained.
[0016] The iteration module provides the evaluation result as a feedback reward to the reinforcement learning agent to guide the reinforcement learning agent to select an obfuscation method in the next iteration. The process of repeatedly selecting an obfuscation method and evaluating the obfuscation method is carried out until an obfuscation method that meets the obfuscation intensity and performance overhead performance in the business application scenario is obtained. The application module applies the obfuscation method to obfuscate the key code file.
[0017] Further, the code processing module of the code virtualization obfuscation device based on reinforcement learning preprocesses the key code file, converts the key code file into an intermediate code file, uses the reinforcement learning agent to select an obfuscation method in the behavior space to obfuscate the intermediate code file, generates an obfuscated intermediate code file, then generates an obfuscated executable program file through a compiler, and uses the compiler to generate an unobfuscated executable program file of the intermediate code file as a benchmark file for evaluating the performance overhead.
[0018] Further, when the code processing module of the code virtualization obfuscation device based on reinforcement learning uses the reinforcement learning agent to select an obfuscation method in the behavior space, the behavior space includes a set composed of virtualization obfuscation methods based on LLVM IR and compiler optimization options. The virtualization obfuscation methods based on LLVM IR include basic block-level virtualization obfuscation methods, instruction-level virtualization obfuscation methods, and combinations between virtualization obfuscation methods.
[0019] Further, the evaluation module of the code virtualization obfuscation device based on reinforcement learning uses the NCD algorithm to calculate the binary code similarity between the unobfuscated executable program file and the obfuscated executable program file, and judges the difference between the executable programs before and after obfuscation according to the similarity. The smaller the difference, the lower the obfuscation intensity.
[0020] The beneficial effects of the present invention are as follows:
[0021] With the help of the reinforcement learning agent, the relationship between the obfuscation intensity of the obfuscation method and the performance overhead is transformed into a reinforcement learning task. In multiple iterations of the reinforcement learning agent, an obfuscation method that meets the actual application scenario is selected, so that the obfuscation method has good obfuscation intensity while meeting the performance constraints, making the code obfuscation method more practical. Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0023] Figure 1 It is a schematic diagram of the application framework of the present invention. Detailed Embodiments
[0024] The following will further illustrate the present invention in conjunction with the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited shall not be construed as limiting the present invention.
[0025] Example 1
[0026] The specific solution proposed by the present invention is as follows:
[0027] The present invention provides a code virtualization obfuscation method based on reinforcement learning, including:
[0028] Step 1: For the key code file test.c in the business application scenario, use the reinforcement learning agent to select an obfuscation method in the behavior space to obfuscate the key code file.
[0029] The behavior space includes a set composed of a virtualization obfuscation method based on LLVM IR and compiler optimization options. The virtualization obfuscation method based on LLVM IR includes a virtualization obfuscation method at the basic block level, an instruction-level virtualization obfuscation method, and a combination between virtualization obfuscation methods. This can ensure the obfuscation strength of the finally obtained obfuscation method. In addition, the role of compiler optimization options in performance optimization and changing the program structure is also utilized to optimize the performance overhead of the selected obfuscation method.
[0030] In Step 1, the key code file test.c can also be preprocessed, converted into an intermediate code file test.bc, use the reinforcement learning agent to select an obfuscation method in the behavior space to obfuscate the intermediate code file test.bc, generate an obfuscated intermediate code file out.bc, then generate an obfuscated executable program file out through the compiler, and use the compiler to generate an unobfuscated executable program file test of the intermediate code file as a benchmark file for evaluating the performance overhead performance.
[0031] Step 2: Evaluate the obfuscation method selected by the reinforcement learning agent from two aspects of obfuscation effect and performance overhead performance. Among them, the NCD algorithm is used to evaluate the obfuscation effect, and the performance overhead performance of the obfuscation method is evaluated by directly comparing the running time and volume size of the executable programs before and after code obfuscation, and the evaluation results are obtained.
[0032] When using the NCD algorithm to evaluate the obfuscation effect, it may include: using the NCD algorithm to calculate the binary code similarity between the unobfuscated executable program file test and the obfuscated executable program file out, and judging the difference between the executable programs before and after obfuscation according to the similarity. The smaller the difference, the lower the obfuscation strength;
[0033] When evaluating the performance overhead of the obfuscation method, compare the running time and size of the obfuscated executable program file with those of the benchmark file. For example, if the obfuscated executable program file has a long running time or a large size, the performance overhead is large. At the same time, the business scenario may also limit the performance overhead. For example, in a task scenario that limits time overhead, a task scenario that limits space overhead, and a task scenario that limits both time and space overhead, the expansion coefficient can be parameterized for each actual task scenario to limit the expansion multiples of time and space. Under the limitation of the expansion multiple, find the obfuscation method according to the performance overhead.
[0034] Step 3: Provide the evaluation result as a feedback reward to the reinforcement learning agent to guide the reinforcement learning agent to select an obfuscation method in the next iteration. Repeat the process of selecting and evaluating the obfuscation method until an obfuscation method that meets the obfuscation strength and performance overhead in the business application scenario is obtained, that is, obtain the optimal obfuscation strategy in the iteration after multiple rounds, and apply the obfuscation method to the obfuscation of the key code file.
[0035] Embodiment 2
[0036] The present invention also provides a code virtualization obfuscation device based on reinforcement learning, including a code processing module, an evaluation module, an iteration module, and an application module.
[0037] The code processing module, for the key code file in the business application scenario, uses the reinforcement learning agent to select an obfuscation method in the behavior space to perform code obfuscation on the key code file.
[0038] The evaluation module evaluates the obfuscation method selected by the reinforcement learning agent from two aspects of obfuscation effect and performance overhead. Among them, the NCD algorithm is used to evaluate the obfuscation effect, and the performance overhead of the obfuscation method is evaluated by directly comparing the running time and size of the executable program before and after code obfuscation, and the evaluation result is obtained.
[0039] The iteration module provides the evaluation result as a feedback reward to the reinforcement learning agent to guide the reinforcement learning agent to select an obfuscation method in the next iteration. Repeat the process of selecting and evaluating the obfuscation method until an obfuscation method that meets the obfuscation strength and performance overhead in the business application scenario is obtained. The application module applies the obfuscation method to the obfuscation of the key code file.
[0040] Regarding the information interaction and execution process among the above modules, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention and will not be elaborated here.
[0041] Similarly, the device of the present invention can utilize a reinforcement learning agent to transform the relationship between the obfuscation strength and performance overhead of the obfuscation method into a reinforcement learning task. During multiple iterations of the reinforcement learning agent, an obfuscation method that meets the actual application scenario is selected, enabling the obfuscation method to have good obfuscation strength while satisfying the performance constraints, thereby making the code obfuscation method more practical.
[0042] The above-described embodiments are merely preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are within the protection scope of the present invention. The protection scope of the present invention is subject to the claims.
Claims
1. A code virtualization obfuscation method based on reinforcement learning, characterized by include: Step 1: For key code files in business application scenarios, use reinforcement learning agents to select obfuscation methods in the behavior space to obfuscate key code files. Step 2: Evaluate the obfuscation method selected by the reinforcement learning agent from two aspects: obfuscation effect and performance overhead. The NCD algorithm is used to evaluate the obfuscation effect. The performance overhead of the obfuscation method is evaluated by directly comparing the running time and size of the executable program before and after code obfuscation, and the evaluation results are obtained. Step 3: Provide the evaluation results as feedback rewards to the reinforcement learning agent to guide the reinforcement learning agent to select the obfuscation method in the next iteration. Repeat the process of selecting and evaluating the obfuscation method until an obfuscation method that meets the obfuscation strength and performance overhead requirements in the business application scenario is obtained, and the obfuscation method is applied to obfuscate key code files.
2. According to the code virtualization obfuscation method based on reinforcement learning in claim 1, it is characterized in that in step 1, the key code file is preprocessed, the key code file is converted into an intermediate code file, the intermediate code file is obfuscated by selecting an obfuscation method in the behavior space by the reinforcement learning agent, and the obfuscated intermediate code file is generated, and then the obfuscated executable program file is generated by the compiler, and the unobfuscated executable program file of the intermediate code file is generated by the compiler as a benchmark file for performance overhead performance evaluation.
3. According to the code virtualization obfuscation method based on reinforcement learning according to claim 1, it is characterized by The behavior space in step 1 includes a set of virtualization obfuscation methods based on LLVM IR and compiler optimization options. The virtualization obfuscation methods based on LLVM IR include basic block-level virtualization obfuscation methods, instruction-level virtualization obfuscation methods, and a combination of virtualization obfuscation methods.
4. According to the code virtualization obfuscation method based on reinforcement learning in claim 2, it is characterized by In step 2, the NCD algorithm is used to evaluate the obfuscation effect, including: using the NCD algorithm to calculate the binary code similarity between the unobfuscated executable program file and the obfuscated executable program file, and judging the difference between the executable program before and after obfuscation based on the similarity. The smaller the difference, the lower the obfuscation intensity.
5. A code virtualization obfuscation device based on reinforcement learning, characterized in that It includes code processing module, evaluation module, iteration module and application module. The code processing module uses reinforcement learning agents to select obfuscation methods in the behavior space to obfuscate key code files in business application scenarios. The evaluation module evaluates the obfuscation method selected by the reinforcement learning agent from two aspects: obfuscation effect and performance overhead. The NCD algorithm is used to evaluate the obfuscation effect. The performance overhead of the obfuscation method is evaluated by directly comparing the running time and size of the executable program before and after code obfuscation, and the evaluation results are obtained. The iteration module provides the evaluation results as feedback rewards to the reinforcement learning agent, guiding the reinforcement learning agent to select the obfuscation method in the next iteration. The process of selecting and evaluating the obfuscation method is repeated until an obfuscation method that meets the obfuscation strength and performance overhead requirements in the business application scenario is obtained. The application module applies the obfuscation method to obfuscate key code files.
6. The code virtualization obfuscation device based on reinforcement learning according to claim 5 is characterized in that the code The processing module preprocesses key code files, converts the key code files into intermediate code files, uses the reinforcement learning agent to select the obfuscation method in the behavior space to confuse the intermediate code files, generates obfuscated intermediate code files, and then generates obfuscated executable program files through the compiler. The compiler is used to generate unobfuscated executable program files of the intermediate code files as benchmark files for performance overhead evaluation.
7. According to the code virtualization obfuscation method based on reinforcement learning in claim 5, it is characterized in that the code When the processing module uses the reinforcement learning agent to select the obfuscation method in the behavior space, the behavior space includes a set of virtualization obfuscation methods based on LLVM IR and compiler optimization options. The virtualization obfuscation method based on LLVM IR includes a basic block-level virtualization obfuscation method, an instruction-level virtualization obfuscation method, and a combination of virtualization obfuscation methods.
8. The code virtualization obfuscation method based on reinforcement learning according to claim 6 is characterized by: The evaluation module uses the NCD algorithm to calculate the binary code similarity between the unobfuscated executable program file and the obfuscated executable program file, and determines the difference between the executable program before and after obfuscation based on the similarity. The smaller the difference, the lower the obfuscation intensity.
Citation Information
Cited By
Universal LLVM optimization method and system based on reinforcement learning
CN121433627A