Compilation method, system and equipment based on confusion mechanism and medium

By analyzing the source code in function structure and reorganizing the control flow, and using the basic block flattening method to generate obfuscated intermediate code, the problem of insufficient defense capabilities caused by the transparency of the control flow in the existing compilation method is solved, and higher anti-attack ability and system security are achieved.

CN120046145APending Publication Date: 2025-05-27CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +3
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411901691.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing compilation methods for master chips with different instruction set architectures make the control flow more transparent to attackers, lack of defense capabilities, and difficult to resist malicious attacks.

Method used

A compilation method based on obfuscation mechanism is proposed. By analyzing the source code function structure, generating function feature information, performing language analysis to generate intermediate code, and using the basic block flattening method to reorganize the control flow, generate target files and convert them into machine code.

Benefits of technology

By performing functional structure analysis and control flow reorganization of the source code before compilation, the complexity and difficulty of prediction of the code are increased, the source code's resistance to attacks is improved, and the system's security and reliability are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046145A_ABST
    Figure CN120046145A_ABST
Patent Text Reader

Abstract

The invention provides a compilation method, system and device based on an obfuscation mechanism and a medium, and the method comprises the steps: carrying out the function structure analysis of an obtained to-be-compiled source code, and obtaining the function feature information of the source code; performing language analysis on the source code according to the function feature information to obtain an intermediate code corresponding to the source code; based on an intermediate code corresponding to the source code, performing control flow recombination on the intermediate code by using a basic block flattening method to generate a target file, and converting the target file into a machine code; according to the method, the function structure analysis is performed on the source code before compiling, so that the control flow structure of the source code can be understood in a finer-grained manner; and a basic block flattening method is introduced to reorganize basic program blocks in the control flow diagram, so that the control flow is more complex and difficult to predict, and the anti-attack capability of source codes is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electric power information security, and particularly relates to a compilation method, system, device and medium based on an obfuscation mechanism. Background Art

[0002] In recent years, the architectures and applications of master control chips have shown a diversified trend. According to their application targets, they can be roughly divided into two categories: basic MCUs (Microcontroller Units) for single applications and complex SOCs (System on Chips) with operating systems and containers or for multi-target applications. Each category can be further subdivided according to different core numbers and different instruction set architectures.

[0003] Generally, master control chips with different instruction set architectures have their own supporting compilation toolchains. Users can only use the language codes supported by the compilation tools and do not have the conditions for diversified compilation, so they cannot achieve mimetic upgrading at the software level. For multi-core SOCs with built-in FPGAs (Field Programmable Gate Arrays) or master control chips with multiple instruction set architectures that will appear in the future, existing compilation tools cannot build mimetic scheduling decision logics between multiple heterogeneous cores;

[0004] Existing compilation methods for master control chips with different instruction set architectures usually consist of three parts: a front end, an optimizer, and a back end. During compilation, the front end is mainly responsible for lexical and syntactic analysis and converts the source code into an abstract syntax tree; the optimizer optimizes the obtained intermediate code on the basis of the front end to make the code more efficient; the back end converts the optimized intermediate code into machine code for their respective instruction set architectures and platforms. LLVM (Low Level Virtual Machine) is a compilation toolchain adopting this three-stage architecture. Front-end developers only need to convert the source code into intermediate code that the optimizer can understand without knowing the working principle of the optimizer and the knowledge of the target machine, which greatly reduces the development difficulty of the compiler. However, in this traditional process of the front end being responsible for parsing, validating, and diagnosing errors in the input code, constructing an abstract syntax tree, then translating the abstract syntax tree into intermediate code, then improving the code through a series of analyses and optimizations, and then inputting the code generator to generate target machine code, the source code and resources are not converted, and the control flow is relatively transparent to attackers, resulting in vulnerabilities in the source code being easily cracked by attackers through static or dynamic analysis means, thus the defense ability is insufficient, reducing the security and reliability of the system. Summary of the Invention

[0005] To solve the problem that in the compilation process of a main control chip composed of different instruction set architectures, the compiler directly translates the source code into target machine code, making the control flow relatively transparent to attackers and resulting in insufficient defense capabilities against malicious attacks, the present invention proposes a compilation method based on an obfuscation mechanism, including:

[0006] Perform function structure analysis on the obtained source code to be compiled to obtain function feature information of the source code;

[0007] According to the function feature information, perform language analysis on the source code to obtain intermediate code corresponding to the source code;

[0008] Based on the intermediate code corresponding to the source code, use the basic block flattening method to perform control flow restructuring on the intermediate code, generate a target file, and convert the target file into machine code.

[0009] Optionally, the performing function structure analysis on the obtained source code to be compiled to obtain function feature information of the source code includes:

[0010] Perform function classification on the obtained source code to be compiled to obtain function classification information to which the source code belongs;

[0011] According to the function classification information, obtain the call relationship of functions in the source code, and generate function feature information of the source code according to the call relationship.

[0012] Optionally, the performing language analysis on the source code according to the function feature information to obtain intermediate code corresponding to the source code includes:

[0013] According to the function feature information, perform lexical analysis on the source code to obtain a lexical sequence corresponding to the source code;

[0014] According to the lexical sequence, use context-free grammar to perform syntax analysis on the source code and output a syntax tree corresponding to the source code;

[0015] Perform semantic analysis on the syntax tree corresponding to the source code to obtain a syntax tree with annotation information, and convert the syntax tree with annotation information into intermediate code.

[0016] Optionally, the performing control flow restructuring on the intermediate code based on the intermediate code corresponding to the source code by using the basic block flattening method to generate a target file includes:

[0017] According to the intermediate code corresponding to the source code, use the dynamic opaque predicate insertion method to perform opaque obfuscation on the intermediate code to obtain obfuscated intermediate code;

[0018] Use the basic block flattening method to perform control flow restructuring on the obfuscated intermediate code to obtain the flattened obfuscated intermediate code;

[0019] Through a pre-set multi-compiler, perform code optimization on the flattened obfuscated intermediate code to generate the target file corresponding to the source code.

[0020] Optionally, the step of performing code optimization on the flattened obfuscated intermediate code through a pre-set multi-compiler to generate the target file corresponding to the source code includes:

[0021] Through a pre-set multi-compiler, perform compilation conversion on the flattened obfuscated intermediate code to obtain each basic program block corresponding to the flattened obfuscated intermediate code;

[0022] Through a code optimizer, analyze the execution time of each basic program block to generate an executable file corresponding to each basic program block;

[0023] Generate a target file according to the executable files corresponding to each basic program block.

[0024] Optionally, the step of generating a target file according to the executable files corresponding to each basic program block includes:

[0025] Execute the executable files corresponding to each basic program block through a performance analyzer and record the running time of each basic program block;

[0026] According to the running time of each basic program block, select the optimal code block corresponding to each basic program block;

[0027] Synthesize the optimal code blocks corresponding to each basic program block into a target file according to the running time.

[0028] Optionally, the function feature information includes one or more of the following: function call relationship, function dependency, return value, function complexity, code reuse, function scale, exception handling, and boundary conditions.

[0029] Based on the same inventive concept, the present invention also provides a compilation system based on an obfuscation mechanism, including:

[0030] A structure analysis module, configured to perform function structure analysis on the obtained source code to be compiled to obtain the function feature information of the source code;

[0031] A language analysis module, configured to perform language analysis on the source code according to the function feature information to obtain the intermediate code corresponding to the source code;

[0032] A control flow restructuring module, which is used to perform control flow restructuring on the intermediate code corresponding to the source code by using the basic block flattening method, generate a target file, and convert the target file into machine code.

[0033] Optionally, the structure analysis module includes:

[0034] A function classification sub-module, which is used to classify the source code to be compiled obtained, and obtain the function classification information to which the source code belongs;

[0035] A function call sub-module, which is used to obtain the call relationship of the functions in the source code according to the function classification information, and generate the function feature information of the source code according to the call relationship.

[0036] Optionally, the language analysis module includes:

[0037] A lexical analysis sub-module, which is used to perform lexical analysis on the source code according to the function feature information to obtain the lexical sequence corresponding to the source code;

[0038] A syntax analysis sub-module, which is used to perform syntax analysis on the source code by using context-free grammar according to the lexical sequence, and output the syntax tree corresponding to the source code;

[0039] A semantic analysis sub-module, which is used to perform semantic analysis on the syntax tree corresponding to the source code, obtain a syntax tree with annotation information, and convert the syntax tree with annotation information into intermediate code.

[0040] Optionally, the control flow restructuring module includes:

[0041] A code obfuscation sub-module, which is used to perform opaque obfuscation on the intermediate code corresponding to the source code by using the dynamic opaque predicate insertion method to obtain obfuscated intermediate code;

[0042] A code flattening sub-module, which is used to perform control flow restructuring on the obfuscated intermediate code by using the basic block flattening method to obtain the obfuscated intermediate code after flattening;

[0043] A code optimization sub-module, which is used to perform code optimization on the obfuscated intermediate code after flattening through a preset multi-compiler to generate the target file corresponding to the source code.

[0044] Optionally, the code optimization sub-module includes:

[0045] A compilation conversion unit, which is used to perform compilation conversion on the obfuscated intermediate code after flattening through a preset multi-compiler to obtain each basic program block corresponding to the obfuscated intermediate code after flattening;

[0046] A time analysis unit, configured to analyze the execution time of each of the basic program blocks through a code optimizer, and generate an executable file corresponding to each of the basic program blocks;

[0047] A target file generation unit, configured to generate a target file according to the executable files corresponding to each of the basic program blocks.

[0048] Optionally, the target file generation unit includes:

[0049] A file execution subunit, configured to execute the executable files corresponding to each of the basic program blocks through a performance analyzer, and record the running time of each of the basic program blocks;

[0050] A code block selection subunit, configured to select an optimal code block corresponding to each of the basic program blocks according to the running time of each of the basic program blocks;

[0051] A file synthesis subunit, configured to synthesize the optimal code blocks corresponding to each basic program block into a target file according to the running time.

[0052] Optionally, the function feature information includes one or more of the following: function call relationship, function dependency, return value, function complexity, code reuse, function scale, exception handling, and boundary conditions.

[0053] On the other hand, the present invention further provides an electronic device, including: at least one processor and a memory; the memory and the processor are connected by a bus;

[0054] The memory is configured to store one or more programs;

[0055] When the one or more programs are executed by the at least one processor, the above-mentioned compilation method based on an obfuscation mechanism is implemented.

[0056] On the other hand, the present invention further provides a computer-readable storage medium, on which an execution program is stored, and when the execution program is executed, the above-mentioned compilation method based on an obfuscation mechanism is implemented.

[0057] Compared with the prior art, the beneficial effects of the present invention are:

[0058] The present invention provides a compilation method, system, device, and medium based on an obfuscation mechanism, including: performing function structure analysis on the obtained source code to be compiled to obtain function feature information of the source code; performing language analysis on the source code according to the function feature information to obtain intermediate code corresponding to the source code; based on the intermediate code corresponding to the source code, using the basic block flattening method to perform control flow reorganization on the intermediate code to generate a target file, and converting the target file into machine code; through function structure analysis of the source code before compilation, the present application can understand the control flow structure of the source code in a finer granularity; then, by introducing the basic block flattening method, the basic program blocks in the control flow graph are reorganized to make the control flow more complex and unpredictable, which is beneficial to enhancing the anti-attack ability of the source code. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a schematic flowchart of a compilation method based on an obfuscation mechanism provided by the present invention;

[0060] Figure 2 It is a schematic execution logic diagram of the source code without using the basic block flattening method in the compilation method based on an obfuscation mechanism provided by the present invention;

[0061] Figure 3 It is a schematic execution logic diagram of the source code using the basic block flattening method in the compilation method based on an obfuscation mechanism provided by the present invention;

[0062] Figure 4 It is a schematic execution framework diagram of multi-compiler fusion compilation in the compilation method based on an obfuscation mechanism provided by the present invention;

[0063] Figure 5 It is a schematic framework diagram of multi-module implementation of compilation in the compilation method based on an obfuscation mechanism provided by the present invention;

[0064] Figure 6 It is a functional implementation diagram of each module when multi-module implementation of compilation is adopted in the compilation method based on an obfuscation mechanism provided by the present invention;

[0065] Figure 7 It is a schematic structural composition diagram of a compilation system based on an obfuscation mechanism provided by the present invention;

[0066] Figure 8 It is a schematic structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] The present invention proposes a compilation method, system, device, and medium based on an obfuscation mechanism. The following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings.

[0068] Example 1:

[0069] A compilation method based on an obfuscation mechanism provided by the present invention has a process schematic diagram as Figure 1 shown, including:

[0070] Step 1: Perform function structure analysis on the obtained source code to be compiled to obtain function feature information of the source code;

[0071] Step 2: According to the function feature information, perform language analysis on the source code to obtain intermediate code corresponding to the source code;

[0072] Step 3: Based on the intermediate code corresponding to the source code, use the basic block flattening method to perform control flow reorganization on the intermediate code, generate a target file, and convert the target file into machine code.

[0073] In modern software development, with the frequent emergence of software vulnerabilities, network security issues have become increasingly severe. Against these vulnerabilities, the means of attack and exploitation are constantly evolving, and traditional security protection measures can no longer completely eliminate various threats. To improve the security and reliability of the system, software diversification has emerged as an effective defense strategy. Software diversification uses different compilation and code restructuring technologies to generate multiple heterogeneous variants of the same application program, thereby increasing the difficulty for attackers to reverse engineer and exploit the software. However, software diversification has not completely eliminated security threats. It can only increase the difficulty of attacking and exploiting software vulnerabilities to a certain extent. The software protection mechanism based on the mimic defense concept does not rely on the confidentiality of a single version, but constructs a set containing multiple heterogeneous variants through diversification methods. These variants can produce different behavioral outputs during software execution, thereby effectively resisting attacks against static or known versions. The commonly used diversification means is to generate a variant set through diversification compilation technology, including processes such as structural analysis, language analysis, and reorganization of intermediate code of the source code. Specifically:

[0074] In one implementation manner, the process of performing function structure analysis on the obtained source code to be compiled in the above Step 1 to obtain function feature information of the source code may include:

[0075] Perform function classification on the obtained source code to be compiled to obtain function classification information to which the source code belongs;

[0076] According to the function classification information, obtain the call relationship of functions in the source code, and generate function feature information of the source code according to the call relationship;

[0077] Exemplarily, the above function feature information may include one or more of the following: function call relationship, function dependency, return value, function complexity, code reuse, function scale, exception handling, and boundary conditions;

[0078] In this implementation, by performing function structure analysis on the source code to be compiled, the function feature information of the source code can be effectively extracted and summarized. This method not only improves the understandability and maintainability of the program source code but also provides a solid foundation for subsequent code optimization and refactoring. By classifying and analyzing functions, a clear function call relationship graph can be formed, helping developers identify the dependencies and complexity between functions, thereby more efficiently managing and optimizing the code structure. In addition, this implementation has advantages in code reuse. By analyzing feature information such as function size and return value, developers can identify and utilize existing functions, reducing repetitive work and improving development efficiency. At the same time, the analysis of exception handling and boundary conditions also ensures the robustness of the program, enabling potential problems emerging during the development process to be discovered and handled in a timely manner. Therefore, through this implementation, not only can it provide a tool for static analysis of the code, but also it can provide the necessary information support for developers during dynamic analysis and performance optimization, thereby significantly improving the overall efficiency and quality of software development.

[0079] After completing the above steps, the rich function feature information obtained will help to deeply understand the structure and behavioral characteristics of the source code, which lays a solid foundation for subsequent language analysis, enabling the source code to be more systematically parsed and prepared for the generation of intermediate code. Therefore, the next step is to use this function feature information to carry out language analysis on the source code to obtain the corresponding intermediate code. Specifically:

[0080] In one implementation, the process of performing language analysis on the source code according to the function feature information in step 2 above to obtain the intermediate code corresponding to the source code may include:

[0081] Perform lexical analysis on the source code according to the function feature information to obtain the lexical sequence corresponding to the source code;

[0082] According to the lexical sequence, perform syntax analysis on the source code using context-free grammar and output the syntax tree corresponding to the source code;

[0083] Perform semantic analysis on the syntax tree corresponding to the source code to obtain a syntax tree with annotation information, and convert the syntax tree with annotation information into intermediate code;

[0084] In this implementation, through systematic language analysis of the source code, an effective conversion from the source code to the intermediate code is achieved, demonstrating the importance and practicality of compilation principles. This process first performs lexical analysis, treating the source code as a multi-line string, scanning each character one by one to identify "word" symbols. This stage is the first step of compilation, and its main task is to convert the character sequence into words in the form of pairs according to the lexical rules of the language, including word categories and word values, forming a complete lexical sequence, laying the foundation for subsequent analysis. After completing lexical analysis, the next syntactic analysis stage uses context-free grammar to parse the lexical sequence and constructs a syntax tree corresponding to the source code. The main task of syntactic analysis is to decompose the word symbol sequence into various syntactic units according to the syntactic rules of the language, such as "expressions" and "statements". This stage not only ensures that the structure of the source program conforms to the syntactic specifications but also can provide diagnostic information in a timely manner when syntax errors are found, enhancing the code correctness and reliability during the development process. Only when both syntax and semantics are correct can the source program be effectively converted into intermediate code to ensure the integrity of the language structure and the logical consistency of the program. When further performing semantic analysis on the syntax tree, the system checks for semantic errors in the source code and collects type information for use in subsequent code generation. The key to this stage lies in type analysis and checking to ensure the correct semantics of the source program. Finally, through the transformation of the syntax tree, the obtained intermediate code, as an execution language between the source code and the machine code, not only has high scalability but also can effectively handle potential problems brought about by diverse transformations of the source code. This process forms the core content of software diverse compilation, achieving flexible processing and efficient management of the source code by utilizing the intermediate code.

[0085] After successfully generating the intermediate code, the next step is to further optimize and transform these intermediate codes to ensure that the final generated target file has higher performance and security. To achieve this goal, a series of control flow restructuring and obfuscation techniques will be adopted based on the intermediate code corresponding to the source code, thereby enhancing the anti-reverse engineering ability of the code. Specifically:

[0086] In one implementation, in step 3 above, based on the intermediate code corresponding to the source code, the process of using the basic block flattening method to perform control flow restructuring on the intermediate code to generate the target file may include:

[0087] According to the intermediate code corresponding to the source code, use the dynamic opaque predicate insertion method to perform opaque obfuscation on the intermediate code to obtain the obfuscated intermediate code;

[0088] Use the basic block flattening method to perform control flow restructuring on the obfuscated intermediate code to obtain the flattened obfuscated intermediate code;

[0089] Through a pre-set multi-compiler, code optimization is performed on the flattened obfuscated intermediate code to generate the target file corresponding to the source code;

[0090] In this implementation method, the opaque predicate in the dynamic opaque predicate insertion method refers to that the calculation result of the predicate v at a certain point P in the program is predictable before obfuscation, that is, known to the program obfuscator or developer, and its result is uncertain for the reverse analyst after obfuscation. When constructing an opaque predicate, if the output of the predicate is always true, it is denoted as P T ; if the output of the predicate is always false, it is denoted as P F ; if the output of the predicate is sometimes true and sometimes false, it is denoted as P ? 。 Through a condition that is always true or always false, the opaque predicate can be used as a condition generator for control flow obfuscation. The opaque predicate technology constructs a "black box" effect by increasing the complexity of the branch structure, enabling the reverse analyst to only analyze the predicate logic by observing the input and output values. Furthermore, by implementing the dynamics of the opaque predicate, the attacker cannot determine the specific flow of the program; Reorganizing the control flow of the obfuscated code through the basic block flattening method means transforming the clones of some basic program blocks in the function and scrambling the order of the basic program blocks into a switch statement, making the logical relationship of the basic program blocks change from linear to parallel, constructing opaque predicates according to certain rules, and finally determining the target basic block of the next jump through the opaque predicate. As shown in Figures 2 - 3 Figure 11 is a comparison diagram before and after processing using the basic block flattening method. Specifically, Figure 2 Figure 13 is the code execution logic flow chart of the source code without using the basic block flattening method. The specific execution steps are as follows:

[0091] Initialization: The variable a is assigned the value 1, and the variable b is assigned the value 2.

[0092] Enter L1: Check the condition a < 10. If the condition is not met (i.e., a >= 10), jump to L4.

[0093] Calculation: If a < 10, then execute b = a + b, adding the values of a and b and assigning the result to b.

[0094] Check b: Subsequently, check the condition b > 10. If this condition is not met, jump to L2; if it is met, execute b =

[0095] b - 1.

[0096] Execute L2: In L2, execute a++, incrementing a by 1 and returning to L1 to continue the loop check.

[0097] Repeat the loop: This process will continue to loop until a reaches 10 or greater, causing the initial L1 condition to be not met.

[0098] End: Finally, when a is greater than or equal to 10, jump to L4 and execute the final operation use(b), using the variable b. Through the above steps, the basic program block continuously adjusts the values of a and b through conditional judgments and loops until the exit condition is reached. Figure 3 The following is the execution logic flowchart for processing the source code using the basic block flattening method. The specific execution steps are as follows:

[0099] Initialization: Set the initial value of the variable swVar to 1. In addition, a is assigned the value 1 and b is assigned the value 2.

[0100] Enter the switch statement: According to the value of swVar, execute the corresponding code block. The initial value is 1, so L1 will be executed. L1: Set a = 1, b = 2, and update swVar to 2.

[0101] L2 (according to the new value of swVar): Check the condition!(a < 10). If the condition holds, update swVar to 6; if not, update it to 3.

[0102] Jump according to the value of swVar:

[0103] If swVar is 6, jump to L3, execute b = b + a and check if b is greater than 10, and update swVar.

[0104] If swVar is 3, jump to L4, execute b-- and update swVar.

[0105] Continue the loop: After updating swVar, return to the initial switch judgment and enter the new code block until the exit condition is met.

[0106] Final operation: When a certain loop is completed, it may reach L6 and "use" the variable b.

[0107] By comparing the execution logics with and without the basic block flattening method, it can be seen that the basic block flattening method provides significant advantages in terms of code optimization and the clarity of the execution logic. First, by reducing the branches and jumps in the control flow, it reduces the complexity of the program, making the execution path more intuitive. Second, the readability and maintainability of the flattened code are significantly enhanced, facilitating developers to quickly understand and modify the code. In addition, the optimized execution process reduces the runtime jumps, improving the overall performance of the program. Finally, this structured code form simplifies subsequent analysis and optimization, helping the compiler to optimize the code more effectively. In summary, the basic block flattening method significantly improves the code quality and execution efficiency; in addition, due to the differences in the compiler developers' understanding and perception of compilation strategies, as well as the different design concepts of the compilers themselves, different compilers on the same architecture will have significant differences in the representation of the same source code segment, which makes it possible to generate multiple heterogeneous terminal application execution bodies. Therefore, in this implementation, the code is fused and compiled by multiple pre-set compilers to draw on the advantages of different compilers and obtain more diverse application execution bodies. During the fusion compilation process, different compilers such as GCC and LLVM can be integrated (GCC stands for GNU Compiler Collection, representing an open-source compiler collection; LLVM stands for Low Level Virtual Machine, representing a low-level virtual machine). Part of the source program is compiled using different compilers, and the various codes are linked to generate a complete application binary file. For different front-ends and back-ends, the optimization phase of LLVM uses a unified intermediate code. If a new programming language or a new instruction set chip needs to be supported, only a new front-end or back-end needs to be implemented, without the need to modify the optimization phase. It can be seen that by performing heterogeneous processing on the software at this stage, multi-target compilation can be achieved for any compilable language and any target machine code. In addition, the abstract syntax tree and the overall design provided by LLVM are human-readable, which is useful for the analysis of heterogeneous code; moreover, LLVM has better modularity and reusability and can be reused by source code analysis tools, refactoring, integrated development environments (Integrated Development Environment, IDE), etc.; at the same time, the heterogeneity of LLVM can be carried out throughout the process, including compile time, link time, load time, and run time; in summary, making full use of these characteristics of LLVM in the design of the compiler can make it more than capable in the implementation of heterogeneous processing and the construction of mimetic data handling resources (Data Handling Resource, DHR) at each stage.

[0108] During the process of generating the target file, special emphasis will be placed on more refined processing through a pre-set multi-compiler. This process aims to ensure that the final target file achieves the best results in terms of execution efficiency and security. The following steps will describe in detail how to use the multi-compiler and the code optimizer to generate a more efficient executable file and the final target file from the flattened obfuscated intermediate code. Specifically:

[0109] In one implementation, the process of optimizing the code of the flattened obfuscated intermediate code through the pre-set multi-compiler to generate the target file corresponding to the source code may include:

[0110] Compile and transform the flattened obfuscated intermediate code through the pre-set multi-compiler to obtain each basic program block corresponding to the flattened obfuscated intermediate code;

[0111] Analyze the execution time of each basic program block through the code optimizer to generate the executable file corresponding to each basic program block;

[0112] Generate the target file according to the executable file corresponding to each basic program block;

[0113] In this implementation, the optimization steps of integrating the multi-compiler are divided into three stages:

[0114] The first stage is to perform various compilation conversions on the source code to form basic program blocks. If there are jumps, the code at the jump location is inlined, and the code in the loop nest is replaced with function calls, etc. For the processed program blocks, data flow analysis and control flow analysis are performed to ensure the logical equivalence of the same program block under different compilers and achieve the alignment of the program blocks.

[0115] The second stage is the analysis stage. In this stage, the analyzer will execute the executable files generated by the code optimizer one by one and perform multiple runs for stable data. The Profiler (i.e., the performance analyzer) collects the analysis information of each program block at the end of each execution and analyzes the execution time of the extracted code blocks. Execute the executable files generated for each code optimizer and collect the execution time.

[0116] The third stage is the synthesis stage. Here, for each code block, the loop execution times collected from each code optimizer are compared, and the code optimizer that produces the best execution code is selected, that is, the optimized code that completes the execution of each program block in the shortest time. Finally, the default compiler is used to link the target files in the selected code optimizers of each code block, as well as the target files generated by the default compiler for the base files. This stage also needs to link the libraries that the code optimizer may have used or supported to generate the code of the complete file;

[0117] As Figure 4 shown in the schematic diagram of the execution framework for multi-compiler fusion compilation, the figure shows the execution framework of multi-compiler fusion during the code optimization process. First, the source program is divided into multiple program blocks (such as program blocks A, B, C, D, and E). Then, these blocks are compiled and optimized in parallel by different compilers (such as GCC, Clang, and Pluto). Each compiler may adopt different optimization strategies for the same program block to generate updated program blocks (such as program blocks A', B', C', etc.). The program blocks after compilation and optimization are finally aggregated into an executable program, and these blocks are integrated into multiple executable program parts (such as executable programs 1, 2, and 3) to form a complete execution path. In this way, multi-compiler fusion optimization can significantly improve the execution efficiency of the code, utilize the advantages of different compilers, and achieve more efficient program performance.

[0118] On the premise of ensuring that the generated target file is not only correct but also can achieve the best performance, a series of performance analyses will be further carried out to select the optimal execution path and code blocks. This process is a key step in achieving the efficient final target file and will directly affect its running speed and resource occupancy. Next, how to further optimize each basic program block through performance analysis and finally synthesize the target file will be described in detail. Specifically:

[0119] In one implementation, the process of generating the target file based on the executable files corresponding to each basic program block may include:

[0120] Execute the executable files corresponding to each basic program block through a performance analyzer and record the running time of each basic program block;

[0121] Select the optimal code block corresponding to each basic program block according to the running time of each basic program block;

[0122] Synthesize the target file by arranging the optimal code blocks corresponding to each basic program block according to the running time;

[0123] In this implementation, by integrating the processes of performance analysis and code optimization, the execution efficiency and resource utilization rate of the program can be significantly improved. First, by using a performance analyzer to separately evaluate each basic program block, the running time of each code segment can be accurately identified. This meticulous evaluation enables developers to understand the performance of each component in the program in a data-driven manner, thus laying the foundation for subsequent optimization decisions. In the stage of selecting the optimal code blocks, by combining real-time data on running time, the program can dynamically adapt to different operating environments and load conditions. This flexibility not only improves the execution efficiency of the code blocks but also leaves more possibilities for future expansion and maintenance. Through this performance-based code selection mechanism, the overall running efficiency of the program will be substantially enhanced, enabling it to better handle big data processing and rapid response requirements. Especially in application scenarios that require high performance, it can provide a better user experience. Finally, the optimized code blocks are synthesized into a target file according to the running time, enabling the entire program to run in an optimal state under different input conditions, minimizing resource waste and performance bottlenecks. This strategy of integrating performance analysis and optimal code selection helps to form a closed-loop of continuous optimization, enabling the software to continuously adapt to new demand changes and technological challenges throughout its life cycle, ultimately driving the entire software industry towards higher performance standards.

[0124] In summary, in view of the problem that during the compilation process of a main control chip composed of different instruction set architectures, the compiler directly translates the source code into target machine code, making the control flow relatively transparent to attackers and resulting in insufficient defense capabilities against malicious attacks, the present invention proposes a compilation method based on a confusion mechanism. The execution process of this method can be reflected by the functions of four modules, as Figure 5 shown, namely: a source code diversification transformation module, a compilation front-end module, an intermediate code diversification transformation module, and a compilation back-end processing module. Specifically, the functions of each module are as Figure 6 shown. The source code diversification transformation module executes the source code diversification strategy to achieve diversification transformation of the program at the source code level, analyzes the function structure, and provides heterogeneous execution bodies for mimic defense design; the compilation front-end module implements functions such as lexical analysis, syntax analysis, and semantic analysis; the intermediate code diversification transformation module mainly reorganizes the control flow by adopting dynamic opaque predicates and basic block flattening methods; the compilation back-end module is mainly used to implement code optimization and target code generation. Through the execution processes of the above modules, a mimic DHR structure can be realized to meet diversified application requirements. And by studying various application scenarios in the power industry, in view of the characteristics of main control chips with different instruction set architectures and operating systems, a compilation method suitable for these diversified environments is summarized. Through this highly adaptable compilation method, the performance can be effectively optimized for different instruction set architectures and operating systems, thus better serving the specific needs of the power industry.

[0125] Example 2:

[0126] Based on the same inventive concept, the present invention also provides a compilation system based on a confusion mechanism. The schematic diagram of the structural composition is as shown in Figure 7 shown, including:

[0127] A structure analysis module for performing function structure analysis on the obtained source code to be compiled to obtain function feature information of the source code;

[0128] A language analysis module for performing language analysis on the source code according to the function feature information to obtain intermediate code corresponding to the source code;

[0129] A control flow reorganization module for performing control flow reorganization on the intermediate code based on the intermediate code corresponding to the source code by using the basic block flattening method to generate a target file and converting the target file into machine code.

[0130] In one implementation, the above-mentioned structure analysis module may include:

[0131] A function classification sub-module for classifying the obtained source code to be compiled to obtain function classification information to which the source code belongs;

[0132] A function call sub-module for obtaining the call relationship of functions in the source code according to the function classification information and generating function feature information of the source code according to the call relationship.

[0133] Exemplarily, the above-mentioned function feature information may include one or more of the following: function call relationship, function dependency, return value, function complexity, code reuse, function scale, exception handling, and boundary conditions.

[0134] In one implementation, the above-mentioned language analysis module may include:

[0135] A lexical analysis sub-module for performing lexical analysis on the source code according to the function feature information to obtain a lexical sequence corresponding to the source code;

[0136] A syntax analysis sub-module for performing syntax analysis on the source code by using context-free grammar according to the lexical sequence and outputting a syntax tree corresponding to the source code;

[0137] A semantic analysis sub-module for performing semantic analysis on the syntax tree corresponding to the source code to obtain a syntax tree with annotation information and converting the syntax tree with annotation information into intermediate code.

[0138] In one implementation, the above-mentioned control flow reorganization module may include:

[0139] A code obfuscation sub-module, which is used to perform opaque obfuscation on the intermediate code according to the intermediate code corresponding to the source code by using the dynamic opaque predicate insertion method to obtain the obfuscated intermediate code;

[0140] A code flattening sub-module, which is used to perform control flow restructuring on the obfuscated intermediate code by using the basic block flattening method to obtain the flattened obfuscated intermediate code;

[0141] A code optimization sub-module, which is used to perform code optimization on the flattened obfuscated intermediate code through a pre-set multi-compiler to generate the target file corresponding to the source code.

[0142] In one implementation, the above code optimization sub-module may include:

[0143] A compilation conversion unit, which is used to perform compilation conversion on the flattened obfuscated intermediate code through a pre-set multi-compiler to obtain each basic program block corresponding to the flattened obfuscated intermediate code;

[0144] A time analysis unit, which is used to analyze the execution time of each basic program block through a code optimizer to generate an executable file corresponding to each basic program block;

[0145] A target file generation unit, which is used to generate a target file according to the executable file corresponding to each basic program block.

[0146] In one implementation, the above target file generation unit may include:

[0147] A file execution sub-unit, which is used to execute the executable file corresponding to each basic program block through a performance analyzer and record the running time of each basic program block;

[0148] A code block selection sub-unit, which is used to select the optimal code block corresponding to each basic program block according to the running time of each basic program block;

[0149] A file synthesis sub-unit, which is used to synthesize the optimal code block corresponding to each basic program block into a target file according to the running time.

[0150] Embodiment 3:

[0151] As Figure 8 shown, the present invention also provides an electronic device, which may be a computer device, a single-chip microcomputer device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, the processor, and the transceiver component are connected through a bus; the memory can be used to store an execution program, and an exemplary execution program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, and the data can be called and / or modified when the instructions are executed.

[0152] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a compilation method based on an obfuscation mechanism in the above embodiments.

[0153] Embodiment 4:

[0154] Based on the same inventive concept, the present invention also provides a readable storage medium, specifically an electronic device-readable storage medium (Memory). The electronic device-readable storage medium is a memory device in the electronic device, used to store programs and data. It can be understood that the storage medium here can include both the built-in storage medium in the electronic device, and of course, can also include the extended storage medium supported by the electronic device. The storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more execution programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory, or a non-volatile memory, such as at least one disk memory. By loading and executing one or more instructions stored in the storage medium by the processor, the steps of a compilation method based on an obfuscation mechanism in the above embodiments can be implemented.

[0155] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0156] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more flows and / or one or more blocks in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or one or more blocks Figure 1 means for implementing the functions specified in one or more flows and / or one or more blocks.

[0157] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in one or more flows and / or one or more blocks. Figure 1 one or more flows and / or one or more blocks Figure 1 means for implementing the functions specified in one or more flows and / or one or more blocks.

[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or one or more blocks. Figure 1 one or more flows and / or one or more blocks Figure 1 steps for implementing the functions specified in one or more flows and / or one or more blocks.

[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit the scope of its protection. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that after reading the present invention, various changes, modifications, or equivalent substitutions can still be made to the specific implementation manners of the application. However, these changes, modifications, or equivalent substitutions are all within the scope of the protection of the claims pending for the application.

Claims

1. A compilation method based on an obfuscation mechanism, characterized in that: include: Performing function structure analysis on the acquired source code to be compiled to obtain function feature information of the source code; Performing language analysis on the source code according to the function feature information to obtain an intermediate code corresponding to the source code; Based on the intermediate code corresponding to the source code, the control flow of the intermediate code is reorganized by using a basic block flattening method to generate a target file, and the target file is converted into a machine code.

2. The method according to claim 1, characterized in that The function structure analysis of the acquired source code to be compiled is performed to obtain function feature information of the source code, including: Performing function classification on the acquired source code to be compiled to obtain function classification information to which the source code belongs; The calling relationship of the functions in the source code is obtained according to the function classification information, and the function feature information of the source code is generated according to the calling relationship.

3. The method according to claim 1, characterized in that The performing language analysis on the source code according to the function feature information to obtain the intermediate code corresponding to the source code includes: According to the function feature information, the source code is subjected to lexical analysis to obtain a lexical sequence corresponding to the source code; According to the lexical sequence, the source code is parsed using a context-free grammar, and a syntax tree corresponding to the source code is output; A semantic analysis is performed on the syntax tree corresponding to the source code to obtain a syntax tree with annotation information, and the syntax tree with annotation information is converted into an intermediate code.

4. The method according to claim 1, characterized in that The intermediate code corresponding to the source code is reorganized by using a basic block flattening method to control the intermediate code to generate a target file, including: According to the intermediate code corresponding to the source code, the intermediate code is opaquely obfuscated by using a dynamic opaque predicate insertion method to obtain an obfuscated intermediate code; Using a basic block flattening method to reorganize the control flow of the obfuscated intermediate code to obtain a flattened obfuscated intermediate code; The flattened obfuscated intermediate code is optimized by a pre-set multi-compiler to generate a target file corresponding to the source code.

5. The method according to claim 4, characterized in that The method of optimizing the flattened obfuscated intermediate code by using a pre-set multi-compiler to generate a target file corresponding to the source code includes: Compile and convert the flattened obfuscated intermediate code through a pre-set multi-compiler to obtain basic program blocks corresponding to the flattened obfuscated intermediate code; Analyzing the execution time of each basic program block through a code optimizer to generate an executable file corresponding to each basic program block; Generate a target file according to the executable files corresponding to each basic program block.

6. The method according to claim 5, characterized in that The step of generating a target file according to the executable files corresponding to the basic program blocks comprises: Executing the executable files corresponding to the basic program blocks through a performance analyzer, and recording the running time of the basic program blocks; Selecting the optimal code block corresponding to each basic program block according to the running time of each basic program block; The optimal code block corresponding to each basic program block is synthesized into a target file according to the running time.

7. The method according to any one of claims 1 to 6, characterized in that: The function feature information includes one or more of the following: function call relationship, function dependency, return value, function complexity, code reuse, function scale, exception handling and boundary conditions.

8. A compilation system based on an obfuscation mechanism, characterized in that: include: A structure analysis module, used to perform function structure analysis on the acquired source code to be compiled, and obtain function feature information of the source code; A language analysis module, used to perform language analysis on the source code according to the function feature information to obtain an intermediate code corresponding to the source code; A control flow reorganization module is used to reorganize the control flow of the intermediate code corresponding to the source code by using a basic block flattening method, generate a target file, and convert the target file into machine code.

9. The system according to claim 8, characterized in that The structural analysis module comprises: A function classification submodule is used to classify the acquired source code to be compiled, and obtain function classification information to which the source code belongs; The function calling submodule is used to obtain the calling relationship of the function in the source code according to the function classification information, and generate the function feature information of the source code according to the calling relationship.

10. The system according to claim 8, characterized in that The language analysis module comprises: A lexical analysis submodule, used for performing lexical analysis on the source code according to the function feature information to obtain a lexical sequence corresponding to the source code; A syntax analysis submodule, configured to perform syntax analysis on the source code using a context-free grammar according to the lexical sequence, and output a syntax tree corresponding to the source code; The semantic analysis submodule is used to perform semantic analysis on the syntax tree corresponding to the source code to obtain a syntax tree with annotation information, and convert the syntax tree with annotation information into an intermediate code.

11. The system according to claim 8, characterized in that The control flow reorganization module comprises: A code obfuscation submodule, used to perform opaque obfuscation on the intermediate code corresponding to the source code by using a dynamic opaque predicate insertion method to obtain an obfuscated intermediate code; A code flattening submodule, used for reorganizing the control flow of the obfuscated intermediate code by using a basic block flattening method to obtain a flattened obfuscated intermediate code; The code optimization submodule is used to optimize the flattened obfuscated intermediate code through a pre-set multi-compiler to generate a target file corresponding to the source code.

12. The system according to claim 11, characterized in that The code optimization submodule includes: A compiling and converting unit, used for compiling and converting the flattened obfuscated intermediate code through a pre-set multi-compiler to obtain basic program blocks corresponding to the flattened obfuscated intermediate code; A time analysis unit, used to analyze the execution time of each basic program block through a code optimizer, and generate an executable file corresponding to each basic program block; The target file generating unit is used to generate a target file according to the executable files corresponding to the basic program blocks.

13. The system of claim 12, wherein: The target file generating unit comprises: A file execution subunit, used for executing the executable files corresponding to the basic program blocks through a performance analyzer, and recording the running time of the basic program blocks; A code block selection subunit, used for selecting the optimal code block corresponding to each basic program block according to the running time of each basic program block; The file synthesis subunit is used to synthesize the optimal code block corresponding to each basic program block into a target file according to the running time.

14. The system according to any one of claims 8 to 13, characterized in that: The function feature information includes one or more of the following: function call relationship, function dependency, return value, function complexity, code reuse, function scale, exception handling and boundary conditions.

15. An electronic device, characterized in that: include: at least one processor and memory; The memory and the processor are connected via a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, a compilation method based on an obfuscation mechanism as claimed in any one of claims 1 to 7 is implemented.

16. A computing device readable storage medium, characterized in that: An execution program is stored thereon, and when the execution program is executed, a compilation method based on an obfuscation mechanism as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Compiler performance optimization method

    CN120491975A

  • A compiler performance optimization method

    CN120491975B