Code intermediate representation pre-training method based on symbolic execution track
By generating symbolic execution trajectories on the control flow graph and using compilation optimization and comparison learning techniques, the robustness problem of existing models when facing semantic equivalent variants is solved, and the understanding of IR structure changes in the code pre-trained model is improved.
Patent Information
- Application Number
- CN202510448019.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
AI Technical Summary
When existing code pre-trained models learn semantic information represented in the middle of the code, they lack local semantic supervision, making it difficult to deal with semantic equivalent variants caused by technologies such as compilation optimization and code obfuscation, resulting in the model not being robust enough when facing semantic equivalent control flow changes.
Symbol execution trajectories are generated by randomly walking on the control flow graph, semantic equivalents but inconsistent symbolic execution trajectory variants are generated using compilation optimization and code obfuscation techniques, and the semantic equivalence between symbolic execution trajectories is learned using Triplet Loss' contrast learning task.
The symbolic execution trajectory method can fully explore different execution branches of the program without actually executing code, reduce learning difficulty, enhance the robustness of neural networks for IR structure changes, and improve the accuracy of the model's semantic understanding.
Smart Images

Figure BDA0005353167050000031 
Figure BDA0005353167050000041 
Figure BDA0005353167050000064
Abstract
Description
Technical Field
[0001] The present invention relates to a method for pre-training code intermediate representation based on symbolic execution traces, which can provide semantic supervision of the execution locality in the pre-training based on code intermediate representation, and belongs to the technical field of code semantic understanding. Background Art
[0002] With the rapid development of code pre-training models, models such as CodeBERT and CodeT5 have achieved remarkable success and attracted wide attention. These pre-training models effectively improve the model's semantic understanding ability of code by training on large-scale unlabeled code datasets, and have achieved remarkable results in downstream tasks in multiple software engineering fields, such as code summarization, code generation, and defect detection. These results indicate that code pre-training models can significantly improve the efficiency and accuracy of programming language processing.
[0003] However, most of the existing source code pre-training models rely on surface text features for learning, lack a deep understanding of code execution semantics, and are easily affected by factors such as language types and text differences. Therefore, some researchers have begun to explore building pre-training models at the level of the intermediate representation language (IR) of compilers. Compared with source code, on the one hand, IR is independent of programming language types and provides a general code understanding platform; on the other hand, IR has a semantic-clear and limited instruction set definition, providing a concise and unified semantic understanding framework.
[0004] During the process of compiling source code into IR, technologies such as compilation optimization and code obfuscation often adjust the structure and execution flow of the code. After these technical processes, a piece of source code may be transformed into multiple semantically equivalent IR variants. Although these variants have no difference in the execution results of the program, there may be significant differences in their presentation forms. Therefore, a robust semantic understanding model should have the ability to generate similar representations for these different equivalent IR variants, so as to accurately understand the essential logic of the program, rather than simply relying on surface differences.
[0005] Existing work has proposed various pre-training techniques to learn the semantic information of IR, but they all have deficiencies in capturing semantic equivalence. Sequence learning based on surface text treats IR as a continuous sequence of tokens, learning the context relevance of the sequence, and is severely affected by text changes. Graph learning based on structural expressions extracts structures such as control flow graphs and data flow graphs from IR, learning the topological features of the graph, and is severely affected by changes in the topological structure. Contrastive learning based on overall similarity generates semantically similar variants for IR through methods such as compilation optimization, optimizing the distance between the representations of similar IRs at the function or program level, and it is difficult to handle a large number of local equivalence combination forms due to the lack of local semantic supervision. Summary of the Invention
[0006] Object of the Invention: To solve the problem of the lack of local semantic supervision in existing IR pre-training techniques, the present invention provides a pre-training method for intermediate representation of code based on symbolic execution traces, generating equivalence locally during execution. By randomly walking on the control flow graph, a function is expanded into a symbolic execution trace, and compilation optimization and code obfuscation techniques are used to generate symbolic execution trace variants that are semantically equivalent but have inconsistent representations. Then, a contrastive learning task is used to learn the semantic equivalence between symbolic execution traces. Through this pre-training method, the present invention attempts to add local semantic equivalence supervision to the IR pre-training model, making it more robust when facing control flow changes with semantic equivalence.
[0007] Technical Solution: A pre-training method for intermediate representation of code based on symbolic execution traces. An execution trace is the path taken during program execution, usually divided into two types: symbolic execution traces and dynamic execution traces. Symbolic execution traces represent variables through symbolic values and are composed of instruction sequences in the program, without actually executing the program. Dynamic traces, on the other hand, assign initial values during program execution and record the specific change values of variables during execution. First, we expand a function into a symbolic execution trace by randomly walking on the control flow graph; during this process, the number of back-edge visits is restricted to avoid repeated visits to loops, and symbolic constraints are used to represent control flow transfer conditions. Then, compilation optimization and code obfuscation techniques are used to generate symbolic execution trace variants that are semantically equivalent but have inconsistent representations. Finally, a contrastive learning task based on Triplet Loss is used to learn the representation similarity of equivalent symbolic execution traces.
[0008] A pre-training system for intermediate representation of code based on symbolic execution traces, comprising:
[0009] Symbolic Tracing Module Based on Random Walk: Expands a function into a symbolic execution trace by randomly walking on the control flow graph, restricts the number of back-edge visits to avoid repeated visits to loops, and uses symbolic constraints to represent control flow transfer conditions;
[0010] Equivalent Mutation Module Based on Compilation Optimization and Code Obfuscation: Generates symbolically executed trace variants that are semantically equivalent but representationally inconsistent through compilation optimization and code obfuscation techniques;
[0011] Contrastive Learning Module Based on Triplet Loss: Learns the representational similarity of equivalent symbolically executed traces using a contrastive learning task based on Triplet Loss.
[0012] The implementation process of the system is the same as the method and will not be elaborated here.
[0013] A computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above-mentioned method for pre-training the intermediate representation of code based on symbolically executed traces.
[0014] A computer-readable storage medium stores a computer program for executing the above-mentioned method for pre-training the intermediate representation of code based on symbolically executed traces.
[0015] Advantageous Effects: Compared with the prior art, the advantages of the method for pre-training the intermediate representation of code based on symbolically executed traces provided by the present invention are as follows:
[0016] 1) Symbolic tracing does not require actual code execution. In actual application scenarios, the code of a function usually contains calls to external functions or access to the program environment. Actual code execution requires constructing a complete machine state, which is difficult to achieve. Symbolic tracing can fully explore different execution branches of the program through simulated execution and completely retain the execution conditions through symbolic constraints.
[0017] 2) Symbolically executed traces can eliminate the influence of control flow on execution semantics. The prior art usually takes functions as the basic unit of semantic learning, but the execution result of a program is usually affected by the specific values of the input. Therefore, semantic understanding needs to consider different execution branches. Symbolically executed traces unfold the control flow of the function, enabling the neural network to only understand the semantics of consecutive instruction sequences and reducing the learning difficulty.
[0018] 3) Symbolic trace mutation enhances the robustness of the neural network to IR structure changes. Affected by technologies such as compilation optimization and code obfuscation, the same source code can correspond to different IR structures. By generating trace variants that are semantically equivalent but different in form through symbolic trace mutation and using contrastive learning techniques for learning, the neural network can better understand the structure changes of the IR. Description of the Drawings
[0019] Figure 1 is the workflow diagram of an embodiment of the present invention. Detailed Embodiments
[0020] The present invention will be further illustrated below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications made by those skilled in the art fall within the scope defined by the appended claims of this application.
[0021] A code intermediate representation pre-training method based on symbolic execution traces. An execution trace is the path taken during program execution, usually divided into two types: symbolic execution traces and dynamic execution traces. Symbolic execution traces represent variables through symbolic values and are composed of instruction sequences in the program, without actually executing the program. Dynamic traces, on the other hand, assign initial values during program execution and record the specific changing values of variables during execution. First, the function is expanded into a symbolic execution trace by randomly walking on the control flow graph; during this period, the number of back edge visits is restricted to avoid repeated access to loops, and symbolic constraints are used to represent control flow transfer conditions. Then, semantic-equivalent but representationally inconsistent symbolic execution trace variants are generated through compilation optimization and code obfuscation techniques. Finally, a contrastive learning task based on Triplet Loss is used to learn the representational similarity of equivalent symbolic execution traces. The pre-training method is described in detail as follows:
[0022] 1) Symbolic tracing based on random walk
[0023] The following gives the pseudocode for generating a symbolic execution trace by random walk. This algorithm takes an IR function p as input and outputs a symbolic execution trace t.
[0024]
[0025]
[0026] Lines 1-5 initialize the local variables required by the algorithm. The list t is used to record the instruction sequence of the basic blocks visited during the walk, the variable time is used to record the current time, the dictionary firstVisit is used to record the first visit time of each basic block, the set visitedBackEdge is used to record the visited back edges, and the basic block block represents the currently visited basic block;
[0027] Line 6 loops through the basic block block until it is empty;
[0028] Lines 7-10 determine whether the basic block block is visited for the first time. If so, record its first visit time;
[0029] Line 11 records the instruction sequence of the basic block block;
[0030] Line 12 initializes the candidate list of accessible successor basic blocks;
[0031] Lines 13 - 17 traverse the successor basic blocks of the basic block block, determine whether it is the basic block pointed to by the visited back edge, and if not, add it to the candidate list;
[0032] Line 18 determines whether the candidate list is empty. If it is not empty, the walk continues;
[0033] Lines 19 - 20 randomly select a basic block from the candidate list as the next access target and record the condition of this control flow transfer;
[0034] Lines 21 - 23 determine whether this access is a back edge access. If so, record it;
[0035] During the walk, symbolic constraints are used to represent the control flow transfer conditions. The current basic block ends with the instruction "bri1"
[0036] <cond>,label <iftrue>,label <iffalse>"At the end, if the next wandering target selection label <iftrue>For the pointed basic block, record the symbolic constraint "assume" <cond>”In the caret execution trajectory, if the next roaming target selection label
[0037] <iffalse>If it points to the basic block, record the symbolic constraint "assume!" <cond>”in the caret execution trace; the current basic block ends with the instruction "switch" <intty> <value>,label <defaultdest> [ <intty> <val>,label <dest>...]” At the end, if the next traversal target selection label <dest>the basic block pointed to, record the symbolic constraint "assume <value> ==
[0038] <val>”into the caret execution trace.
[0039] 2) Equivalent mutation based on compilation optimization and code obfuscation
[0040] Considering that the symbolic execution trace is a sequence of continuously executed instructions without control flow structures, the following semantic equivalent transformation operations are selected to mutate the symbolic execution trace:
[0041] 2.1) Constant Propagation, which propagates constants known at compile time to various parts of the program to eliminate unnecessary variable assignments and calculations. The unfolding of the control flow may make more variables become known constants, and eliminating unnecessary calculations through propagation can help the model understand the arithmetic calculation results.
[0042] 2.2) Dead Code Elimination, which deletes useless variables, unreachable code paths, or calculations that never affect the program state (i.e., dead code) through static analysis. The unfolding of the control flow may make the calculations of some variables useless, and eliminating them can help the model understand the use of variables.
[0043] 2.3) Instruction Combining, which combines multiple inefficient instructions into fewer equivalent instructions by reducing unnecessary instructions (e.g., a = 2 * t; b = 2 * a; → b = 4 * a;). Combining instructions can help the model understand the arithmetic equivalence between instructions.
[0044] 2.4) Mem2reg, which promotes frequently accessed variables from memory (stack) to registers to reduce the overhead of memory access. The conversion between memory and registers can help the model understand the equivalent use between memory addresses and temporary variables.
[0045] 2.5) Instruction Substitution, which unfolds one instruction into several equivalent instructions according to predefined rules (e.g., a = b + c; → t = -c; a = b - t;). Instruction Substitution can be regarded as the reverse operation of Instruction Combining and can also help the model understand the arithmetic equivalence between instructions.
[0046] 3) Contrastive learning based on Triplet Loss
[0047] Use contrastive learning to optimize the representation similarity between equivalent symbolic execution traces to learn their semantic equivalence. By minimizing the distance between similar samples and maximizing the distance between dissimilar samples, effective feature representations are learned.
[0048] With the triplet (t, t + , t - ) As the basic unit of model input, where \(t, t\) + , \(t\) - respectively represent the symbolic execution trace, the symbolic execution trace equivalent to \(t\), and the symbolic execution trace not equivalent to \(t\). Assume that a mini - batch during model training is
[0049] which contains \(n\) randomly sampled triple inputs. Also assume that the symbolic execution trace obtains a vector representation after being embedded by the model. Then calculate the loss function of contrastive learning as follows:
[0050]
[0051] where \(d\) represents the dimension of the vector representation , \(l(t\) i ) represents the label of the symbolic execution trace \(t\) i , \(l(t\) i ) \(\neq l(t\) j ) means that the symbolic execution trace \(t\) i and the symbolic execution trace \(t\) j are not equivalent.
[0052] Taking the construction of an LLVM IR pre - trained model for binary similarity detection as an example, illustrate the specific application of the pre - training method of code intermediate representation based on symbolic execution traces. Binary similarity detection is a basic binary analysis ability, and its goal is to measure the similarity between two binary functions. This application consists of the following three stages:
[0053] 1) Data processing stage
[0054] In this stage, decompile the binary file into the LLVM IR required by the model and calculate the labels required for contrastive learning of symbolic execution traces.
[0055] 1. LLVM IR dataset collection
[0056] In this step, several widely used open-source C / C++ projects on GitHub are first collected, and then they are separately compiled into binary files using different compilation environments. Among them, there are five machine platforms (x86, x64, arm32, arm64, mips32), six compilation optimization options (O0, O1, O2, O3, Os, Ofast), and eight compiler versions (gcc-4.9.4, gcc-5.5.0, gcc-7.3.0, gcc-9.4.0, clang-4.0.0, clang-5.0.2, clang-7.0.1, clang-9.0.1). Then, RetDec is used to decompile them into LLVM IR files.
[0057] 2. LLVM IR Symbol Tracing
[0058] In this step, symbol tracing is performed on LLVM IR functions to obtain symbol execution traces. Given the control flow graph of an LLVM IR function, the entry basic block of the function is used as the starting point for traversal. Each time, a basic block that is not pointed to by an already visited back edge is randomly selected from the successor basic blocks of the current basic block as the next traversal target, and the function call record is used to control the conditions for control flow transfer. Finally, the instruction sequences of the basic blocks on the traversal path are sequentially concatenated as the symbol execution trace.
[0059] Two types of LLVM IR control flow transfer instructions are mainly considered:
[0060] i.br instruction: Its format is "br i1 <cond>,label <iftrue>,label <iffalse>", whose semantics is according to the conditional variable <cond>The true value selects the next basic block to execute. If true, jump to the label <iftrue>Execute the basic block pointed to; if false, jump to the label <iffalse>The basic block execution pointed to. When walking, if the next walking target selects a label
[0061] <iftrue>If it is the basic block pointed to, then insert the function call "call i1 @assumeTrue(i1 <cond>)” into the symbolic execution trace; if the next traversal target selects a label <iffalse>For the pointed basic block, record the symbolic constraint "calli1
[0062] @assumeFalse(i1 <cond>) into the caret execution trace.
[0063] ii. switch instruction: Its format is "switch" <intty> <value>,label <defaultdest> [ <intty> <val>,label <dest>...]”, whose semantics is according to the conditional variable <value>The value selects the next basic block to execute, provided that its value matches one of the constants in the jump table <val>If they are equal, jump to the label at the constant index <dest>The basic block pointed to is pointed to; if they are all not equal, jump to the label <defaultdest>The execution of the pointed basic block. During the walk, if the next walk target selects a label <dest>For the basic block pointed to, insert the function call "call i1 @assumeEqual( <intty> <value> , <intty> <val> ) <value> == <val>” into the symbolic execution trace.
[0064] 3. Mutation of the LLVM IR Symbolic Execution Trace
[0065] In this step, the symbolic execution trace of LLVM IR is mutated to generate semantically equivalent symbolic execution trace variants. The existing compilation optimization phases and code obfuscation phases in the LLVM compilation system are used for mutation, including constant propagation (constprop), dead code elimination (adce), instruction combination (instcombine), memory to register (mem2reg), and instruction substitution (instsubstitution).
[0066] 2) Model Training Phase
[0067] In this phase, a model for binary similarity detection is pre-trained.
[0068] 1. Model Construction
[0069] Since the symbolic execution trace is a sequential instruction sequence without control flow structures, the sequence model Transformer is used to learn the token sequence of the symbolic execution trace. The special token [CLS] is added to the beginning of the token sequence of the symbolic execution trace, and its final vector representation is used as the vector representation of the symbolic execution trace.
[0070] 2. Input Representation
[0071] The input of the model is a sequence of tokens. First, a Function Pass in the LLVM compilation system is implemented to output an LLVM IR function as a sequence of tokens. Since many important semantic information and structures are lost or deconstructed during the compilation of source code to binary code (for example, function names are discarded and function formal parameters are converted to registers or memory addresses), it becomes less important to decompile and obtain information such as function names and variable names in LLVM IR, and it will instead pose an obstacle to semantic understanding. Therefore, the LLVM IR input is simplified by only retaining those parts related to semantics to reduce the understanding burden of the model. Specifically, given an LLVM IR instruction, first remove the meta-annotations, byte alignment, etc. information, and only retain the three parts of the opcode, operands, and result value of the instruction. For example, the instruction "store i1 false,i1*%storemerge,align 1,!insn.addr!3950” is simplified to "store i1 false i1*%storemerge”. Then standardize the variable names and function names. The variable names are uniformly named in the form of v{i}, and the function names are uniformly named in the form of f{i}. For example, the instruction "store i1 false i1*%storemerge.reg2mem” is standardized to "store i1 false i1*%v0”. Finally, to reduce the length of the instruction and improve the speed of the model processing the sequence, the variable types in the instruction are extracted into a separate parallel sequence. For example, the instruction "store i1 false i1*%v0” is split into the token sequence "store false%v0” and the type sequence " <null>i1 i1*”。
[0072] 3. Two-stage training
[0073] Learn the representation of LLVM IR functions through two-stage training of symbolic execution trace learning and function transfer learning:
[0074] i. Symbolic execution trace learning: The input to the model at this stage is the token sequence of the symbolic execution trace. We use the masked language model task to capture the context relevance of the symbolic execution trace. The masked language model learns the context information of the sequence by randomly masking tokens in the input sequence and having the model recover them. Suppose a basic block is composed of a token sequence {t1, t2, …, t l}, denoted as represents the other tokens in the sequence except t i . M is the set of sequence positions of the masked tokens, and its objective function can be expressed as:
[0075]
[0076] Learn the similarity of symbolic execution trace representations using the contrastive learning task. The model takes a triple (t, t + , t - ) as the basic unit of input, where t, t + , t - represent the symbolic execution trace, the symbolic execution trace equivalent to t, and the symbolic execution trace not equivalent to t, respectively. Suppose a mini-batch during model training is which contains n randomly sampled triple inputs. Also suppose that the symbolic execution trace after being embedded by the model obtains a vector representation Then the loss function of contrastive learning is calculated as follows:
[0077]
[0078] where d represents the dimension of the vector representation , l(t i ) represents the label of the symbolic execution trace t i , and l(t i ) ≠ l(t j ) means that the symbolic execution trace t i and the symbolic execution trace t j are not equivalent.
[0079] The final loss function of this stage is the sum of the above two training tasks, that is:
[0080] L = L MLM + L STC
[0081] ii. Function transfer learning: The input of the model at this stage is the token sequence of the function, and the model is trained using the same loss function as the contrastive learning with symbolic execution traces.
[0082] 3) Model inference stage
[0083] In this stage, the similarity between two given binary functions is calculated. First, we decompile them into LLVM IR through RetDec, then input them into the pre-trained model to obtain their vector representations, and finally calculate the cosine similarity of the vector representations of the two functions as the similarity. Suppose the vector representations of the two functions are E1 and E2, then the similarity calculation is as follows:
[0084]
[0085] Obviously, those skilled in the art should understand that each step of the above-mentioned pre-training method for code intermediate representation based on symbolic execution traces in the embodiments of the present invention or each module of the pre-training system for code intermediate representation based on symbolic execution traces can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.< / null> < / val> < / value> < / val> < / intty> < / value> < / intty> < / dest> < / defaultdest> < / dest> < / val> < / value> < / dest> < / val> < / intty> < / defaultdest> < / value> < / intty> < / cond> < / iffalse> < / cond> < / iftrue> < / iffalse> < / iftrue> < / cond> < / iffalse> < / iftrue> < / cond> < / val> < / value> < / dest> < / dest> < / val> < / intty> < / defaultdest> < / value> < / intty> < / cond> < / iffalse> < / cond> < / iftrue> < / iffalse> < / iftrue> < / cond>
Claims
1. A pre-training method for intermediate representation of code based on symbolic execution traces, characterized in that Including: 1) Symbol tracing based on random walk; Expand the function into a symbolic execution trace by randomly walking on the control flow graph, limit the number of back edge visits to avoid repeated loop visits, and use symbolic constraints to represent the control flow transfer conditions; 2) Equivalent mutation based on compilation optimization and code obfuscation; Generate symbolic execution trace variants that are semantically equivalent but have inconsistent representations through compilation optimization and code obfuscation techniques; 3) Contrastive learning based on Triplet Loss; Use the contrastive learning task based on Triplet Loss to learn the representation similarity of equivalent symbolic execution traces.
2. The method for pre-training the intermediate representation of code based on the symbolic execution trace according to claim 1, wherein The pre-training method is based on the intermediate representation of the LLVM compilation system and is used to guide the symbolic execution trace representation learning of the pre-training model based on the LLVM intermediate representation.
3. The code intermediate representation pre-training method based on symbolic execution traces according to claim 1, characterized in that, In the above 1), given the control flow graph of a function, use the entry basic block of the function as the starting point of the walk. Each time, randomly select a basic block that is not pointed to by an already visited back edge from the successor basic blocks of the current basic block as the next walk target, and use symbolic constraints to record the conditions of the control flow transfer. Finally, sequentially splice the instruction sequences of the basic blocks on the walk path as the symbolic execution trace.
4. The method for pre-training the intermediate representation of code based on symbolic execution traces according to claim 1 or 2, characterized in that In the random walk, before the walk starts, initialize the time variable time to 0 and the set of visited back edges visitedBackEdges to an empty set; whenever a basic block b is first visited, record its first visit time as firstVisit(b) = time and increment the time variable time = time + 1; when selecting the next walk target, first obtain all the successor basic blocks of the current basic block b. The basic block succ that satisfies the condition that the edge (b, succ) is not included in the set of visited back edges visitedBackEdges is the basic block pointed to by an unvisited back edge. Randomly select one of them as the next walk target; if the selected basic block satisfies the condition firstVisit(succ) < firstVisit(b), then record the back edge of this visit and add the edge (b, succ) to the set of visited back edges visitedBackEdges; if there are no successor basic blocks that satisfy such conditions, it means that the exit basic block of the function has been reached or an endless loop has been entered, and this walk ends.
5. The method for pre-training the intermediate representation of code based on symbolic execution traces according to claim 1 or 2, wherein In the random walk, two control flow transfer conditions are considered: the current basic block ends with an instruction "br i1 <cond>,label <iftrue>,label <iffalse>"At the end, if the next roaming target selection label <iftrue>the basic block it points to, record the symbolic constraint "assume <cond>"In the caret execution trajectory, if the next traversal target selection label <iffalse>For the pointed basic block, record the symbolic constraint "assume! <cond>”in the caret execution trace; the current basic block ends with the instruction "switch <intty> <value>,label <defaultdest> [ <intty> <val>,label <dest>...]” At the end, if the next roaming target selection label <dest>the basic block pointed to, record the symbolic constraint "assume <value> == <val>"Insert into the symbolic execution trace.< / val> < / value> < / dest> < / dest> < / val> < / intty> < / defaultdest> < / value> < / intty> < / cond> < / iffalse> < / cond> < / iftrue> < / iffalse> < / iftrue> < / cond> 6. The method for pre-training the intermediate representation of code based on the symbolic execution trace according to claim 1, characterized in that In the above 2), select transformation operations in compilation optimization and code obfuscation to perform semantically equivalent mutations on the symbolic execution trace, including five transformation operations: constant propagation, dead code elimination, instruction merging, memory-to-register conversion, and instruction replacement.
7. The method for pre-training the intermediate representation of code based on the symbolic execution trace according to claim 1, wherein In the above (3), the triple (t, t + , t - ) is used as the basic unit of the neural network input, where t, t + , t - represent the symbolic execution trace, the symbolic execution trace equivalent to t, and the symbolic execution trace not equivalent to t respectively; let a sample batch during neural network training be which contains n randomly sampled triple inputs; the symbolic execution trace obtains a vector representation after being embedded by the neural network The loss function of contrastive learning is calculated as follows: Among them, d represents the dimension of the vector representation ; l(t i ) represents the label of t i of the symbolic execution trace; l(t i ) ≠ l(t j ) indicates that the symbolic execution trace t i and the symbolic execution trace t j are not equivalent.
8. A code intermediate representation pre-training system based on symbolic execution traces, characterized in that, Including: Symbol tracing module based on random walk: Expand the function into a symbolic execution trace by randomly walking on the control flow graph, limit the number of back edge visits to avoid repeated loop visits, and use symbolic constraints to represent the control flow transfer conditions; Equivalent mutation module based on compilation optimization and code obfuscation: Generate symbolic execution trace variants that are semantically equivalent but have inconsistent representations through compilation optimization and code obfuscation techniques; Contrastive learning module based on Triplet Loss: Use the contrastive learning task based on Triplet Loss to learn the representational similarity of equivalent symbolic execution traces.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the pre-training method for the intermediate code representation based on symbolic execution traces as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program that executes the pre-training method for the intermediate code representation based on symbolic execution traces as described in any one of claims 1-7.