Code intermediate representation pre-training method based on semantic similarity
By introducing the input and output equivalence definition and semantic similarity learning of basic blocks in IR pre-training, the problem of insufficient local semantic supervision in the prior art is solved, and the robustness and generalization ability of the neural network are improved.
Patent Information
- Application Number
- CN202510447641.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
AI Technical Summary
The existing IR pre-training technology lacks local semantic supervision of the intermediate representation of the code, resulting in insufficient robustness in the face of changes in basic block instructions of semantic equivalents, and prone to error propagation.
By defining the input and output equivalence of the basic block, quantifying the semantic similarity, and using the mean square error loss function for learning, combining the LLVM compilation system and the program validator Boogie to train the semantic similarity of the basic block.
It improves the robustness and generalization ability of neural networks to IR basic block changes, can accurately understand the essential logic of the program, and reduce error propagation.
Smart Images

Figure CN120295637A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a pre-training method for code intermediate representation based on semantic similarity, which can provide semantic supervision at the local level of basic blocks in pre-training based on code intermediate representation, including the definition of basic block semantic similarity based on input-output equivalence and the learning pipeline of basic block semantic similarity. This pre-training method enables the model to have better robustness to local semantic changes and belongs to the technical field of pre-training in deep learning. Background Art
[0002] With the rapid development of code pre-training models, models such as PLBART and UniXcoder have achieved remarkable results and attracted wide attention. These models have significantly improved the model's ability to understand code semantics through pre-training on large-scale unlabeled code datasets and have performed outstandingly in downstream tasks in multiple software engineering fields, such as code summarization, code generation, and defect detection. Their success shows that code pre-training models can significantly improve the efficiency and accuracy of programming language understanding tasks.
[0003] However, most code pre-training models only use or mainly use software source code as training data, which has certain limitations: lacking a deep understanding of code execution semantics and being easily affected by factors such as language types and programming style differences. For this reason, some researchers have begun to explore building pre-training models at the level of intermediate representation language (IR). Compared with source code, on the one hand, IR is independent of programming language types and provides a general understanding platform; on the other hand, IR has a set of instruction definitions with clear semantics and a limited number, providing a concise and unified understanding framework.
[0004] During the process of compiling source code into IR, techniques such as compilation optimization and code obfuscation often adjust the structure and execution flow of the code to improve program performance or protect software property rights. After these technical processes, a source code may be transformed into multiple semantically equivalent IR variants. Although these variants have no difference in the execution results of the program, they may have significant differences in their presentation forms. Therefore, a robust semantic understanding model should be able to generate similar representations for these different equivalent IR variants, so as to accurately understand the essential logic of the program, rather than relying solely on surface differences.
[0005] Among them, a relatively common semantic equivalence variant is different instruction sequences that calculate the same expression in a basic block. Existing IR pre-training techniques all have deficiencies in dealing with this problem. Sequence learning based on surface text treats IR as a continuous sequence of tokens and is severely affected by syntactic changes that are semantically irrelevant. Graph learning based on structural expressions extracts structures such as control flow graphs and data flow graphs from IR and learns the topological features of the graphs, which is severely affected by computational changes of semantic equivalence. Contrastive learning based on overall similarity generates semantically similar variants for IR through methods such as compilation optimization, and optimizes the distance between the representations of similar IRs at the function or program level. Due to the lack of local semantic supervision, it is difficult to handle a large number of equivalent combination forms and is prone to the problem of error propagation. Summary of the Invention
[0006] Object of the Invention: To solve the problem of the lack of local semantic supervision in existing IR pre-training techniques, the present invention introduces semantic difference quantification locally in basic blocks, defines the proportion of equivalent output variables based on input-output equivalence as the semantic similarity of two basic blocks, and uses the mean squared error loss function to learn the similarity of basic block representations. Through this pre-training method, the present invention attempts to add local semantic equivalence supervision to the IR pre-training model, making it more robust when facing semantic equivalent basic block instruction changes.
[0007] Technical Solution: A code intermediate representation pre-training method based on semantic similarity. First, it defines the semantic similarity of basic blocks based on input-output equivalence. By enumerating the equivalence relationships of input variables, checking the equivalence relationships of output variables, and calculating the proportion of the largest equivalent output variables as the semantic similarity of basic blocks. Then, it designs a learning pipeline based on the semantic similarity, including extracting basic blocks from the code intermediate representation; translating the instruction semantics of basic blocks into the intermediate verification language of a program verifier; sampling candidate basic block pairs to be calculated for similarity; using the program verifier to calculate the similarity of candidate pairs in parallel; modeling the learning of similarity as a regression task and training the model using the mean squared error loss. The pre-training method is based on the intermediate representation of the LLVM compilation system and the program verifier Boogie, and is used to guide the basic block representation learning of the pre-training model based on the LLVM intermediate representation. The pre-training method is described in detail as follows:
[0008] 1) Definition of the semantic similarity of basic blocks based on input-output equivalence, including the following content:
[0009] 1.1) Give basic semantic notations
[0010] A program state σ consists of a pair (l, values), where l ∈ Loc is a program location (i.e., a moment in program execution), and values: Var → Val is a function from variables to values that maps the variables of the program at location l to their concrete values. The set of all possible states of program P is denoted as Σ P .
[0011] An execution trace is a sequence of states <σ1, σ2, …, σ n >>, which describes the state transition path of a single execution of the program. Here, denotes a sequence consisting of any number of elements in Σ P . The set of all possible traces of program P is denoted as Define the function first: and the function last: to return the first and last states in the trace, respectively
[0012] Based on this, considering all possible input value combinations of basic block b, b can be represented as a set of a series of traces, that is Additionally, define in(b), out(b), and var(b) to represent the input variables, output variables, and the set of all variables of basic block b, respectively. Among them, input variables are variables used directly without being defined, output variables are variables directly defined, and all variables are the union of the former two
[0013] 1.2) Define variable matching
[0014] Given two basic blocks b1 and b2, taking the variables var(b1) and var(b2) in the two basic blocks as two independent point sets, considering the legal equivalent possibilities between variables (i.e., variable type compatibility), connect the nodes in the two point sets. In this way, a bipartite graph can be formed. The matching on this bipartite graph is called the variable matching between the two basic blocks, denoted as γ. It is easy to see that In this matching, any variable is connected to at most one variable in another basic block, which describes the possible equivalent relationships between variables in the two basic blocks. Later, based on the matching γ, the equivalence between inputs will be assumed and the equivalence between outputs will be verified
[0015] 1.3) Then give the definitions of state and trace equivalence
[0016] Given two states σ1 and σ2 and a matching γ, if there is then these two states are said to be equivalent with respect to the matching γ, denoted as σ1 ≡ γσ2. Intuitively, this means that the program assigns the same value to the corresponding variables in γ in these two states. Further, given two traces π1, π2, if their last states are equivalent for the match γ, i.e., last(π1) ≡ γ last(π2), the two traces are said to be equivalent for the match γ, denoted as π1 ≡ γ π2.
[0017] 1.4) Define the input-output equivalence of basic blocks
[0018] Given two basic blocks b1, b2 and a match γ, if:
[0019] i. The input variables of basic blocks b1, b2 are all connected in the match γ;
[0020] ii. Any pair of input-equivalent trace paths is equivalent for the match γ, i.e.,
[0021]
[0022]
[0023] The two basic blocks are said to be equivalent for the match γ. Intuitively, this means that there are equivalent matches for the inputs of basic blocks b1, b2 in the match γ, and the same output values are assigned to the corresponding variables in the match γ in all possible executions.
[0024] 1.5) By extending the above input-output equivalence, the semantic similarity of basic blocks can be defined
[0025] Based on the above definition of input-output equivalence, if the proportion of equivalent output variables in the match γ is considered, the semantic similarity of basic blocks can be measured, and its definition is as follows:
[0026]
[0027] where out(b1) γ , out(b2) γ represent the sets of output variables of basic blocks b1 and b2 included in the match γ, respectively, and are defined as Intuitively, this definition of semantic similarity enumerates the possibilities of the match γ that makes basic blocks b1, b2 equivalent, selects the match with the most equivalent output variables, and defines the proportion of equivalent output variables in all output variables as the semantic similarity of the two basic blocks.
[0028] 2) Semantic similarity learning pipeline
[0029] 2.1) Extract basic blocks from the intermediate representation of the code
[0030] Use the tool llvm-extract in the LLVM compilation system to extract basic blocks from the LLVM IR, and extract each basic block into a function. llvm-extract will automatically identify local variables used directly without definition in the basic block as the formal parameters of the extracted function.
[0031] 2.2) Translate the semantic meaning of basic block instructions into an intermediate verification language
[0032] Translate the instructions of the extracted basic blocks from the LLVM IR format into the intermediate verification language BoogieIVL of Boogie. In terms of type translation, translate the integer type of LLVM IR into the integer type of Boogie IVL; translate the floating-point type of LLVM IR into the real type of Boogie IVL; translate the user-defined structure type of LLVM IR into the special custom type "struct" of BoogieIVL. In terms of instruction translation, translate the integer arithmetic instructions of LLVM IR into the built-in integer operations of BoogieIVL; translate other instructions of LLVM IR into uninterpreted functions in Boogie IVL. In terms of function calls, translate all external function calls into uninterpreted functions in Boogie IVL.
[0033] An uninterpreted function is a concept in program verification, usually used to model and abstract functions in a program. It does not care about the specific implementation of the function, models the function as a black box, and only cares about the input-output relationship of the function. Only when the arguments of the uninterpreted function calls are the same, the verifier will consider the call results to be the same. In fact, the semantics of function calls will be partially reflected in the code for argument preparation and return value usage. Therefore, for two function calls, if their argument preparation processes are exactly the same and the return values are used in the same way, the two are regarded as equivalent.
[0034] 2.3) Sample pairs of basic blocks for which similarity is to be calculated
[0035] Sample pairs of basic blocks for which similarity is to be calculated from the set of basic blocks. In the traditional neural network training process, the composition of a mini-batch (i.e., which basic blocks are used together for gradient calculation and update) is usually randomly sampled during training. If pairs are sampled completely randomly among all basic blocks, it will lead to a low hit rate during training (the basic block pairs are not in the same mini-batch). Therefore, first sample the composition of each training round's mini-batch before training, and then sample basic block pairs separately within each mini-batch for similarity calculation.
[0036] 2.4) Parallelly call the program verifier to calculate similarity
[0037] Parallelly call Boogie to calculate the similarity of basic block pairs. Encode variable matching as a query to the program verifier to calculate the semantic similarity of basic blocks. The query mainly consists of three parts, including the equivalence assumption of input variables under the matching γ, the instruction semantic constraints of basic blocks, and the equivalence assertion of output variables under the matching γ. After inputting the query into Boogie, Boogie will return the correctness of the assertion statement in the query, that is, the equivalence of output variables, and based on this, the semantic similarity can be calculated according to the aforementioned formula. It should be noted that to reduce the search space of the matching γ, only the case where the matching γ is a maximum matching is considered in the actual query.
[0038] 2.5) Similarity regression learning based on mean squared error loss
[0039] Suppose a dataset of basic block semantic similarities of size N is obtained by sampling calculation Among them, sim i are respectively the basic block pair and the semantic similarity of the i-th data. Use the mean squared error loss function to train the model:
[0040]
[0041] Among them, M represents the embedding model of basic blocks, which converts the basic block input into a vector form; Equation calculates the cosine similarity of the representations of two basic blocks.
[0042] Technical effects: Compared with the prior art, the pre-training method for intermediate representation of code based on semantic similarity provided by the present invention has remarkable effects in the following aspects:
[0043] 1) Provide an accurate understanding of the semantics of basic blocks. By introducing the definition of basic block semantic similarity based on input-output equivalence, the semantic similarity between different basic blocks can be quantified, providing more refined similarity guidance for neural network training, making its performance more accurate during the learning process, and avoiding error propagation caused by the lack of local semantic equivalence awareness.
[0044] 2) Improve the robustness of the neural network to changes in IR basic blocks. By providing local semantic supervision for basic blocks, the neural network can accurately capture semantically equivalent basic block variants after technical processes such as code optimization and obfuscation, and has a better understanding ability of the semantics of basic blocks.
[0045] 3) Enhance the generalization ability of the neural network. Through training for different semantic variants, the neural network can not only adapt to known code changes, but also better handle unknown semantic variants, thus improving the generalization ability in actual applications. Description of the Drawings
[0046] Figure 1 This is the flowchart of the embodiment of the present invention. Detailed implementation manner
[0047] The present invention will be further clarified below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications of the present invention by those skilled in the art fall within the scope defined by the appended claims of this application.
[0048] A code intermediate representation pre-training method based on semantic similarity defines the semantic similarity of basic blocks based on input-output equivalence. By enumerating the equivalence relationships of input variables, verifying the equivalence relationships of output variables, and calculating the proportion of the maximum equivalent output variables, the semantic similarity of basic blocks is obtained. Based on the learning pipeline defined by the semantic similarity of basic blocks, it specifically includes extracting basic blocks from the code intermediate representation; translating the instruction semantics of basic blocks into the intermediate verification language of a program verifier; sampling candidate basic block pairs to be calculated for similarity; using the program verifier to calculate the similarity of candidate pairs in parallel; modeling the learning of similarity as a regression task, and training the model using the mean square error loss. The pre-training method is described in detail as follows:
[0049] 1) LLVM IR basic block extraction
[0050] In this step, the llvm-extract tool of the LLVM compilation system is used to extract each basic block of each function in the LLVM IR file into a separate function. Llvm-extract will automatically identify the local variables that are used directly without definition in the basic block as the formal parameters of the extracted function. Then, define the formal parameters of the extracted function as the input variables of the original basic block, and the local variables of the extracted function as the output variables of the original basic block. Only basic blocks with more than 5 instructions are extracted when extracting basic blocks to ensure that the extracted basic blocks have a certain degree of complexity and representativeness.
[0051] 2) Translation of LLVM IR to Boogie IVL
[0052] In this step, rules are manually constructed to implement a Function Pass of the LLVM compilation system, so as to translate the basic blocks extracted as functions into the Boogie IVL form. In terms of type translation, the integer types in LLVM IR are translated into integer types in Boogie IVL; the floating-point types in LLVM IR are translated into real types in Boogie IVL; the user-defined structure types in LLVM IR are translated into the special custom type "struct" in Boogie IVL. In terms of instruction translation, the integer arithmetic instructions in LLVM IR are translated into built-in integer operations in Boogie IVL; other instructions in LLVM IR are translated into uninterpreted functions in Boogie IVL. In terms of function calls, all external function calls are translated into uninterpreted functions in Boogie IVL.
[0053] 3) Sample basic block pairs for which the similarity is to be calculated
[0054] In this step, first, the function composition of each batch is sampled according to the batch size of the training, and then several basic block pairs are randomly sampled within the set of basic blocks of the functions in each batch as candidate pairs for which the semantic similarity is to be calculated.
[0055] 4) Call Boogie in parallel to calculate the similarity
[0056] In this step, several processes are enabled in parallel to call Boogie to calculate the semantic similarity. The following gives the pseudocode for using Boogie to query and calculate the semantic similarity. This algorithm takes the Boogie IVL functions of two basic blocks b1 and b2 as input and outputs their semantic similarity.
[0057]
[0058] Line 1 initializes the variable maxSim to record the maximum semantic similarity.
[0059] Lines 2 - 3 enumerate all possible variable matches and create a new Boogie procedure p as a query carrier for each match.
[0060] Lines 4 - 6 add the assumption of input equality in the match, that is, the assume statement, to achieve equivalent constraints on the input.
[0061] Line 7 adds the instruction sequence of the basic block to the procedure p.
[0062] Lines 8 - 10 add the assertion of output equality in the match, that is, the assert statement, to verify their equivalence at the end of the program.
[0063] Submit the process p to Boogie for solving at line 11.
[0064] At lines 12 - 14, determine whether two basic blocks are equivalent for matching. If they are equivalent, calculate their semantic similarity and update the maximum value.
[0065] 5) Neural network training
[0066] In this step, the neural network model is trained based on the existing data. Assume that in step 3 mentioned above, a batch of functions {f1, f2, …, f n} of size n is sampled, and m pairs of basic block pairs are sampled from them to calculate the semantic similarity. We use four pre - training tasks to train the model, including:
[0067] 1. Semantic Similarity Regression (SSR) of basic blocks
[0068] Assume that the basic block obtains the representation after passing through the model. We use the mean squared error loss to learn the cosine similarity of the basic block representations. The objective function can be expressed as:
[0069]
[0070] 2. Masked Language Model (MLM)
[0071] The Masked Language Model learns the context information of the sequence by randomly masking the tokens in the input sequence and letting the model restore them. Assume that the basic block is composed of a token sequence {t1, t2, …, t l} of length l, and denotes the other tokens in the sequence except t i . M is the set of sequence positions of the masked tokens. The objective function can be expressed as:
[0072]
[0073] The base of log is 2.
[0074] 3. Flow - Type Prediction (FTP)
[0075] Flow - Type Prediction randomly masks the edge features in the control - flow graph and lets the model predict to learn the topological features of the control - flow graph. Assume that the control - flow graph contains s basic - block nodes in total, and the edge features are {φ 11 , φ 12 , …, φ ss}. The edge φ ij obtains the representation e ij after passing through the model. We use a fully - connected layer to predict the original features of the edge, that is, where W, b are the parameters of the fully - connected layer. Assume Gmask Denote the obfuscated control flow graph as, and the set of obfuscated edge features as F. The objective function can be expressed as:
[0076]
[0077] The base of the log is 2.
[0078] 4. Momentum Contrastive Learning (MoCo)
[0079] Momentum contrastive learning is a contrastive learning technique. It maintains a dynamic and large negative sample library through the momentum update mechanism, narrows the representation distance between the original sample and the positive sample, and widens the representation distance between the original sample and the negative sample to learn the similarity of representations. For each function f in a mini-batch i , we randomly sample a corresponding function generated under different compilation optimization options in step 1 as the positive sample f i '; use the negative sample library maintained in momentum contrastive learning as the negative sample. Assume that the function f i j obtained through the model has the representation The set of negative sample representations maintained by momentum contrastive learning is K. The objective function can be expressed as:
[0080]
[0081] where τ is the temperature parameter of momentum contrastive learning, and the base of the log is 2.
[0082] Finally, the objective function for model training is the sum of the above four loss functions:
[0083] L = L SSR + L MLM + L FTP + L MoCo .
Claims
1. A pre-training method for code intermediate representation based on semantic similarity, characterized in that First, the semantic similarity of basic blocks is defined based on input-output equivalence. By enumerating the equivalence relations of input variables, checking the equivalence relations of output variables, and calculating the proportion of the maximum equivalent output variables, the semantic similarity of basic blocks is obtained. Then, a learning pipeline is designed based on the semantic similarity, including extracting basic blocks from the intermediate representation of the code; translating the instruction semantics of basic blocks into the intermediate verification language of the program verifier; sampling candidate basic block pairs whose similarity needs to be calculated; using the program verifier to calculate the similarity of candidate pairs in parallel; modeling the learning of similarity as a regression task, and training the model using the mean squared error loss.
2. The method for pre-training the intermediate representation of code based on semantic similarity according to claim 1, wherein The pre-training method is based on the intermediate representation of the LLVM compilation system and the program verifier Boogie, and is used to guide the basic block representation learning of the pre-trained model based on the LLVM intermediate representation.
3. The method for pre-training the intermediate representation of code based on semantic similarity according to claim 1, characterized in that A program state σ consists of a pair (l, values), where l ∈ Loc is a certain program location and values: Var → Val is a function from variables to values that maps the variables of the program at location l to their concrete values; the set of all possible states of program P is denoted as Σ P ; an execution trace is a sequence of states <σ1, σ2, …, σ n >>, which describes the state transition path of a single execution of the program; the set of all possible traces of program P is denoted as Define the function and the function to return the first and last program states in the trace respectively; define the basic block to consist of all possible traces of its inputs; define the functions out(b) and var(b) to return the output variables and all variables of the basic block respectively; define the variable matching to be the matching of variables in basic block b1 and variables in basic block b2; the semantic similarity of basic blocks is defined as follows: where π1≡ γ π2 means that trajectory π1 and trajectory π2 are equivalent under matching γ, that is, the final program states of trajectory π1 and trajectory π2 assign the same values to the corresponding variables in matching γ; out(b1) γ , out(b2) γ respectively represent the sets of output variables included in basic blocks b1 and b2 in matching γ.
4. The method for pre-training the intermediate representation of code based on semantic similarity according to claim 3, wherein In the definition of the variable matching γ, taking out(b1) and out(b2) as two independent point sets, considering the legal equivalent possibilities between variables, connecting the nodes in the two point sets to form a bipartite graph; the matching γ is the matching on this bipartite graph, where each variable is connected to at most one variable in another basic block.
5. The method for pre-training code intermediate representation based on semantic similarity according to claim 1, wherein The basic block semantic similarity learning pipeline consists of the following five steps: 1) Extract basic blocks from the intermediate representation of the code; 2) Translate the instruction semantics of basic blocks into the intermediate verification language; 3) Sample basic block pairs whose similarity needs to be calculated; 4) Call the program verifier in parallel to calculate the similarity; 5) Similarity regression learning based on the mean squared error loss.
6. The method for pre-training the intermediate representation of code based on semantic similarity according to claim 5, characterized in that, In step 2), the basic block instructions of the LLVM intermediate representation are translated into the intermediate verification language Boogie IVL of Boogie; in terms of type translation, the integer type of the LLVM intermediate representation is translated into the integer type of Boogie IVL, the floating-point type of the LLVM intermediate representation is translated into the real type of Boogie IVL, and the user-defined structure type of the LLVM intermediate representation is translated into the special user-defined type "struct" of Boogie IVL; in terms of instruction translation, the integer arithmetic instructions of the LLVM intermediate representation are translated into the built-in integer operations of Boogie IVL, and the other instructions of the LLVM intermediate representation are translated into uninterpreted functions in Boogie IVL.
7. The method for pre-training the intermediate representation of code based on semantic similarity according to claim 5, wherein In step 4), first sample the composition of each training round before model training, that is, which basic blocks will be placed in the same batch for learning; then sample basic block pairs whose similarity needs to be calculated from the sampled batches.
8. The method for pre-training the intermediate representation of code based on semantic similarity according to claim 5, characterized in that, In the above (5), it is assumed that a basic block semantic similarity data set of size N is obtained through sampling calculation Among them, sim i respectively, the basic block pair and semantic similarity of the i-th piece of data; the similarity regression learning optimizes the following loss function: Among them, M represents the embedded neural network of the basic block, which converts the basic block input into a vector form; Equation calculates the cosine similarity of the vector representations of two basic blocks.
Citation Information
Cited By
Defense method and device for context injection attack of large language model
CN121598428A