A compiler optimization method and system based on node semantic enhancement
Through the compiler optimization method based on node semantic enhancement, the trained code optimization model and dense data flow graph are used to solve the problems of slow optimization speed and insufficient information utilization during the compiler optimization process, and more efficient code optimization is achieved.
Patent Information
- Application Number
- CN202510258424.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The existing compiler optimization technology is slow in finding optimization speed, and the use of intermediate representation code information is insufficient, resulting in limited optimization effect.
The compiler optimization method based on node semantic enhancement is adopted. By obtaining the intermediate representation code, a dense data flow diagram is generated, and a trained code optimization model is used, combining the compression and restore module and the structural feature extraction module, code optimization is performed to generate the optimized target code.
While removing code redundant information, the necessary structural information is retained, which improves the encoder optimization capability and significantly improves the optimization speed and effect.
Smart Images

Figure CN119739375B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of encoder optimization, and in particular to a compiler optimization method and system based on node semantic enhancement. Background Art
[0002] Compiler optimization has been a long-standing challenge in computer science. Its goal is to generate a new program that is semantically equivalent to the original program but with lower resource consumption and shorter running time by performing a series of transformations on the intermediate code. Code optimization usually includes multiple optimization passes, such as loop unrolling, common subexpression removal, and constant folding. According to their scope, they can be divided into instruction-level, basic block-level, function-level, module-level and other optimization levels. Different optimization methods and combinations may lead to different optimization results. The field of machine learning has made many attempts to solve this problem, focusing on optimization selection and phase sorting. Optimization selection focuses on determining which optimization content should be adopted, while phase sorting focuses on selecting the combination of optimizations. The LLVM (Low Level Virtual Machine) compiler is one of the most famous examples. LLVM integrates the front-ends of multiple high-level programming languages and provides a wide range of optimization passes for improving the performance of intermediate codes. By selecting different combinations of optimization passes, LLVM defines multiple optimization levels, such as -O0, -O1, -O2, -Oz, etc. However, LLVM IR usually needs to be manually updated to adapt to new programming language features and hardware architectures. Through automated optimization through deep learning models, the workload of manual maintenance can be reduced and overall efficiency can be improved.
[0003] Traditional compiler optimization techniques usually rely on static analysis and heuristic rules. Machine learning methods used representation-based methods to solve problems in the early days, such as using the static characteristics of the code as a representative of the program or dynamically predicting the optimization sequence based on the performance counters of the code. Heuristic methods are algorithms that produce reasonable results for problems based on experience, and are usually applied to optimization pass processes such as inlining and register allocation. However, there are some inherent problems with this type of method that have been difficult to solve. The first is that machine learning methods have a scarcity problem in the use of information about the IR (intermediate representation) code itself. Using a certain type of feature to represent the entire code will miss some information. Secondly, although the heuristic method is currently the best, the optimization speed is very slow due to the huge search space. It usually takes several minutes to find a better solution, which is unacceptable for the actual generation environment.
[0004] Therefore, finding a method that can fully utilize the intermediate representation code information and improve the compiler optimization speed is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the invention
[0005] The present invention provides a compiler optimization method and system based on node semantic enhancement, which are used to solve the defects of very slow optimization speed and insufficient utilization of code information in the compiler optimization process in the prior art, and achieve the removal of redundant information in the code while retaining necessary structural information, and improve the optimization capability of the encoder.
[0006] The present invention provides a compiler optimization method based on node semantic enhancement, comprising:
[0007] Obtaining an intermediate representation code before optimization, and processing the intermediate representation code before optimization to obtain a target code to be optimized;
[0008] Performing data annotation on the target code to be optimized to generate a dense data flow graph containing operation type information;
[0009] The target code to be optimized and the dense data flow graph are input into a trained code optimization model, the compression and restoration module and the structural feature extraction module of the trained code optimization model extract the static features of the code, and the node semantic enhancement module of the trained code optimization model performs code optimization based on the static features of the code to obtain the optimized target code.
[0010] According to a compiler optimization method based on node semantic enhancement provided by the present invention, the training process of the trained code optimization model is as follows:
[0011] Build a code optimization model based on the Transformer encoder-decoder architecture;
[0012] The code optimization ability is trained using the negative log-likelihood loss function, and the inter-node relationship and operation type prediction ability is trained using the multi-classification binary cross entropy loss function.
[0013] According to a compiler optimization method based on node semantic enhancement provided by the present invention, the loss function of the code optimization model is:
[0014] ;
[0015] ;
[0016] ;
[0017] in, represents the loss function of the code optimization model, represents the negative log-likelihood loss function, represents the hyperparameter, represents the multi-classification binary cross entropy loss function, L Indicates the length of the optimized target code, represents the marker at the dth position in the optimized target code, Indicates that given a word sequence c and a sequence Generated under the conditions The probability of represents the true label between node i and node j in the dense data flow graph under category k, k represents the number of the operation category, C is the total number of categories, is the probability that the code optimization model predicts that node i and node j belong to the kth class.
[0018] According to a compiler optimization method based on node semantic enhancement provided by the present invention, the code static features include code text sequence features and dense data flow node features, and the steps of the compression and restoration module are:
[0019] The target code to be optimized is subjected to standardized compression, and the compressed code is converted into a word sequence;
[0020] The word sequence is input into the decoder of the code optimization model to generate the encoding features of the code sequence.
[0021] According to a compiler optimization method based on node semantic enhancement provided by the present invention, the processing steps of the node semantic enhancement module are:
[0022] Perform feature splicing on the code text sequence feature and the dense data stream node feature to obtain a comprehensive feature;
[0023] Based on the dense data flow graph, the adjacency matrix of the directed connection relationship between nodes is constructed, and the shortest path topology structure of the adjacency matrix of the directed connection relationship between nodes is calculated by the Dijkstra algorithm.
[0024] The hop count information between nodes in the shortest path topology structure is mapped to different number ranges for positive data flow and negative data flow respectively through a nonlinear mapping function to generate a distance matrix, and a bidirectional relative data flow node position offset is constructed based on the distance matrix;
[0025] A masking matrix is constructed according to the distance matrix, and the bidirectional relative data flow node position bias and node masking mechanism are used to perform attention calculation and perception semantic information on the comprehensive features to generate an optimized target code.
[0026] According to a compiler optimization method based on node semantic enhancement provided by the present invention, the node masking mechanism is:
[0027] Mark the data flow node positions that have direct or indirect relationships in the distance matrix as 1;
[0028] When there is no direct or indirect dependency between data flow nodes, or there is no mapping relationship between nodes and code sequences, the corresponding attention score position is marked as negative infinity.
[0029] According to a compiler optimization method based on node semantic enhancement provided by the present invention, the perceived semantic information specifically includes:
[0030] The comprehensive features are input into the encoder of the code optimization model, and multi-head self-attention is calculated based on the node position bias of the bidirectional relative data flow and the node masking mechanism;
[0031] Normalize the layers of multi-head self-attention to obtain intermediate features; the calculation formula is:
[0032] ;
[0033] ;
[0034] in, represents the intermediate feature, is the layer normalization operation, represents the multi-head self-attention mechanism, Represents comprehensive characteristics, represents the code text sequence features, Represents the node characteristics of dense data flow;
[0035] The intermediate features are processed by a two-layer feedforward network in the encoder of the code optimization model and the layer is normalized to obtain the final feature representation; the calculation formula is:
[0036] ;
[0037] ;
[0038] in, Represents the two-layer feed-forward network in the encoder of the code optimization model, represents the activation function, x is the feature vector passed into the feedforward network, and is the weight matrix of the first layer feedforward network and the second layer feedforward network, and is the bias parameter of the first layer feedforward network and the second layer feedforward network, represents the final feature representation;
[0039] The final feature representation is input into the encoder of the code optimization model, and the optimized target code is generated through the hidden state of the last layer of the decoder.
[0040] The present invention also provides a compiler optimization system based on node semantic enhancement, which implements the compiler optimization method as described above, including:
[0041] A code acquisition module, used to acquire the intermediate representation code before optimization, and process the intermediate representation code before optimization to obtain the target code to be optimized;
[0042] A data annotation module, used to perform data annotation on the target code to be optimized and generate a dense data flow graph containing operation type information;
[0043] The code optimization module is used to input the target code to be optimized and the dense data flow graph into the trained code optimization model, the compression and restoration module and the structural feature extraction module of the trained code optimization model extract the static features of the code, and the node semantic enhancement module of the trained code optimization model performs code optimization based on the static features of the code to obtain the optimized target code.
[0044] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any one of the compiler optimization methods described above is implemented.
[0045] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the compiler optimization method described above is implemented.
[0046] The present invention provides a compiler optimization method and system based on node semantic enhancement, which realizes compiler optimization by combining a compression and restoration module and a structural feature extraction module with node semantic enhancement, removes redundant information in the code while retaining necessary structural information, improves the utilization efficiency of code information through a structure-aware self-attention mechanism, and improves the encoder optimization capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0048] Figure 1 It is a flowchart of the compiler optimization method provided by the present invention;
[0049] Figure 2 It is a flowchart of the compiler optimization method provided by the present invention;
[0050] Figure 3 is an example diagram of structural information of the compiler optimization method provided by the present invention;
[0051] Figure 4 It is an example diagram of a node semantic enhancement module of the compiler optimization method provided by the present invention;
[0052] Figure 5 It is a perceptual semantic information framework diagram of the compiler optimization method provided by the present invention;
[0053] Figure 6 is an application example diagram of the compiler optimization method provided by the present invention;
[0054] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0056] like Figure 1 and Figure 2 As shown, the present invention provides a compiler optimization method based on node semantic enhancement, comprising the following steps:
[0057] S1. Obtaining an intermediate representation code before optimization, and processing the intermediate representation code before optimization to obtain a target code to be optimized;
[0058] S2. Data annotation is performed on the target code to be optimized to generate a dense data flow graph containing operation type information;
[0059] S3. Input the target code to be optimized and the dense data flow graph into a trained code optimization model. The compression and restoration module and the structural feature extraction module of the trained code optimization model extract the static features of the code. The node semantic enhancement module of the trained code optimization model performs code optimization based on the static features of the code to obtain the optimized target code.
[0060] It can be understood that obtaining the intermediate representation code before optimization is to collect a large number of C language functions from the code hosting platform and the C language functions in the LLVM compiler test suite, and convert these functions into the intermediate representation (IR) code before optimization.
[0061] Among them, part of the data used in the present invention comes from the open source platform Angha, including some code snippets obtained from the GitHub platform. The obtained data is screened, and the code snippets that can be correctly executed by the compiler are retained and recorded as the target code to be optimized.
[0062] In one embodiment of the present invention, the intermediate representation code before optimization is processed, and a machine learning optimization algorithm can be used to explore an optimization path combination that is equal to or better than the compiler's Oz optimization level. After obtaining the optimization path, data with insignificant optimization effect is removed, and non-essential text segments that do not affect code execution are deleted from the original IR data and the optimized IR data. The data is validated using the LLVM compiler to obtain an optimized target code, wherein the optimized target code is equal to or better than the compiler's code in terms of code size optimization.
[0063] It can be understood that LLVM's Oz level is a special optimization level provided by the LLVM compiler. Among LLVM's multiple optimization levels (such as -O0, -O1, -O2, etc.), the Oz level is specifically optimized for generating the minimum code size. When the optimized code size is equal to or less than the code size obtained by using the Oz level optimization, the optimization is considered to be effective.
[0064] Specifically, the CompilerGym tool was used to run the heuristic method on multiple CPU machines, with the search space set to 128, covering 128 optimization options of the LLVM compiler, and the LLVM Oz-level optimization size as the reward function. CompilerGym is a platform developed by Meta for compiler researchers to conduct optimization experiments.
[0065] In one embodiment of the present invention, data labeling includes: using a lexical analyzer to parse the target code to be optimized, identifying and labeling the data dependency and control flow information in the target code to be optimized, and constructing a dense data flow graph (DDFG) containing complete semantic information.
[0066] Among them, the data flow relationship refers to the dependency relationship between variables in the code, such as the data transfer generated by assignment, calculation and other operations. The operation types usually include: arithmetic operations (addition, subtraction, multiplication and division), logical operations (AND, OR, NOT), comparison operations (greater than, less than or equal to), assignment operations, function calls, etc.
[0067] Specifically, a lexical analyzer is used to identify various types of operation instructions in the code, analyze the data flow relationship between instructions and mark the operation type;
[0068] For the jump instructions between different basic blocks, a jump relationship diagram of the program control flow is constructed; a basic block refers to a sequence of code that is executed sequentially in a program. Except for the last instruction, it does not contain any jump instructions and will not enter the code from the middle. Jump instructions include conditional jumps (if-else), loop jumps (for / while), function calls, etc.
[0069] The data flow relationship and jump relationship diagram are combined to generate a dense data flow diagram, which includes the data flow relationship, operation type information and jump relationship of the code.
[0070] In one embodiment of the present invention, the training process of the trained code optimization model is as follows:
[0071] Build a code optimization model based on the Transformer encoder-decoder architecture;
[0072] The negative log-likelihood loss function is used to train the code optimization capability, and the multi-classification binary cross entropy loss function is used to train the node relationship and operation type prediction capability. The loss function of the code optimization model is:
[0073] ;
[0074] ;
[0075] ;
[0076] in, represents the loss function of the code optimization model, represents the negative log-likelihood loss function, represents the hyperparameter, represents the multi-classification binary cross entropy loss function, L Indicates the length of the optimized target code, represents the marker at the dth position in the optimized target code, Indicates that given a word sequence c and a sequence Generated under the conditions The probability of represents the true label between node i and node j in the dense data flow graph under category k, k represents the number of the operation category, C is the total number of categories, is the probability that the code optimization model predicts that node i and node j belong to the kth class.
[0077] The present invention adopts a multi-task learning training method and uses the optimized target code to calculate the loss of the code optimization model. Through the combination of negative log-likelihood loss function and multi-classification binary cross entropy loss function, the code generation ability and node relationship prediction ability are optimized at the same time, which enhances the semantic perception ability of the code optimization model in the decoding stage and makes the generated optimized code more accurate.
[0078] In one embodiment of the present invention, the code static features include code text sequence features and dense data flow node features, and the steps of the compression and restoration module are:
[0079] The target code to be optimized is subjected to standardized compression, and the compressed code is converted into a word sequence;
[0080] The word sequence is input into the decoder of the code optimization model to generate the encoding features of the code sequence.
[0081] Specifically, the compression steps are as follows:
[0082] 1) Remove naming redundancy: Since node and variable names in the intermediate representation usually do not contain actual semantic information, compressing the text in the instruction sequence will not affect the semantic integrity of the code. The compression process of the CR (Compression Restoration) module mainly standardizes the names of variables and basic blocks in the intermediate representation code before optimization.
[0083] 2) Variable name reconstruction: For all variable names that appear in the intermediate representation code before optimization, replace them with regular expressions and use the same variable names. Replace with Represents numbers and increases gradually in the order of appearance.
[0084] 3) Basic block reconstruction: For basic blocks, use the form The standard format is renamed, where Indicates numbers, increasing in sequence.
[0085] 4) Other refactoring methods: In addition to variable names and basic blocks, the CR module also standardizes other code elements to eliminate naming redundancy and maintain the integrity of code semantics.
[0086] The compression and restoration module of the present invention performs standardized compression processing on the intermediate representation through regular expressions, effectively removing redundant information in the code while retaining necessary structural information. It not only reduces the volume of the input code and reduces the generation pressure of the code optimization model, but also improves the generalization ability of the code optimization model through a unified expression method.
[0087] Among them, the compression and restoration module is used to compress and restore the code without losing the functional and semantic integrity of the code, as shown in Table 1:
[0088] Table 1 Comparison table of sample codes for compression and restoration modules
[0089]
[0090] As shown in Table 1, the unprocessed intermediate code in code sample 1 includes the function definition @gen and four basic blocks (the basic block organization method is label %...). The compression and restoration module compresses the intermediate code, removes the variable name (such as %2), uses common variable names (such as I1, I2) to represent instructions, optimizes the format, retains the code logic and semantic integrity, and obtains a compressed sample. The compressed sample is then restored, the instruction position and execution order are adjusted, and the code logic is improved to obtain code sample 2.
[0091] like Figure 3 As shown, DFG in B1 and DFG in B3 represent two basic blocks in the dense data flow graph, Para represents a parameter node, call represents a call node, icmp.eq represents a conditional judgment node, br, br.T, br.F represent different branch nodes, and the arrows between the nodes represent the data flow relationship. 2 represents the temporary parameter variable %0 in basic block 2, @lua represents the function name lua, I1 2 Indicates the first instruction in basic block 2, I2 1 Indicates the second instruction in basic block 1, I4 1 Indicates the 4th instruction in basic block 1, @ngx_pool indicates the function name ngx_pool, I1 3 Indicates the first instruction in basic block 3, B5 3 Represents label B5 in basic block 3. First, the node adjacency matrix is constructed according to the data flow graph, and then the Dijkstra algorithm is used to calculate the shortest path between nodes. Then the shortest path is transformed by the nonlinear mapping function f, and positive mapping and negative mapping are performed in the upper triangular area and lower triangular area of the matrix respectively, and finally the distance matrix D and node position embedding information are obtained.
[0092] In one embodiment of the present invention, the steps of the structural feature extraction module are as follows:
[0093] The encoded features of the code sequence are processed through the word embedding layer to obtain the code text sequence features;
[0094] According to the dense data flow graph, an adjacency matrix of the directed connection relationship between nodes and an adjacency matrix representing the mapping relationship between nodes and code sequences are constructed, and input into the code optimization model. After processing through the node embedding test, the dense data flow node features are obtained.
[0095] Preferably, the encoding features of the code sequence are processed by word segmentation using the BPE (Byte Pair Encoding) algorithm to obtain a word sequence , the code text sequence features are obtained through the word embedding layer , where each represents a word token after being split by the BPE algorithm, and m is the maximum length of the sequence; two key matrices are constructed from the data flow information in the dense data flow graph: the adjacency matrix of the directed connection relationship between nodes and the adjacency matrix representing the mapping relationship between nodes and code sequences. The adjacency matrix of the directed connection relationship between nodes and the adjacency matrix representing the mapping relationship between nodes and code sequences are input into the code optimization model, and the dense data flow node features are obtained through a node embedding layer. , where the dense data flow node sequence ,in Represents a data flow node extracted from the code, n is the maximum number of nodes, and the dimension of the word embedding layer is consistent with the output dimension of the node embedding layer.
[0096] like Figure 4 As shown, in one embodiment of the present invention, the processing steps of the node semantic enhancement module are:
[0097] Perform feature splicing on the code text sequence feature and the dense data stream node feature to obtain a comprehensive feature;
[0098] Based on the dense data flow graph, an adjacency matrix of directed connection relationships between nodes is constructed, and the shortest path topology structure of the adjacency matrix of directed connection relationships between nodes is calculated by the Dijkstra algorithm;
[0099] The hop count information between nodes in the shortest path topology structure is mapped to different number ranges for positive data flow and negative data flow respectively through a nonlinear mapping function to generate a distance matrix, and a bidirectional relative data flow node position offset is constructed based on the distance matrix;
[0100] A masking matrix is constructed according to the distance matrix, and the bidirectional relative data flow node position bias and node masking mechanism are used to perform attention calculation and perception semantic information on the comprehensive features to generate an optimized target code.
[0101] Specifically, the calculation formula for the shortest path topology is:
[0102]
[0103] in, The adjacency matrix represents the directed connection relationship between data flow nodes. n is the maximum number of nodes. Represents the shortest path topology between nodes, represents the shortest path length from node i to node j, that is, the element in the i-th row and j-th column in the shortest path topology matrix. Represents a function that uses Dijkstra's algorithm to calculate the shortest path. Representation Node and nodes The connection relationship between them, i represents the index of the starting node, and j represents the index of the target node.
[0104] It can be understood that when performing data stream-aware self-attention calculations, a masking mechanism based on the distance matrix is introduced to set the attention scores of nodes that do not have direct or indirect dependencies to negative infinity, ensuring that the code optimization model only focuses on semantically related node connections; finally, these structured information are fused in the encoder through the structure-aware self-attention mechanism, and the hidden state of the last layer of the decoder is used to generate the optimized code sequence.
[0105] The node semantic enhancement module of the present invention calculates the shortest path topology of data flow nodes and introduces a bidirectional relative position bias, and combines the masking mechanism to guide the model to focus on semantically related node connections, thereby enhancing the code optimization model's ability to understand the code structure, thereby improving the optimization effect of the code optimization model.
[0106] Furthermore, due to The represented topological structure is directional. The upper triangular area represents positive data flow, which is represented by positive values, and the lower triangular area represents negative data flow, which is represented by negative values. The expression of the distance matrix M is as follows:
[0107] Positive area:
[0108]
[0109] Negative area:
[0110]
[0111] in, Represents the element value of the uth row and vth column in the mapped distance matrix, represents the shortest path length from node i to node j, Indicates the maximum number of hops.
[0112] It can be understood that the number 8 represents one quarter of the maximum number. The node distance within 8 hops in the upper triangle area is mapped to positions numbered 17 to 24, and the subsequent numbers increase logarithmically; the node distance within 8 hops in the lower triangle area is mapped to positions numbered 1 to 8, and the subsequent numbers increase logarithmically.
[0113] Furthermore, the node shadowing mask mechanism is:
[0114] Mark the data flow node positions with direct or indirect relationships in the distance matrix as 1;
[0115] When there is no direct or indirect dependency between data flow nodes, or there is no mapping relationship between nodes and code sequences, the corresponding attention score position is marked as negative infinity.
[0116] like Figure 6 As shown in the figure, by splicing the code text sequence features and the dense data flow node features, the input code is semantically perceived based on the adjacency matrix of the directed connection relationship between nodes and the adjacency matrix representing the mapping relationship between nodes and code sequences, and the optimized target code is obtained. The perceived semantic information specifically includes:
[0117] The comprehensive features are input into the encoder of the code optimization model, and multi-head self-attention is calculated based on the node position bias of the bidirectional relative data flow and the node masking mechanism;
[0118] Normalize the layers of multi-head self-attention to obtain intermediate features; the calculation formula is:
[0119] ;
[0120] ;
[0121] in, represents the intermediate feature, is the layer normalization operation, represents the multi-head self-attention mechanism, Represents comprehensive characteristics, represents the code text sequence features, Represents the node characteristics of dense data flow;
[0122] The intermediate features are processed by a two-layer feedforward network in the encoder of the code optimization model and the layer is normalized to obtain the final feature representation; the calculation formula is:
[0123] ;
[0124] ;
[0125] in, Represents the two-layer feed-forward network in the encoder of the code optimization model, represents the activation function, x is the feature vector passed into the feedforward network, and is the weight matrix of the first layer feedforward network and the second layer feedforward network, and is the bias parameter of the first layer feedforward network and the second layer feedforward network, represents the final feature representation;
[0126] The final feature representation is input into the encoder of the code optimization model, and the optimized target code is generated through the hidden state of the last layer of the decoder.
[0127] In one embodiment of the present invention, the masked attention mechanism can filter out irrelevant nodes to ensure that the code optimization model only focuses on semantically relevant node connections. If there is no mapping relationship between them (that is, the value of the corresponding position in the distance matrix is 0), the attention score is set to negative infinity. The calculation formula of the attention score is:
[0128]
[0129] in, Represents a data flow node and The attention score between and Represents two data flow nodes to be calculated for attention, and Respectively represent nodes and The feature vector representation of and are the query transformation matrix and the key transformation matrix respectively, T represents the matrix transpose operation, and Respectively represent nodes and The mapping relationship matrix with the word sequence C, Representation Node arrive The weight bias corresponding to the number of hops between Representation Node and The shortest path length between .
[0130] Among them, the construction process of the mapping relationship matrix includes: first, extracting the position index information of each data flow node in the target code to be optimized, then using the word segmenter to tokenize the target code to be optimized to obtain a complete token sequence, and at the same time obtaining the local token representation corresponding to each node based on the position index information, and finally constructing a binary mapping matrix through the position correspondence between the node token and the complete code sequence token, which is used to characterize the correspondence between the node and the word sequence.
[0131] The present invention realizes compiler optimization by combining the compression and restoration module and the structural feature extraction module with node semantic enhancement, thereby removing redundant information in the code while retaining the necessary structural information. The utilization efficiency of code information is improved through the structure-aware self-attention mechanism, thereby improving the encoder optimization capability.
[0132] In one embodiment of the present invention, the optimization capabilities of other traditional LLVM compiler optimization methods, heuristic optimization methods, Transformer model-based optimization methods, and Transformer model-based optimization methods that introduce the NSE module are compared, as shown in Table 2:
[0133] Table 2 Comparison of code size optimization capabilities
[0134]
[0135] Table 2 shows the comparison results of different optimization methods in terms of code size optimization capabilities. The number of instructions under each optimization method is counted for the semantically complete code generated by the model. It can be seen that the performance benchmark of the traditional LLVM compiler optimization method is 1.0. The optimization method based on the Transformer model has achieved an optimization capability of 1.080 without using the NSE module. After the introduction of the NSE module, it is further improved to 1.106, which enhances the optimization effect of the model. Compared with the heuristic method of machine learning, this method achieves 74.6% of its effect in optimizing code size. Compared with the LLVM compiler optimization method, although the heuristic optimization method is often more effective, it is less efficient and has a slower processing speed. The optimization method combining deep learning and the NSE module is faster and has more potential.
[0136] like Figure 6 As shown, specifically, a specific embodiment is used to illustrate that the model is integrated into the LLVM compiler. The specific implementation process is as follows:
[0137] 1) Compiler front-end integration: The compiler front-end generates unoptimized IR code at the input end, which is used as the input of the model after being standardized by the CR module;
[0138] 2) Model optimization processing: The trained code optimization model receives the standardized IR code and generates optimized code. During this process, the optimized code output by the trained code optimization model will be restored to the LLVM IR standard format by the CR module;
[0139] 3) Code verification mechanism: This embodiment uses the Alive2 tool to perform strict correctness verification. The Alive2 tool verifies the semantic integrity of the optimized code and checks its semantic equivalence with the original code;
[0140] 4) Error handling mechanism: trigger fallback processing when Alive2 verification reports semantic inequality or detects potential undefined behavior;
[0141] 5) Fallback process: Adopt the atomic processing principle, including saving a complete backup of the original IR code. When Alive2 verification fails, immediately mark the optimization invalid and restore the original IR representation of the code segment, and call the original LLVM optimization channel for processing;
[0142] 6) Backend code generation: After the optimized code verified by Alive2 is restored through the CR module, the optimized code is converted and the backend code, i.e. the final target code, is generated.
[0143] The present invention provides a compiler optimization system based on node semantic enhancement, which implements the compiler optimization method as described above, including:
[0144] A code acquisition module, used to acquire the intermediate representation code before optimization, and process the intermediate representation code before optimization to obtain the target code to be optimized;
[0145] A data annotation module, used to perform data annotation on the target code to be optimized and generate a dense data flow graph containing operation type information;
[0146] The code optimization module is used to input the target code to be optimized and the dense data flow graph into the trained code optimization model, the compression and restoration module and the structural feature extraction module of the trained code optimization model extract the static features of the code, and the node semantic enhancement module of the trained code optimization model performs code optimization based on the static features of the code to obtain the optimized target code.
[0147] The compiler optimization device provided by the present invention is described below. The compiler optimization device described below and the compiler optimization method described above can be referenced to each other.
[0148] Figure 7 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 7As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730 and a communication bus 740, wherein the processor 710, the communication interface 720 and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logic instructions in the memory 730 to execute the compiler optimization method, which includes: obtaining the intermediate representation code before optimization, and processing the intermediate representation code before optimization to obtain the target code to be optimized; data annotation of the target code to be optimized to generate a dense data flow graph containing operation type information; inputting the target code to be optimized and the dense data flow graph into the trained code optimization model, the compression and restoration module and the structural feature extraction module of the trained code optimization model extract the static features of the code, and the node semantic enhancement module of the trained code optimization model performs code optimization based on the static features of the code to obtain the optimized target code.
[0149] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0150] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the compiler optimization method provided by the above methods, which includes: obtaining an intermediate representation code before optimization, and processing the intermediate representation code before optimization to obtain a target code to be optimized; performing data annotation on the target code to be optimized to generate a dense data flow graph containing operation type information; inputting the target code to be optimized and the dense data flow graph into a trained code optimization model, the compression and restoration module and the structural feature extraction module of the trained code optimization model extract code static features, and the node semantic enhancement module of the trained code optimization model performs code optimization based on the code static features to obtain the optimized target code.
[0151] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the compiler optimization method provided by the above-mentioned methods, the method comprising: obtaining an intermediate representation code before optimization, and processing the intermediate representation code before optimization to obtain a target code to be optimized; performing data annotation on the target code to be optimized to generate a dense data flow graph containing operation type information; inputting the target code to be optimized and the dense data flow graph into a trained code optimization model, the compression and restoration module and the structural feature extraction module of the trained code optimization model extracting code static features, and the node semantic enhancement module of the trained code optimization model performing code optimization based on the code static features to obtain the optimized target code.
[0152] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0153] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A compiler optimization method based on node semantic enhancement, characterized in that: include: Obtaining an intermediate representation code before optimization, and processing the intermediate representation code before optimization to obtain a target code to be optimized; Performing data annotation on the target code to be optimized to generate a dense data flow graph containing operation type information; wherein the dense data flow graph contains the data flow relationship, operation type information and jump relationship of the code; The target code to be optimized and the dense data flow graph are input into a trained code optimization model, the compression and restoration module and the structural feature extraction module of the trained code optimization model extract code static features, and the node semantic enhancement module of the trained code optimization model performs code optimization based on the code static features to obtain an optimized target code; The code static features include code text sequence features and dense data flow node features, and the steps of the compression and restoration module are: standardizing and compressing the target code to be optimized, converting the compressed code into a word sequence; inputting the word sequence into a decoder of the code optimization model to generate coding features of the code sequence; Among them, the processing steps of the node semantic enhancement module are: feature splicing of the code text sequence features and the dense data flow node features to obtain comprehensive features; constructing an adjacency matrix of the directed connection relationship between nodes based on the dense data flow graph, and calculating the shortest path topology structure of the adjacency matrix of the directed connection relationship between nodes through the Dijkstra algorithm; using the hop count information between nodes in the shortest path topology structure to map the positive data flow and the negative data flow to different numbering ranges through a nonlinear mapping function to generate a distance matrix, and constructing a bidirectional relative data flow node position bias based on the distance matrix; constructing a shielding mask matrix according to the distance matrix, using the bidirectional relative data flow node position bias and the node shielding mask mechanism to perform attention calculation and perception semantic information on the comprehensive features, and generate an optimized target code.
2. A compiler optimization method based on node semantic enhancement according to claim 1, characterized in that: The training process of the trained code optimization model is as follows: Build a code optimization model based on the Transformer encoder-decoder architecture; The code optimization ability is trained using the negative log-likelihood loss function, and the inter-node relationship and operation type prediction ability is trained using the multi-classification binary cross entropy loss function.
3. A compiler optimization method based on node semantic enhancement according to claim 2, characterized in that: The loss function of the code optimization model is: ; ; ; in, represents the loss function of the code optimization model, represents the negative log-likelihood loss function, represents the hyperparameter, represents the multi-classification binary cross entropy loss function, L represents the length of the optimized target code, represents the marker at the dth position in the optimized target code, Indicates that given a word sequence c and a sequence Generated under the conditions The probability of represents the true label between node i and node j in the dense data flow graph under category k, k represents the number of the operation category, C is the total number of categories, is the probability that the code optimization model predicts that node i and node j belong to the kth class.
4. The compiler optimization method based on node semantic enhancement according to claim 1, characterized in that: The node masking mechanism is: Mark the data flow node positions that have direct or indirect relationships in the distance matrix as 1; When there is no direct or indirect dependency between data flow nodes, or there is no mapping relationship between nodes and code sequences, the corresponding attention score position is marked as negative infinity.
5. The compiler optimization method based on node semantic enhancement according to claim 4, characterized in that: The perceptual semantic information specifically includes: Inputting the comprehensive features into an encoder of a code optimization model, and calculating multi-head self-attention based on a bidirectional relative data flow node position bias and a node masking mask mechanism; Normalize the layers of multi-head self-attention to obtain intermediate features; the calculation formula is: ; ; in, represents the intermediate feature, is the layer normalization operation, represents the multi-head self-attention mechanism, Represents comprehensive characteristics, represents the code text sequence features, Represents the node characteristics of dense data flow; The intermediate features are processed by a two-layer feedforward network in the encoder of the code optimization model and the layer is normalized to obtain the final feature representation; the calculation formula is: ; ; in, Represents the two-layer feed-forward network in the encoder of the code optimization model, represents the activation function, x is the feature vector passed into the feedforward network, and is the weight matrix of the first layer feedforward network and the second layer feedforward network, and is the bias parameter of the first layer feedforward network and the second layer feedforward network, represents the final feature representation; The final feature representation is input into the encoder of the code optimization model, and the optimized target code is generated through the hidden state of the last layer of the decoder.
6. A compiler optimization system based on node semantic enhancement, wherein the compiler optimization system based on node semantic enhancement is used to implement the compiler optimization method based on node semantic enhancement according to claim 1, characterized in that: Implementing the compiler optimization method according to any one of claims 1 to 5, comprising: A code acquisition module, used to acquire the intermediate representation code before optimization, and process the intermediate representation code before optimization to obtain the target code to be optimized; A data annotation module, used to perform data annotation on the target code to be optimized and generate a dense data flow graph containing operation type information; The code optimization module is used to input the target code to be optimized and the dense data flow graph into the trained code optimization model, the compression and restoration module and the structural feature extraction module of the trained code optimization model extract the static features of the code, and the node semantic enhancement module of the trained code optimization model performs code optimization based on the static features of the code to obtain the optimized target code.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the compiler optimization method according to any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the compiler optimization method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Virtual compiler data enhancement-based assembly code search and performance optimization method
CN118444891A
Optimised code generation
EP1387265A1