Assembly instruction rearrangement optimization system and method
By combining graph neural network and reinforcement learning framework assembly instruction rearrangement optimization system, the limitations of traditional compiler optimization technology in complex processor architecture are solved, efficient instruction orchestration and resource utilization are achieved, program execution efficiency is improved, highly adaptable, and development costs are reduced.
Patent Information
- Application Number
- CN202510409647.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional compiler optimization technology is difficult to find global optimal solutions on complex processor architectures and large-scale instruction sets, and the optimization effect is limited by the quality of manual rule design and cannot be improved quickly, resulting in inefficient program execution.
Combining the graph neural network and reinforcement learning framework, by building an assembly instruction rearrangement optimization system, using graph neural network to deeply model the instruction dependency relationship and feature matrix, and optimizing the deep reinforcement learning DQN method to generate an efficient instruction orchestration plan.
It has achieved high utilization rate of processor resources and significantly improved program operation efficiency, reduced dependence on manual rule design, strong adaptability, suitable for different hardware platforms and instruction sets, and reduced development costs.
Smart Images

Figure CN120295634A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of optimizing the arrangement of processor assembly instructions in computer science, and particularly relates to an assembly instruction rearrangement optimization system and method. Background Art
[0002] The function of a compiler is to convert high-level language program code into target platform code or low-level machine code that can be directly executed. Its main functions include lexical analysis, syntax analysis, intermediate code generation, optimization, and target code generation, etc. Among them, the optimization stage is one of the core functions and difficulties of the compiler, including a variety of optimization methods and means, aiming to improve the execution efficiency of the target program. Traditional compiler optimization technologies mainly rely on manually written rules and heuristic algorithms, such as loop unrolling, instruction scheduling, register allocation, etc. These methods are mostly based on the summary of experience of fixed patterns and have certain limitations. When facing complex processor architectures and large-scale instruction sets, traditional heuristic methods are difficult to find the global optimal solution, and the optimization effect is limited by the quality of manual rule design, and it is difficult to rapidly improve with the development of technology.
[0003] With the continuous development of hardware and software technologies, the demand for computing performance is also getting higher and higher. Automated and intelligent optimization methods can effectively cope with the challenges brought by modern computing, can discover optimization opportunities that are difficult to find manually, reduce the optimization burden on programmers, and shorten the development cycle, which is an important trend for future development.
[0004] Machine learning is one of the core technologies of artificial intelligence. By constructing mathematical models and algorithms, computer systems can automatically learn and improve from data. Among them, reinforcement learning is an important branch of machine learning. The uniqueness of reinforcement learning lies in that it emphasizes trial-and-error learning. In the process of the agent interacting with the environment, it continuously tries different actions and learns the optimal strategy based on the rewards or punishments feedback by the environment. Reinforcement learning has been successfully applied in many fields such as game AI, robot control, and intelligent scheduling.
[0005] In the traditional Q-learning algorithm, Q-values are usually stored in a table. However, when the state space and action space are very large, this table will become extremely huge and even impossible to store. This "curse of dimensionality" severely limits the application scope of Q-learning. DQN (Deep Q-Network) is an improvement of Q-learning, which uses a deep neural network to approximate the Q-value function, thus overcoming the shortcomings of Q-learning in dealing with large state space and action space. Compared with Q-learning, DQN can effectively handle complex high-dimensional inputs such as images, graph structures, and voices, and has strong generalization ability, making it easier to be migrated and applied.
[0006] A graph neural network (GNN) is a neural network used to process graph data. Graph data is a type of non-Euclidean data composed of nodes and edges, which is used to represent entities and their relationships. By extending the idea of neural networks to graph structures, the core idea of GNN is to exchange and aggregate information between graph nodes through a "message passing mechanism", so as to learn the features of nodes, edges, and the whole graph in the graph, and then complete downstream tasks such as classification, clustering, and link prediction. GNN has a wide range of applications in many fields, such as social network analysis, computer vision, recommendation systems, etc. The input of the GNN network is the structure of the graph (nodes and edges) and its features. The output of GNN can be in various forms: node-level output, edge-level output, and global-level output. These outputs capture the local or global structure information in the graph, including the dependencies between nodes and nodes, thus helping us better understand the inherent meaning of graph data and enabling machine learning algorithms to process and apply this graph data. Summary of the Invention
[0007] Aiming at the problems of the prior art, the present invention provides an optimized system and method for rearrangement of assembly instructions. The method aims to automatically generate efficient instruction arrangements, improve the program execution efficiency, and at the same time reduce the dependence on manual rule design through the autonomous learning ability of reinforcement learning. Combining a graph neural network with a reinforcement learning framework, the method has strong versatility and adaptability, and only needs a small amount of adjustment to adapt to different hardware platforms and instruction sets. In addition, the method can ensure the correctness and stability of the optimization process in the application of Matrix FT-Matrix accelerator assembly instructions.
[0008] In order to solve the problems of the prior art, the present invention adopts the following technical solutions:
[0009] An assembly instruction rearrangement optimization system, the system includes an assembly instruction preprocessing module, a first assembly instruction analysis module, a second assembly instruction analysis module, an assembly instruction compression unit, and an assembly instruction optimization unit; the instruction compression unit includes a first instruction status compression module and a second instruction status compression module; where:
[0010] The instruction preprocessing module divides the assembly code into basic blocks according to lexical, syntactic, and semantic rules;
[0011] The first instruction analysis module constructs an instruction adjacency matrix based on the mutual logical relationship of the code sequences in the basic blocks;
[0012] The second instruction analysis module extracts feature vectors in the basic blocks to construct an instruction feature matrix;
[0013] The assembly instruction compression unit is constructed by training a graph neural network according to the instruction adjacency matrix and the instruction feature matrix according to the following formula;
[0014]
[0015] Where: represents the feature representation of node v at the k-th layer in the model; N(v) is the set of neighbor nodes of node v
[0016] union;
[0017] The first instruction status compression module calculates and generates a global action status feature vector according to the adjacency matrix and the feature matrix according to a dynamic filtering function;
[0018] The second instruction status compression module generates a global spatial status feature vector by fusing a single graph convolutional layer with a fully connected layer;
[0019] The assembly instruction optimization unit is constructed by initializing and training a reinforcement learning network according to the instruction adjacency matrix and the instruction feature matrix according to the following function;
[0020]
[0021] Where: represents the loss function; represents the expectation; r represents the immediate reward; γ represents the discount factor; Q target represents the target network; Q represents the q-value network; s represents the current state; a represents the action taken in the current state; s′ represents the next state transferred to after taking action a in the current state s; a′ represents the possible action in the next state s′;
[0022] The instruction optimization module outputs the optimal action vector corresponding to the global state feature vector from the global action feature vector through a reward mechanism.
[0023] Furthermore, the second instruction state compression module generates a global spatial state feature vector by performing a fully connected layer fusion process on the single-graph convolutional layer, which includes:
[0024] Normalize the instruction adjacency matrix of the single-graph convolutional layer through the following formula:
[0025]
[0026] Where: represents the adjacency matrix with self-loops added. D represents the degree matrix, which is a diagonal matrix, and its diagonal element D(i,i) represents the degree of the node; is the square root of the inverse matrix of the degree matrix; and if and only if there exists at least one node such that 1 and Given the adjacency matrix Where: and if and only if there exists such that and
[0027] Calculate the multi-layer fully connected layer for different single-graph convolutional layers through the following formula:
[0028]
[0029] Where: H l is the result calculated at the l-th layer, and the input H of the 0-th layer 0 is the instruction feature matrix; Θ l is the updatable and weight-shared weight matrix at the l-th layer; σ represents the activation function;
[0030] Obtain the global state feature vector by fusing the multi-layer fully connected layer through the following formula
[0031]
[0032] Where: R(H) represents the function for compressing the intermediate state matrix H; N represents the number of rows of the intermediate state matrix;
[0033] represents the feature of the i-th row of the intermediate state matrix H; σ represents the activation function; represents the obtained global state feature vector.
[0034] Furthermore, the instruction optimization module outputs the optimal action vector corresponding to the global spatial state feature vector from the global action feature vector through a reward mechanism, which includes:
[0035] Construct an experience replay pool for the interaction between the global spatial state feature vector and the global action state, including: the current state s, the action a taken, the reward r obtained, and the quadruple {sars'} composed of the next state s'.
[0036] Update the reward after executing the action through the following formula:
[0037] r = ω1ΔT + ω2U + ω3V
[0038] Where: ΔT represents the reduction in execution time, U represents the resource utilization rate, V is whether to select the node on the longest path, and ω1, ω2, ω3 are weight parameters;
[0039] Select the corresponding global spatial state feature vector output from the global action feature vector according to the updated reward value.
[0040] The present invention can also adopt the following technical solution, including the following steps:
[0041] Assembly instruction preprocessing stage:
[0042] Segment the assembly code into basic blocks according to lexical, syntactic, and semantic rules;
[0043] Construct an instruction adjacency matrix according to the mutual logical relationship of the code sequences in the basic blocks;
[0044] Extract the feature vectors in the basic blocks to construct an instruction feature matrix;
[0045] Train the graph neural network through the instruction adjacency matrix and the instruction feature matrix according to the following formula to construct an assembly instruction compression unit;
[0046]
[0047] Where: represents the feature representation of node v in the k-th layer of the model; N(v) is the set of neighbor nodes of node v;
[0048] Initialize and train the reinforcement learning network through the instruction adjacency matrix and the instruction feature matrix according to the following function to construct an assembly instruction optimization unit;
[0049]
[0050] Where: represents the loss function; represents the expectation; r represents the immediate reward; γ represents the discount factor; Q targetLet \(N\) denote the target network; \(Q\) denote the \(Q\)-value network; \(s\) denote the current state; \(a\) denote the action taken in the current state; \(s'\) denote the next state transferred to after taking action \(a\) in the current state \(s\); \(a'\) denote the possible action in the next state \(s'\).
[0051] Assembly instruction optimization stage:
[0052] Calculate and generate the global action state feature vector according to the dynamic filtering function for the adjacency matrix and the feature matrix;
[0053] Generate the global spatial state feature vector by fusing the single graph convolutional layer with the fully connected layer;
[0054] Initialize and train the reinforcement learning network through the instruction adjacency matrix and the instruction feature matrix to construct the assembly instruction optimization unit;
[0055] Select the optimal action vector corresponding to the global state feature vector from the global action feature vector through the reward mechanism and output it.
[0056] Beneficial effects
[0057] By combining reinforcement learning and graph neural network technologies, the present invention proposes an innovative method for optimizing assembly instructions. The graph neural network is used to deeply model the instruction dependence relationship and the feature matrix, greatly compressing the state and action spaces, and enhancing the feature expression ability through self-supervised learning, realizing the efficient representation of complex instruction sets. On this basis, using the DQN method in deep reinforcement learning, by constructing an experience replay pool, introducing a target network and a \(Q\)-value network, the training efficiency and optimization effect are further improved.
[0058] Compared with the traditional compiler optimization method, the present invention can break out of the constraints of the traditional "fetch - compute - store" model in specific scenarios, generate an efficient assembly instruction program through more compact instruction arrangement, thereby achieving higher utilization of processor resources and a significant improvement in program running efficiency. In contrast, this method has a higher degree of automation, can reduce the dependence on manual tuning by domain experts, and significantly reduce the development cost. In terms of adaptability, the present invention has good scalability and decoupling characteristics. Its core architecture design can be flexibly adjusted, and only some modules need to be modified according to the characteristics of different platforms to complete the adaptation, taking into account both performance and scalability. This technology is widely applicable to fields such as high-performance computing, embedded systems, and heterogeneous computing platforms, providing a new idea and strong support for the efficient implementation of complex computing tasks. Description of the drawings
[0059] Figure 1 It is a schematic structural diagram of an assembly instruction rearrangement optimization system and method of the present invention.
[0060] Figure 2 Schematic diagram of the assembly instruction compression process in an assembly instruction rearrangement optimization system of the present invention. Specific embodiments
[0061] The following will Figure 1 make the following description of the present invention with reference to the attached
[0062] · As Figure 1 shown, the present invention provides an assembly instruction rearrangement optimization system based on graph neural network state compression and reinforcement learning. The system includes an assembly instruction preprocessing module 101, a first assembly instruction analysis module 102, a second assembly instruction analysis module 103, an assembly instruction compression unit 200, and an assembly instruction optimization unit 300; the instruction compression unit 200 includes a first instruction state compression module 201 and a second instruction state compression module 202; the assembly code is parsed into graph structure data (feature matrix, adjacency matrix). The state compression stage (GNN network) of the present invention uses GNN to transform the graph structure data into a global vector representation, significantly reducing the state space. The optimization stage in the present invention is reinforcement learning (DQN network)
[0063] uses DQN and an experience replay pool for optimization, combines the state representation compressed by GNN, outputs an optimized instruction arrangement scheme, and realizes efficient instruction optimization.
[0064] The instruction preprocessing module 101 divides the assembly code into basic blocks according to lexical, syntactic, and semantic rules; the first instruction analysis module 102 constructs an instruction adjacency matrix according to the mutual logical relationship of the code sequences in the basic blocks; the second instruction analysis module 103 extracts feature vectors in the basic blocks to construct an instruction feature matrix; specific content:
[0065] Assembly code parsing
[0066] In an actual scenario, the initially available data is a series of operators written in assembly instructions. A compiler needs to be written to extract and understand these assembly instruction codes for use in subsequent stages. The compiler includes three stages: lexical analysis, syntactic analysis, and semantic analysis.
[0067] · Lexical analysis: The assembly code is decomposed into individual lexical units (called tokens), such as mnemonic tokens, operand tokens, register tokens, operation instruction tokens, immediate number tokens, label tokens, etc. The purpose is to identify keywords, identifiers, constants, etc. in the assembly code.
[0068] · Syntactic analysis: According to the syntactic rules of the assembly code, the tokens identified in the lexical units are combined into legal syntactic structures, such as expressions, operation instructions, etc.
[0069] ·Semantic analysis: It includes checking the syntax structure, etc., and constructing the structure of the entire program. Since all operators have been verified for correctness, there is no need for an error handling module.
[0070] Partitioning of basic blocks
[0071] First, all assembly code needs to be split into different basic blocks. A basic block refers to a sequence of code with a tight logical connection. This block of code has exactly one entry point (i.e., only one possible execution starting point) and one exit point (i.e., only one possible execution ending point). All instructions in a basic block will be executed in sequence without jumping or branching midway, so it is the smallest execution unit in the program control flow. Among them, there are three types of entry points for basic blocks: the first instruction in the code segment, the target statement of a conditional jump or unconditional jump, and the next statement of a conditional jump statement.
[0072] When the program is partitioned into basic blocks, if the basic blocks are regarded as basic unit nodes, the relationship of being predecessors and successors to each other in the program execution flow between basic blocks can be regarded as a directed edge existing between two basic blocks. Then the entire program can be transformed into a directed graph, called the control flow graph (CFG). Between basic blocks, common optimization methods for the arrangement of assembly instructions include dead code elimination, code movement, strength reduction, etc. Among them, the automated optimization method of assembly code combined with a reinforcement learning model will be used within basic blocks.
[0073] The specific method for partitioning basic blocks is as follows: Start scanning from the beginning of the code. When encountering the target instruction of a jump or unconditional jump, it is partitioned into a basic block. For conditional jumps, it is partitioned according to the jump condition. For the loop body, the loop body is regarded as a basic block.
[0074] Construction of the adjacency matrix
[0075] In a program, there are complex dependencies between instructions. This is also the case in the basic blocks of the assembled code that have been partitioned. These dependencies can be divided into the following three types.
[0076] ·RAW (Read After Write): An instruction reads a value written by another instruction. For example, instruction 1 writes to register R1, and instruction 2 reads the value of register R1.
[0077] ·WAR (Write After Read): An instruction writes a value, and another instruction has read this value before.
[0078] ·WAW (Write After Write): Two instructions both write to the same location.
[0079] To facilitate the modeling and analysis of dependencies within a basic block, we can represent the instructions in a basic block as a directed graph. First, assign a unique number to each instruction within the basic block to identify the nodes in the graph. The nodes in the graph represent instructions, and the directed edges represent the dependencies between instructions. For example, if the output of instruction A is the input of instruction B, then there will be a directed edge in the graph from node A to node B. Other types of dependencies are modeled in a similar manner.
[0080] The adjacency matrix is a commonly used method for representing graphs. For a graph structure with n nodes, we can construct an n×n matrix A. If there is an edge from node i to node j, then A(i,j) will be set to 1; otherwise, it will be 0. In this way, the dependencies between the instructions in the entire basic block can be represented by an adjacency matrix.
[0081] The specific steps for constructing the adjacency matrix are as follows: Traverse all pairs of instructions within the basic block and check whether there is a data dependency between them. If a dependency is found, mark the corresponding element position in the adjacency matrix as 1. This approach can effectively express and analyze the complex instruction dependencies within a basic block.
[0082] Construction of the Feature Matrix
[0083] The features of an instruction include the following aspects:
[0084] · Instruction type: Arithmetic instructions, logical instructions, memory access instructions, etc.
[0085] · Operand type: Immediate numbers, registers, memory addresses, etc.
[0086] · Number of operands: The quantity of operands in the instruction.
[0087] · Instruction cycle: The length of time required for the instruction to complete the calculation.
[0088] · Resource occupancy: The hardware resources required for the instruction to execute.
[0089] · In-degree of the instruction: The number of instructions that the instruction depends on.
[0090] · Out-degree of the instruction: The number of other instructions that depend on the instruction.
[0091] · Longest path: Whether the instruction is on the longest path of the program.
[0092] By stacking the feature vectors of each instruction row by row, a feature matrix can be obtained. Each row represents an instruction, and each column represents a feature. Moreover, the row number of each instruction in the feature matrix should correspond to its number in the adjacency matrix.
[0093] In the data preprocessing stage, two core data structures, namely the feature matrix and the adjacency matrix, are generated through the parsing of assembly code and the partitioning of basic blocks. In the assembly code, each instruction in a basic block is assigned a unique number, and these numbers will correspond one by one in the adjacency matrix and the feature matrix. The adjacency matrix is a two-dimensional matrix with dimensions of n×n, which is used to represent the data dependency relationships between instructions. If there is a data dependency relationship between the i-th instruction and the j-th instruction, the element A(i,j) at the corresponding position in the adjacency matrix will be set to 1. The feature matrix converts various features of each instruction (such as instruction type, operand type, instruction cycle, etc.) into a vector representation. Each row represents an instruction, and each column represents a certain feature of the instruction, and finally a feature matrix with multiple dimensions is formed. And the number of rows of the feature matrix is also n.
[0094] By converting the basic blocks of the assembly code into an adjacency matrix and a feature matrix, the assembly code problem can be transformed into a graph theory and machine learning problem, and then deeper analysis and optimization work can be carried out through the GNN network and the DQN network.
[0095] This constructs the assembly instruction compression unit 200 for the training of the graph neural network according to the following formula by using the instruction adjacency matrix and the instruction feature matrix; the first instruction state compression module 201 calculates and generates a global action state feature vector according to the dynamic filtering function for the adjacency matrix and the feature matrix; the second instruction state compression module 202 generates a global spatial state feature vector through the full connection layer fusion processing of the single graph convolutional layer;
[0096] In the optimization of the state space, according to the structural characteristics of the operator basic block, a GNN-based encoder is designed to map the adjacency matrix and the feature matrix into a low-dimensional global feature vector, thereby effectively compressing the state space and reducing the computational complexity. In the optimization of the action space, according to the hardware characteristics of the target platform, the actions that may generate low rewards are eliminated through the dynamic filtering function, reducing unnecessary computations. In the initial stage of the operator, for the load operation that can be directly executed without depending on other instructions, it is preferentially included in the action set to reduce the scale of the action space and improve the optimization efficiency.
[0097] In the figure, (X,A) represents the input feature matrix and adjacency matrix, and the subsequent connected ε represents the network structure for feature extraction and message passing of the input. Next, the most critical part will be specifically introduced.
[0098]
[0099] In the first line of formula (1), Denote the adjacency matrix with self-loops added. D represents the degree matrix, which is a diagonal matrix, and the diagonal element D(i,i) represents the degree of the node. The specific calculation formula is as follows:
[0100] is the square root of the inverse matrix of the degree matrix. Then the function of the first formula is to normalize the adjacency matrix, which can weaken the influence of high-degree nodes and enable the graph convolutional neural network to better learn the local features and global structure of the nodes. The matrix obtained through the above transformation is usually called the symmetric normalized Laplacian matrix.
[0101] Similarly, the calculation processes of the second and third formulas are the same. The following gives the definition of the matrix . Given the adjacency matrix where and if and only if there exists at least one node (select one of the nodes and call it j) such that and The definition of is similar but in the opposite direction. Given the adjacency matrix where and if and only if there exists at least one node (select one of the nodes and call it j) such that and
[0102]
[0103] Formula (4) shows how the data is specifically calculated in different graph convolutional layers. H l is the result calculated in the l-th layer, and H 0 = X. Θ l is the updatable and weight-shared weight matrix in the l-th layer. σ represents the activation function. In the basic graph convolutional layer, the calculation of the second-order in-degree and out-degree matrix is added, considering the directionality of the edges in the graph structure, and more deeper information in the graph structure can be extracted.
[0104] Connect multiple fully connected layers after the graph convolutional layer, record the vector at this time as H, and then through the function
[0105] obtain the final global vector output This vector can be used for the subsequent deep Q-network (DQN) in reinforcement learning and represents the state.
[0106] Initialize and train the reinforcement learning network to construct an assembly instruction optimization unit by using the instruction adjacency matrix and the instruction feature matrix according to the following function; the instruction optimization module selects the optimal action vector corresponding to the global state feature vector from the global action feature vector through a reward mechanism and outputs it. The update formula for the Q value is as follows:
[0107]
[0108] Where: Q(s t ,a t ) is the Q value of taking action a t in state s t ; α is the learning rate; r t+1 is the immediate reward obtained after taking action a t in state s t ; γ is the discount factor; is the maximum Q value of all possible actions in the next state s t+1 ;
[0109] The loss function is as shown in the formula.
[0110]
[0111] In the training process, quadruple data is randomly sampled from the experience replay pool. The current Q value and the target Q value are calculated using the Q value network and the target network respectively, and the parameters of the Q value network are updated according to the loss function. The parameters of the target network are periodically synchronized with the Q value network to ensure the stability of model training. Finally, the DQN model can predict the Q value of the optimal action in any state and select the best action by maximizing the Q value, realizing the optimal scheduling of assembly instructions.
[0112] · Action generation:
[0113] Find nodes with an in-degree of 0. In the data preprocessing stage, a feature matrix and an adjacency matrix are generated. The adjacency matrix is used to represent the dependency relationship between instructions. Nodes with an in-degree of 0 indicate that these nodes (assembly instructions) do not depend on the execution results of other nodes and can be executed directly. Traverse each row of the adjacency matrix and count the in-degree of each node (i.e., the number of elements with a value of 1 in that row). Find the nodes with an in-degree of 0, and these nodes form the initial action set A valid .
[0114] Mark illegal actions. Combine the hardware resources required by the nodes selected in the action to mark illegal actions. Illegal actions refer to those actions that cannot be executed due to resource limitations. For each node with an in-degree of 0, check whether the required hardware resources are available. If the resources required by a certain node are not available, remove it from the action set A valid .
[0115] Mark inefficient actions. Combine the number of nodes with an in-degree of 0 in the current state and mark those actions that clearly cannot obtain large rewards. For example, some nodes may contribute less to performance improvement or have a longer execution time. For each node with an in-degree of 0, evaluate its contribution to performance optimization, resource utilization, and whether it is on the longest path. According to predefined rules (such as excessive execution time or minimal contribution to performance improvement), mark these nodes as inefficient actions. Remove these inefficient actions from the action set A valid Remove these inefficient actions from the action set A
[0116] · Sampling strategy: The sampling method is uniform random sampling. Randomly select an action a from the action set A valid Execute action a and observe the feedback from the environment, including the reward r and the next state s'.
[0117] · Design of the reward function: The role of the reward function is to measure the optimization effect of the selected action. It can be divided into the reward for performance optimization, the reward for resource utilization, and the reward for selecting nodes on the longest path
[0118] r = ω1ΔT + ω2U + ω3V
[0119] where ΔT represents the reduction in execution time, U represents the resource utilization rate, and V represents whether the node on the longest path is selected
[0120] ω1, ω2, ω3 are weight parameters. Setting different parameters can guide the model to be more inclined to select actions that obtain greater rewards in a certain aspect
[0121] 1. Reward for performance optimization. ΔT represents the reduction in execution time. If the execution time is reduced, the reward is positive; otherwise, it is negative. Calculate ΔT by comparing the execution times of the current state and the next state
[0122] 2. Reward for resource utilization. U represents the resource utilization rate. If the resource utilization rate is increased, the reward is positive; otherwise, it is negative. Calculate U by comparing the resource utilization rates of the current state and the next state
[0123] 3. Reward for the longest path. V represents whether the node on the longest path is selected. If the node on the longest path is selected, then V = 1; otherwise, V = 0. Check whether the currently selected action is on the longest path
[0124] · Generation of quadruples: The samples in the experience replay pool are in the form of {sars'} quadruples, where a represents the current state, r represents the reward value after executing the action, s, s' represent the current state and the next state reached after executing action a, and s, s' are the global vectors obtained after being extracted by the GNN model
[0125] · Storage and Update: Store each quadruple into the experience replay pool and set a threshold in combination with the computing power of the laboratory. When the capacity of the pool reaches the set threshold, use the FIFO (First In First Out) mechanism to remove old samples. Initialize a queue as the experience replay pool. Each time a new quadruple is generated, add it to the queue. If the queue length exceeds the set threshold, the oldest sample in the queue.
[0126] 1) DQN Network Structure
[0127] The core of the DQN model is two neural networks:
[0128] · Q Network (Q-Network): Used to estimate the Q-values generated by taking various actions in a given state. The Q-value represents the expectation of the long-term cumulative reward that the system can obtain after taking this action in this state. The input is the feature representation of the current state, and the output is a Q-value vector.
[0129] · Target Network: Regularly copy parameters from the Q-value network and is used to calculate the target Q-value to stabilize the training process
[0130] training process
[0131] The present invention is implemented based on the LLVM framework. LLVM is a modular, flexible, and highly extensible compiler infrastructure, which is widely used in the development of modern compilers. Its core consists of a front end, intermediate code optimization, and a back end, where: Intermediate code optimization is the key link to improve the program execution performance. Integrate the optimization method into the intermediate code optimization stage of LLVM as an independent Pass. During the optimization process, use the above-mentioned action generation, sampling strategy, reward function design, quadruple generation, storage and update, and GNN network and DQN network structure to automatically optimize the arrangement of assembly instructions.
[0132] During the implementation process, by analyzing the basic block information generated by LLVM IR (Intermediate Representation), convert it into a graph structure suitable for the method of the present invention to process, including a feature matrix and an adjacency matrix. The optimization process takes this as the input, trains and optimizes through a joint model of reinforcement learning and graph neural network, and finally outputs an optimized instruction sequence. The optimized result can be docked into the LLVM back-end process, replace the assembly instruction part in the IR, and finally can be used by the back end to generate target code with better performance.
[0133] In addition, integrating this method into the LLVM framework not only improves the usability of the method but also fully demonstrates its good scalability and portability. Since the LLVM framework supports multiple language front-ends and hardware architecture back-ends, the design of this method can easily adapt to different processor architectures and instruction set requirements. By simply making appropriate adjustments to the key Passes in the framework, it can be extended to other platforms, providing support for multiple fields such as high-performance computing and embedded systems. Such an architecture design not only meets the high requirements of modern computing tasks for optimization efficiency but also demonstrates the wide applicability and flexibility of this method in practical applications.
[0134] Although the present invention has been described above, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many variations without departing from the purpose of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. An assembly instruction rearrangement optimization system, characterized in that, The system includes an assembly instruction preprocessing module, a first assembly instruction analysis module, a second assembly instruction analysis module, an assembly instruction compression unit, and an assembly instruction optimization unit; the instruction compression unit includes a first instruction status compression module and a second instruction status compression module; where: The instruction preprocessing module splits the assembly code into basic blocks according to lexical, syntactic, and semantic rules. The first instruction analysis module constructs an instruction adjacency matrix based on the mutual logical relationships of the code sequences in the basic blocks. The second instruction analysis module extracts feature vectors in the basic blocks to construct an instruction feature matrix. The graph neural network is trained using the instruction adjacency matrix and the instruction feature matrix according to the following formula to construct the assembly instruction compression unit. Wherein: represents the feature representation of node v at the k-th layer in the model; N(v) is the set of neighbor nodes of node v; The first instruction status compression module calculates and generates a global action status feature vector for the adjacency matrix and the feature matrix according to a dynamic filtering function. The second instruction status compression module generates a global spatial status feature vector by fusing a single graph convolutional layer with a fully connected layer. The reinforcement learning network is initialized and trained using the instruction adjacency matrix and the instruction feature matrix according to the following function to construct the assembly instruction optimization unit. Wherein: represents the loss function; represents the expectation; r represents the immediate reward; γ represents the discount factor; Q target represents the target network; Q represents the q-value network; s represents the current state; a represents the action taken in the current state; s′ represents the next state transferred to after taking action a in the current state s; a′ represents the possible action in the next state s′; The instruction optimization module selects the optimal action vector corresponding to the global state feature vector from the global action feature vector through a reward mechanism and outputs it.
2. The optimized system for rearranging assembly instructions according to claim 1, wherein The second instruction status compression module generates a global spatial status feature vector by fusing a single graph convolutional layer with a fully connected layer. The process includes: Normalize the instruction adjacency matrix of the single graph convolutional layer according to the following formula: Wherein: represents the adjacency matrix with self-loops added. D represents the degree matrix, which is a diagonal matrix, and its diagonal element D(i,i) represents the degree of the node; is the square root of the inverse matrix of the degree matrix; and if and only if there exists at least one node such that and Given the adjacency matrix Wherein: and if and only if there exists such that and Perform multi-layer fully connected layer calculations on different single graph convolutional layers according to the following formula: Where: H l is the result calculated at the l-th layer, and the input H at the 0-th layer 0 is the instruction feature matrix; Θ l is the updatable and weight-shared weight matrix at the l-th layer; σ represents the activation function; Obtain the global status feature vector by fusing the multi-layer fully connected layers according to the following formula Where: R(H) represents a function for compressing the intermediate state matrix H; N represents the number of rows of the intermediate state matrix; represents the feature of the i-th row of the intermediate state matrix H; σ represents the activation function; represents the obtained global state feature vector.
3. The optimized system for rearranging assembly instructions according to claim 1, wherein: The process by which the instruction optimization module selects the optimal action vector corresponding to the global spatial status feature vector from the global action feature vector through a reward mechanism includes: Construct an experience replay pool for the interaction between the global spatial status feature vector and the global action status, including: a quadruple {sars′} composed of the current state s, the action a taken, the reward r obtained, and the next state s′ Update the reward after performing the action according to the following formula: r = ω1ΔT + ω2U + ω3V where: ΔT represents the reduction in execution time, U represents the resource utilization rate, V indicates whether to select nodes on the longest path, and ω1, ω2, ω3 are weight parameters; Select and output the corresponding global spatial status feature vector from the global action feature vector according to the updated reward value.
4. The method for optimizing the rearrangement of assembly instructions according to the system described in any one of claims 1-3, characterized in that: It includes the following steps: Assembly instruction preprocessing stage: Split the assembly code into basic blocks according to lexical, syntactic, and semantic rules. Construct an instruction adjacency matrix based on the mutual logical relationships of the code sequences in the basic blocks. Extract feature vectors in the basic blocks to construct an instruction feature matrix. The graph neural network is trained using the instruction adjacency matrix and the instruction feature matrix according to the following formula to construct the assembly instruction compression unit. Wherein: represents the feature representation of node v at the k-th layer in the model; N(v) is the set of neighbor nodes of node v; The reinforcement learning network is initialized and trained using the instruction adjacency matrix and the instruction feature matrix according to the following function to construct the assembly instruction optimization unit. Wherein: represents the loss function; represents the expectation; r represents the immediate reward; γ represents the discount factor; Q target represents the target network; Q represents the q-value network; s represents the current state; a represents the action taken in the current state; s′ represents the next state transferred to after taking the action a in the current state s; a′ represents the possible action in the next state s′; Assembly instruction optimization stage: Calculate and generate a global action status feature vector for the adjacency matrix and the feature matrix according to a dynamic filtering function. Generate a global spatial state feature vector by fusing a single-image convolutional layer with a fully connected layer; Construct an assembly instruction optimization unit by initializing and training a reinforcement learning network using an instruction adjacency matrix and an instruction feature matrix; Output an optimal action vector corresponding to the global state feature vector by selecting from the global action feature vector through a reward mechanism.