Compiler backend instruction scheduling optimization method and system based on mamba state space model

CN122653580APending Publication Date: 2026-08-28CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610788383.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0005](1)静态规则难以适配异构微架构:列表调度的优先级函数依赖手工设计的静态代价模型,无法自适应区分Intel Golden Cove、AMD Zen 4、ARM Cortex-X4等现代高性能核心的精细微架构特性,导致跨平台性能损失

Benefits of technology

[0109] The microarchitecture descriptor vector is constructed by reading parameters such as the target CPU's issue width, number of functional units, and latency table from the LLVM ScheduleModel. This allows a single trained model to adapt to different CPUs (Intel/AMD/ARM) by switching vectors, significantly reducing the engineering cost of multi-platform deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122653580A_ABST
    Figure CN122653580A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of compiler optimization, and specifically discloses a compiler back-end instruction scheduling optimization method and system based on a Mamba state space model, which comprises the following steps: obtaining an instruction sequence in a compiler back-end machine basic block, constructing a scheduling directed acyclic graph, extracting a multi-dimensional feature vector and an instruction embedding matrix of each instruction, and encoding a target CPU micro-architecture descriptor into a global context vector; inputting the instruction embedding matrix into a Mamba selective state space model to output a scheduling priority score of each instruction; iteratively decoding in each scheduling slot to generate a legal instruction scheduling arrangement that satisfies all data dependency constraints; feeding the legal instruction scheduling arrangement back to the compiler back-end and adjusting the priority selection process in the original list scheduling to realize instruction scheduling optimization. According to the technical scheme, the instruction scheduling problem is formalized into a sequence-to-sequence optimization problem, a high-quality instruction scheduling arrangement is obtained, and the real-time requirement in the compilation period is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of compiler optimization technology, and relates to a compiler backend instruction scheduling optimization method and system based on the Mamba state space model. Background Technology

[0002] Modern processors generally feature deep pipelines and high parallelism, so whether the target code can fully utilize instruction-level parallelism often directly affects processor efficiency. The instruction scheduling process in the compiler backend needs to rearrange the instruction order within basic blocks without compromising program semantics, in order to minimize pipeline wait and improve final runtime performance.

[0003] In existing compiler implementations, systems such as GCC and LLVM / Clang typically still rely primarily on list-based scheduling mechanisms based on dependency graphs. Recent research has also explored using reinforcement learning (RL) or graph neural networks (GNNs) to assist compiler optimization decisions. For example, Transformer-based code models (such as CodeBERT, CodeT5, and LLM Compiler) have demonstrated the capabilities of large language models in source code understanding.

[0004] However, existing technologies still have the following drawbacks:

[0005] (1) Static rules are difficult to adapt to heterogeneous microarchitectures: The priority function of list scheduling relies on a manually designed static cost model, which cannot adaptively distinguish the fine microarchitecture characteristics of modern high-performance cores such as Intel Golden Cove, AMD Zen 4, and ARM Cortex-X4, resulting in cross-platform performance loss.

[0006] (2) Insufficient coverage of the heuristic search space: List scheduling is a greedy algorithm that only makes local optimal decisions and cannot explore a better global scheduling scheme in the exponentially ordered space.

[0007] (3) Transformer complexity bottleneck: Attention computation is too slow for reasoning on long instruction sequences containing hundreds of instructions, failing to meet compile-time real-time requirements.

[0008] (4) GNN's long-range dependency modeling is insufficient: The message passing depth of graph neural networks is limited, and the ability to capture long-link scheduling constraints that span multiple layers of dependencies is weak, and there is an oversmoothing problem.

[0009] (5) Limitations of LLM’s understanding of IR: Existing large language models have weak control flow reasoning capabilities for compiler IR and cannot be reliably used for instruction-level scheduling decisions. Summary of the Invention

[0010] The purpose of this invention is to address the aforementioned problems in existing technologies by proposing a compiler backend instruction scheduling optimization method and system based on the Mamba state space model.

[0011] To achieve the above objectives, the basic solution of this invention is: a compiler backend instruction scheduling optimization method based on the Mamba state space model, comprising the following steps:

[0012] Obtain the instruction sequence from the compiler backend machine basic block, analyze the dependencies between instructions, and construct a scheduling directed acyclic graph;

[0013] The scheduling directed acyclic graph is topology-aware serialization is performed to extract the multi-dimensional feature vector and instruction embedding matrix of each instruction, and the target CPU microarchitecture descriptor is encoded into a global context vector.

[0014] The instruction embedding matrix and global context vector are input into the Mamba selective state space model to complete global sequence modeling in linear time complexity and output the scheduling priority score of each instruction.

[0015] By using a dependency-constraint-aware pointer network, iterative decoding is performed in each scheduling slot to generate a valid instruction scheduling arrangement that satisfies all data dependency constraints.

[0016] The generated legal instruction scheduling arrangement is fed back to the compiler backend, and the priority selection process in the original list scheduling is adjusted accordingly to achieve instruction scheduling optimization.

[0017] The working principle and beneficial effects of this basic scheme are as follows: This technical scheme formalizes the instruction scheduling problem into a sequence-to-sequence optimization problem. Leveraging the linear complexity advantage of the Mamba state-space model, it achieves efficient global modeling of long instruction sequences. It generates legal scheduling sequences that satisfy data dependency constraints and can directly feed this information back to the compiler backend to adjust the original scheduling decisions, thereby improving instruction-level parallelism.

[0018] Furthermore, the method for obtaining the instruction sequences in the compiler's backend machine basic blocks, analyzing the dependencies between instructions, and constructing a directed acyclic graph of the schedule is as follows:

[0019] For the n MachineInstr instruction sequences in the LLVM MachineBasicBlock The analysis examines register read / write relationships and memory alias relationships between instruction sequences, including: true dependencies (RAW: read after write), anti-dependencies (WAR: write after read), output dependencies (WAW: write after write), and memory alias dependencies.

[0020] Construct a Directed Acyclic Graph (DAG) for scheduling :

[0021] DAG ,

[0022] Among them, vertex set Each edge in the directed edge set E corresponds one-to-one with the instruction. Indication of instructions right There is a delay. Dependence on each cycle.

[0023] Identify RAW, WAR, WAW, and memory alias dependencies, and construct a scheduling DAG with delay annotations to provide an accurate dependency constraint basis for subsequent feature extraction and scheduling decisions.

[0024] Furthermore, the steps for performing topology-aware serialization on the scheduled directed acyclic graph and extracting the multidimensional feature vector and instruction embedding matrix for each instruction are as follows:

[0025] Using the longest path algorithm for directed acyclic graphs (DAGs), the critical path depth (CPD) of each instruction sequence is calculated:

[0026] ,

[0027] in, It is the edge The number of dependency delay cycles, if If there is no successor node (out-degree is 0), then CPD( If ) = 0; There is a successor node, traversal All successor nodes Take the maximum value of "edge delay plus successor node CPD value" as the value. The critical path depth (CPD) can be calculated sequentially in reverse topological order (from the endpoint to the starting point) to obtain the CPD value for each instruction.

[0028] Extract multidimensional feature vectors from each MachineInstr instruction sequence. This includes: opcode embedding, operand type encoding, one-hot encoding of target functional units, normalized critical path depth (CPD) value, dependency graph structure context, ready time slots, and latency sensitivity coefficients, specifically:

[0029] Opcode embedding: via a learnable embedding table Map machine opcodes to dense vectors, dimension ;

[0030] Operand type encoding: One-hot concatenation of register type (integer / floating-point / vector), immediate flag, and memory operand flag; dimension ;

[0031] One-hot encoding of the target functional unit refers to the encoding of the execution port required by the identifier instruction, and its dimensions are... Equal to the number of functional unit types of the target platform;

[0032] Critical path depth normalized value The range is [0,1]; where CPD( ) represents the critical path depth of the k-th instruction, which is the longest weighted path length from this instruction to the endpoint of the DAG; It is the kth instruction, that is, the kth MachineInstr node in the instruction sequence; This represents the maximum CPD of all instructions in the current basic block, which is the total length of the critical path in the entire DAG;

[0033] Dependency Graph Structure Context: Multi-hot encoding of the predecessor node set from the edge set E of the constructed scheduling DAG Multi-hot coding of successor node sets Binary markers for determining whether a path is on the critical path Specifically:

[0034] judge Whether it is on the critical path is determined by the condition: IsCPD ( )=1, if CPD( ) + From the root node to Longest path = If an instruction satisfies both "longest path to destination" and "longest path from origin", it means that it is on the global critical path; the result is a binary label, 0 or 1, which is directly used as a scalar component and concatenated into the feature vector.

[0035] Readiness time slot estimation: based on the earliest launchable period of the current partial schedule. Normalized to [0,1];

[0036] Delay sensitivity coefficient: When constructing the scheduling DAG, the latency sensitivity coefficient for each edge The latency value lat is read and marked from the instruction latency table (InstrItineraries / SchedReadWrite) of the LLVM ScheduleModel, and then iterated through. Take the maximum delay value (lat_max) from all outgoing edges:

[0037] ,

[0038] ,

[0039] For leaf nodes with an out-degree of 0 or nodes with a CPD of 0, LS( ) = 0, LS( ) is an instruction The delay sensitivity coefficient;

[0040] Concatenate all components to obtain the original feature vector of each instruction sequence. ,in , Assign opcode embedding dimension; form instruction embedding matrix d is the total dimension of the feature vector, which is equal to the sum of the dimensions of each component. The dimension for embedding the opcode is set to 128; The dimension for encoding operand types is set to 16. The dimension of the one-hot encoding for the target functional unit is equal to the number of functional unit types on the target platform.

[0041] Extracting multi-dimensional feature vectors enables the model to fully perceive the semantic, structural, and temporal characteristics of instructions, thereby improving the accuracy of scheduling decisions.

[0042] Furthermore, the method for encoding the target CPU microarchitecture descriptor into a global context vector is as follows:

[0043] Perform an improved topological sort on the DAG to generate candidate sequences. The DAG structural context information (predecessor set, successor set, critical path marker) of each node is concatenated into a multi-dimensional feature vector to form a structure-aware input embedding matrix. ,in , For structural context dimension;

[0044] Microarchitectural parameters of the target CPU are extracted from the LLVM target description file (ScheduleModel) to form a microarchitectural feature vector a, including: superscalar issue width (IssueWidth), number of each functional unit type, etc. Delay table entries for each functional unit The out-of-order execution window size (ROBSize), rearrangement buffer depth, and L1 instruction cache line size are normalized and then concatenated to form the original microarchitecture vector. Then, it is projected as a context vector through a fully connected network. :

[0045] ,

[0046] in, LayerNorm represents the learnable parameters and is used for layer normalization.

[0047] By injecting cross-attention mechanism into the Mamba-IS model, conditional reasoning for different target microarchitectures can be achieved.

[0048] Explicitly encoding the DAG structure context information into the input embedding matrix compensates for the lack of control flow understanding in pure sequence models. The target CPU microarchitecture parameters are encoded as global context vectors and injected into the model through cross-attention, enabling a single model to adapt to multiple heterogeneous microarchitectures (Intel / AMD / ARM) without requiring retraining for each hardware platform.

[0049] Furthermore, the steps for inputting the instruction embedding matrix and global context vector into the Mamba selective state-space model to complete global sequence modeling with linear time complexity and output the scheduling priority score for each instruction are as follows:

[0050] The continuous-time state-space equation of the S6 operator in the Mamba selective state-space model is:

[0051] ,

[0052] ,

[0053] in, It is a learnable diagonal matrix. = 512 is the latent space dimension of the model, corresponding to the feature dimension of each position in the input sequence; = 64 represents the hidden state dimension, i.e., the hidden variables in the state-space model. The dimensions; matrices B(·) and C(·) are both derived from the current input. The linear projection is dynamically generated, where D is the scalar jump connectivity coefficient. The instruction embedding at time t is x; This is the derivative of the hidden state with respect to time, i.e., the rate of change of the hidden state over continuous time. This represents the output vector at time t, corresponding to the encoded representation of the instruction at that time. This represents the hidden state vector at time t, with dimension . =64, carrying contextual information about the historical sequence;

[0054] After zero-order hold (ZOH) discretization, the step size Δ is determined by the input: ), thus obtaining the discrete recursive form:

[0055] ,

[0056] ,

[0057] ,

[0058] ,

[0059] Wherein, the state transition matrix Input matrix and output matrix All are from the current input Dynamically generated; The state transition matrix is ​​the discretized state transition matrix, which is the input when processing the τth instruction. Dynamically generated, controlling the hidden state The degree of preservation at the current time step, The input matrix is ​​the discretized form, derived from the current input. Dynamically generated, controlling the current input The degree of injection into the hidden state, The output matrix is ​​derived from the current input. Dynamically generated, controlling the hidden state How to map to output , Let be the hidden state vector at the τth instruction, with dimension . =64, carrying contextual information about the historical sequence. This is the output vector at the τ-th instruction, corresponding to the encoded representation of the current instruction; The discretization step size controls the granularity of discretization of the continuous-time state-space equations, determined by the current input. Dynamically generated, each instruction corresponds to a different step size, which is always positive through the softplus function; softplus is the activation function, softplus(x)=log(1+eˣ), ensuring that Δ is always positive; The input is a learnable linear projection matrix. Mapped to the pre-activation value of the step size; Let τ be the input vector at the τ-th instruction, which is the embedded representation of the current instruction; For learnable bias terms;

[0060] The selective mechanism enables the Mamba selective state space model to adaptively and selectively retain or forget scheduling history information in the preceding hidden state according to the current instruction context, effectively modeling long-link data dependencies.

[0061] The Mamba selective state space consists of an input projection layer and L=6 stacked Mamba Blocks. The input projection layer linearly maps the instruction embedding matrix F to the latent space.

[0062] ,

[0063] in, The output of the input projection layer is the initial latent representation matrix after the instruction sequence is mapped to the latent space, which serves as the input to the first layer Mamba Block; F is the instruction embedding matrix with shape n×d, obtained from the feature extraction stage; The learnable linear projection weight matrix has a shape of d× Mapping the feature vector from the original dimension d to the latent space dimension ; The bias term is a learnable term with dimension . n is the total number of instructions in the current basic block; =512: Latent space dimension, that is, the representation dimension of each instruction in the latent space;

[0064] The calculation process for the l-th layer (l=1,…,6) Mamba Block is as follows:

[0065] ,

[0066] ,

[0067] ,

[0068] Where SSM(·) is the S6 selective scan operator, and σ is the Sigmoid function. , The LayerNorm is a learnable projection matrix, and its value is RMSNorm. This represents the element-wise product of the SSM output and the gate signal in the l-th Mamba Block, i.e., the intermediate representation after gated operation, with a shape of n× , For the first The output of the Mamba Block layer is obtained after residual concatenation, and has a shape of n× , This represents the intermediate projection result of the l-th layer, i.e. through The representation after linear projection, the transition amount before residual connection, has a shape of n× ;

[0069] After the outputs of the Mamba Blocks in layers 3 and 6, a lightweight microarchitecture cross-attention layer (CAA) is inserted respectively:

[0070] ,

[0071] in, This represents the instruction implicit representation matrix obtained after the output of the l-th layer (l=3 or l=6) Mamba Block is processed by the microarchitecture cross-attention layer, and serves as the input to the next layer of Mamba Block; For a learnable query projection matrix, Mapped to attention query; For the learnable key projection matrix, Mapped to attention key; For the learnable Value projection matrix, Mapped to attention value; For the target CPU microarchitecture context vector, dimension =64, broadcast as Key and Value to each position in the sequence, injecting microarchitectural information into the encoding of each instruction; Q (Query): Query matrix, output by the l-th layer Mamba Block. through The projection yields a vector representing each instruction "questioning" the microarchitecture information; K (Key): the key matrix, derived from the microarchitecture context vector. through The projection yields an "index" representing microarchitectural information; V (Value): a value matrix derived from the microarchitectural context vector. through The projection yields the actual content representing the microarchitectural information, which is then weighted by attention weights and injected into the encoding of each instruction.

[0072] Context vector The Key and Value for attention are broadcast to each position in the sequence, enabling instruction encoding to incorporate global information about the target microarchitecture and achieve cross-platform conditional reasoning.

[0073] After global average pooling and output MLP, the scheduling priority score vector of each instruction is obtained. :

[0074] ,

[0075] in, No. The output of the layer Mamba Block, P represents an n-dimensional real vector, where n is the number of instructions in the current basic block.

[0076] The selective mechanism of input dependencies enables the model to adaptively retain or forget historical states, effectively capturing long-link data dependencies. The microarchitectural cross-attention layer enables instruction encoding to fuse target hardware information, achieving cross-platform conditional reasoning.

[0077] Furthermore, through a dependency-constraint-aware pointer network, iterative decoding is performed in each scheduling slot to generate a valid instruction scheduling permutation that satisfies all data dependency constraints. The specific method is as follows:

[0078] Check if the predecessor of each instruction in the directed acyclic graph (DAG) G has been fully scheduled, and add the instructions that meet the conditions to the dynamic ready set. :

[0079] ,

[0080] in, These are candidate instructions that need to be determined whether to be added to the ready set. for The predecessor instruction, that is, there is an edge in the DAG ( → The instructions, For edge ( , The number of dependency latency cycles on ) Describes the set of directed edges of a DAG. This represents the set of instructions that have been scheduled up to the s-th scheduling slot. When all the front-ends are in it, Only then can it be added to the ready set. ;

[0081] Calculate attention score :

[0082] ,

[0083] ,

[0084] in, Encode the Mamba output of the i-th instruction. Let GRU be the hidden state of the decoder at step s. , , These are learnable parameters;

[0085] Select the instruction with the highest score when there are no resource conflicts. :

[0086] ,

[0087] like If a functional unit conflict exists between a scheduled instruction and another instruction (the demand on the same port exceeds capacity in the same cycle), a Beam Search with a beam width of k=4 is initiated. From the ready set, k candidate paths with the highest scores are selected, and the first instruction of the path with no conflict and the highest cumulative score is chosen as the first instruction. ;

[0088] Will Join the scheduled set Update the decoder hidden state. , into the scheduling slot ;

[0089] After decoding, output a valid scheduling sequence that satisfies all data dependencies and resource constraints. , For scheduling and arrangement, This is the index of the instruction corresponding to the nth scheduling slot in the original instruction sequence, i.e., the original number of the last scheduled instruction.

[0090] Handle resource conflicts between functional units, avoid greedy solutions from getting trapped in local optima, and generate legal and high-quality scheduling arrangements.

[0091] Furthermore, it also includes a two-stage training step for the Mamba selective state-space model, specifically:

[0092] Phase 1 (Supervised Pre-training): Using LLVM MCA (LLVM Machine Code Analyzer) simulation IPC (Instructions Per Cycle) as the ranking quality label, cross-entropy loss is employed. Supervised learning of the pointer network decoding process:

[0093] ,

[0094] in, For the expert scheduling sequence in the 1st The instruction index of each scheduling slot, where n is the total number of scheduling slots, i.e., the total number of instructions; the optimizer uses... Train for 100 epochs with a batch size of 128;

[0095] Phase Two (PPO Reinforcement Learning Fine-tuning): The pre-trained model is fine-tuned using the IPC improvement measured by the measured PMU performance counter as the reward. The reward function... Defined as:

[0096] ,

[0097] Among them, IPC_Mamba-IS( ) represents the measured IPC of the program when using the scheduling arrangement π, and IPC_ListScheduler represents the measured IPC of the LLVM default list scheduling;

[0098] The PPO (Proximity Policy Optimization) algorithm is used to prune parameters. Value function coefficients Entropy regularity coefficient The training lasted for 20 epochs; the PPO objective function was... for:

[0099] ,

[0100] in, Importance sampling ratio, For generalized advantage estimation, For the value function loss, This is the policy entropy regularization term; Importance sampling ratio measures the ratio of the probability of the old and new strategies choosing the same action in the same state. For the current policy in state Select action The probability of; For the old strategy before the update, in the state Select action The probability of θ; θ is the learnable parameter of the current policy; The parameters for the old strategy before the update; The action selected for step s, i.e. the instruction selected for the s-th scheduling slot; The state at step s is a comprehensive representation of the current ready set, the scheduled set, and the decoder's hidden state. This represents the truncation function, which will... The value is limited to the range of [1-ε, 1+ε] to prevent the difference between the old and new strategies from being too large, which would lead to training instability. ε=0.2 is the pruning parameter.

[0101] The first stage is supervised pre-training (LLVM MCA simulation IPC as labels) to quickly learn the basic scheduling strategy. The second stage is PPO reinforcement learning fine-tuning (measured PMU performance counter as reward) to further improve the scheduling quality and make the model output closer to the real hardware performance.

[0102] The present invention also provides a compiler backend instruction scheduling optimization system based on the method described in the present invention, comprising:

[0103] The program representation and serialization module is used to build the scheduling DAG, extract instruction features and complete topology-related serialization processing, while generating a microarchitecture context representation;

[0104] The model computation module is used to output instruction priority scores based on the Mamba Block stacking structure and microarchitecture cross-attention mechanism;

[0105] The scheduling decoding module is used to combine dynamic ready set constraints and iterative decoding process to obtain a valid scheduling sequence;

[0106] The LLVM integrated feedback module is used to apply the scheduling results to the LLVM backend via the MachineSchedStrategy interface and supports the training data generation process.

[0107] This system comprises four main modules: program representation, model computation, scheduling decoding, and LLVM integration, forming an end-to-end automated scheduling optimization closed loop.

[0108] Furthermore, the program characterization and serialization module includes a microarchitecture context injection unit, which is used to read parameter information such as the target CPU's issue width, number of functional units, and delay table from the LLVM ScheduleModel to construct a microarchitecture descriptor vector. .

[0109] The microarchitecture descriptor vector is constructed by reading parameters such as the target CPU's issue width, number of functional units, and latency table from the LLVM ScheduleModel. This allows a single trained model to adapt to different CPUs (Intel / AMD / ARM) by switching vectors, significantly reducing the engineering cost of multi-platform deployment. Attached Figure Description

[0110] Figure 1 This is a flowchart illustrating the compiler backend instruction scheduling optimization method based on the Mamba state space model of this invention.

[0111] Figure 2 This is a flowchart illustrating the program representation and serialization module of the compiler back-end instruction scheduling optimization system of the present invention;

[0112] Figure 3 This is a schematic diagram of the structure of the Mamba selective state space model in the compiler backend instruction scheduling optimization method based on the Mamba state space model of this invention;

[0113] Figure 4 This is a flowchart illustrating the S6 operator of the compiler backend instruction scheduling optimization method based on the Mamba state space model of this invention.

[0114] Figure 5 This is a flowchart illustrating the scheduling and decoding module of the compiler backend instruction scheduling optimization system of the present invention;

[0115] Figure 6 This is a schematic diagram of the two-stage training process of the Mamba selective state space model in the compiler backend instruction scheduling optimization method based on the Mamba state space model of this invention.

[0116] Figure 7This is a comparison chart of IPC of the compiler backend instruction scheduling optimization method based on the Mamba state space model of this invention with the baseline method on SPEC CPU 2017. Detailed Implementation

[0117] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0118] In the description of this invention, it should be understood that the terms "longitudinal", "lateral", "up", "down", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0119] In the description of this invention, unless otherwise specified and limited, it should be noted that the terms "installation", "connection" and "linking" should be interpreted broadly. For example, they can refer to mechanical or electrical connections, or internal connections between two components. They can be direct connections or indirect connections through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms according to the specific circumstances.

[0120] This invention discloses a compiler backend instruction scheduling optimization method based on the Mamba Selective State Space Model (SSM). The instruction scheduling problem is formalized as a structured sequence-to-sequence optimization problem. It takes the instruction dependency graph and microarchitectural context as input, the Mamba linear sequence model as encoder, the dependency constraint-aware pointer network as decoder, and LLVM MachineSchedStrategy as deployment interface to learn and output a high-quality instruction scheduling arrangement end-to-end.

[0121] like Figure 1 As shown, the compiler backend instruction scheduling optimization method based on the Mamba state space model includes the following steps:

[0122] Obtain the instruction sequence (MachineInstr) in the machine basic block of the compiler backend (LLVM), analyze the dependencies between instructions: true dependencies (RAW), anti-dependencies (WAR), output dependencies (WAW), and memory alias dependencies, and construct a scheduling directed acyclic graph (DAG).

[0123] Topology-aware serialization is performed on the directed acyclic graph of the scheduling system, extracting the multidimensional feature vector and instruction embedding matrix for each instruction, and encoding the target CPU microarchitecture descriptor into a global context vector (the global context vector is injected into the model through a cross-attention mechanism; specifically, the global context vector...). In the CAA layer following the output of the Mamba Block in layers 3 and 6, the Key and Value for attention are broadcast to each position in the sequence and cross-attention is performed with the implicit representation of each instruction. This integrates the microarchitecture information of the target CPU into the encoding of each instruction, enabling conditional inference of a single model for different target microarchitectures.

[0124] Input the instruction embedding matrix and global context vector into the Mamba (Selective State Space Model, S6) selective state space model to complete global sequence modeling in linear time complexity and output the scheduling priority score of each instruction;

[0125] By using a dependency-constraint-aware pointer network, iterative decoding is performed in each scheduling slot to generate a valid instruction scheduling arrangement that satisfies all data dependency constraints.

[0126] The generated valid instruction scheduling order is fed back to the compiler backend (via LLVM's MachineSchedStrategy interface), and the priority selection process in the original list scheduling is adjusted accordingly. The score of each instruction in the priority score vector P output by Mamba-IS is directly used as the sorting key for pickNode(), replacing the original heuristic scoring. Each scheduling slot still selects instructions from the current ready set, but the selection criterion is changed to the score of the corresponding instruction in P; the instruction with the highest score is emitted. The goal is to ensure that the final instruction order produced by list scheduling is consistent with the permutation π calculated by Mamba-IS; overwriting the pickNode() score with the P vector is the means to achieve this goal. This optimizes instruction scheduling.

[0127] This invention overcomes the shortcomings of traditional list-based static cost models that cannot adapt to heterogeneous microarchitectures, and the Transformer method. The compiler addresses the real-time bottleneck caused by complexity, and the limitations of existing large language models in understanding the intermediate representation control flow of the compiler. The resulting compiled code achieves a significant improvement in IPC on the SPEC CPU 2017 benchmark, with a single basic block inference latency of less than 5 milliseconds.

[0128] In a preferred embodiment of the present invention, the method for obtaining the instruction sequence in the compiler backend machine basic block, analyzing the dependencies between instructions, and constructing a scheduling directed acyclic graph is as follows:

[0129] For the n MachineInstr instruction sequences in the LLVM MachineBasicBlock The analysis examines register read / write relationships and memory alias relationships between instruction sequences, including: true dependencies (RAW: read after write), anti-dependencies (WAR: write after read), output dependencies (WAW: write after write), and memory alias dependencies.

[0130] Construct a Directed Acyclic Graph (DAG) for scheduling :

[0131] DAG ,

[0132] Among them, vertex set Each edge in the directed edge set E corresponds one-to-one with the instruction. Indication of instructions right There is a delay. Dependence on each cycle.

[0133] Systematic defects discovered in the study of large language model understanding compilers (IR) include low accuracy in control flow graph reconstruction tasks (GPT-4 full accuracy of only 39 / 164, DeepSeek R1 of 57 / 77) and a pass rate of less than 0.37 in instruction-level execution inference.

[0134] This invention explicitly encodes control flow-aware features (DAG-dependent topology and critical path markings) into the instruction feature vector during the serialization stage, thereby compensating for the lack of control flow understanding in pure sequence models from the data structure level and avoiding reliance on the implicit understanding of LLVM IR by LLM.

[0135] In a preferred embodiment of the present invention, the steps of performing topology-aware serialization on the scheduling directed acyclic graph and extracting the multidimensional feature vector and instruction embedding matrix of each instruction are as follows:

[0136] Using the longest path algorithm for Directed Acyclic Graphs (DAGs), the Critical Path Depth (CPD) of each instruction sequence is calculated to provide a foundation for subsequent feature extraction. For each node vᵢ, its CPD is defined as the length of the longest weighted path from that node to the endpoint of the DAG, and the recursive formula is as follows:

[0137] ,

[0138] in, It is the edge The number of dependency delay cycles, if If there is no successor node (out-degree is 0), then CPD( If ) = 0; If there is a successor node, calculate according to the recursive formula mentioned above and traverse. All successor nodes Take the maximum value of "edge delay plus successor node CPD value" as the value. The critical path depth is calculated in reverse topological order. The CPD values ​​of all its successor nodes have been obtained and can be reused directly; by calculating in reverse topological order (from the end point to the start point), the CPD value of each instruction can be obtained.

[0139] Extract multidimensional feature vectors from each MachineInstr instruction sequence. This includes: opcode embedding, operand type encoding, one-hot encoding of target functional units, normalized critical path depth (CPD) value, dependency graph structure context, ready time slots, and latency sensitivity coefficients, specifically:

[0140] Opcode embedding: via a learnable embedding table (Embedded tables can be learned) It is a shape of |V op | × The learnable matrix, where |V op | represents the total number of machine opcodes on the target platform. =128 is the embedding dimension. This matrix is ​​updated via gradient descent during model training and is fixed after training. It maps machine opcodes to dense vectors (each instruction's opcode is first mapped to an integer index id, then the corresponding row is directly retrieved from the embedding table). = [id] ∈ That is, using the opcode index as the row number, the corresponding 128-dimensional dense vector is obtained by looking up the table in the embedding table, which serves as the feature representation of the instruction opcode. ;

[0141] Operand type encoding: One-hot concatenation of register type (integer / floating-point / vector), immediate flag, and memory operand flag; dimension ;

[0142] Suppose an instruction has the following operand characteristics: the register type is an integer register, there are no immediate values, and there are memory operands. Then, the one-hot encoding of each component is as follows:

[0143] Register type (3-bit one-hot): Integer = 1, Floating point = 0, Vector = 0, Register type encoding is [1, 0, 0];

[0144] Immediate value flag (1 bit): No immediate value, the immediate value flag is encoded as [0];

[0145] Memory operand flag (1 bit): If there is a memory operand, the memory operand flag is encoded as [1];

[0146] After concatenation, a vector of [1, 0, 0, 0, 1, …] is obtained, and the remaining bits are padded to dimension 16. This results in the operand type encoding vector for the instruction. .

[0147] One-hot encoding of target functional units: identifies the execution port required by the instruction (integer ALU, floating-point ALU, Load, Store, branch, vector, etc.), dimension Equal to the number of functional unit types of the target platform;

[0148] Critical path depth normalized value The range is [0,1]; where CPD( ) represents the critical path depth of the k-th instruction, which is the longest weighted path length from this instruction to the endpoint of the DAG; It is the kth instruction, that is, the kth MachineInstr node in the instruction sequence; This represents the maximum CPD of all instructions in the current basic block, which is the total length of the critical path in the entire DAG;

[0149] Dependency Graph Structure Context: Multi-hot encoding of the predecessor node set from the edge set E of the constructed scheduling DAG Multi-hot coding of successor node sets Traverse all edges. The starting points of the edges that are the endpoints form the predecessor set. The endpoints of edges that serve as starting points form the successor set. The encoding method is a multi-hot vector of length n, where the positions of predecessor / successor nodes are set to 1, and the rest to 0. A binary marker is used to determine whether a node is on the critical path. Specifically:

[0150] judge Whether it is on the critical path is determined by the condition: IsCPD ( )=1, if CPD( ) + From the root node to Longest path = If an instruction satisfies both "longest path to destination" and "longest path from origin", it means that it is on the global critical path; the result is a binary label, 0 or 1, which is directly used as a scalar component and concatenated into the feature vector.

[0151] All three datasets are derived from existing results from the DAG construction and CPD calculation phases, requiring no additional calculations.

[0152] Readiness time slot estimation: based on the earliest launchable period of the current partial schedule. Normalized to [0,1];

[0153] Delay sensitivity coefficient: When constructing the scheduling DAG, the latency sensitivity coefficient for each edge The latency value lat is read and marked from the instruction latency table (InstrItineraries / SchedReadWrite) of the LLVM ScheduleModel, and then iterated through. Take the maximum delay value (lat_max) from all outgoing edges:

[0154] ,

[0155] ,

[0156] For leaf nodes with an out-degree of 0 or nodes with a CPD of 0, LS( ) = 0, LS( ) is an instruction The delay sensitivity coefficient;

[0157] Concatenate all components to obtain the original feature vector of each instruction sequence. ,in , Assign opcode embedding dimension; form instruction embedding matrix d is the total dimension of the feature vector, which is equal to the sum of the dimensions of each component. The dimension for embedding the opcode is set to 128; The dimension for encoding operand types is set to 16. The dimension of the one-hot encoding for the target functional unit is equal to the number of functional unit types on the target platform.

[0158] In a preferred embodiment of the present invention, the method for encoding the target CPU microarchitecture descriptor into a global context vector is as follows:

[0159] Perform an improved topological sort on the DAG to generate candidate sequences. (When breaking a topological tie, prioritize nodes with larger CPDs.) Concatenate the DAG structural context information (predecessor set, successor set, critical path markers) of each node into a multi-dimensional feature vector to form a structure-aware input embedding matrix. ,in , For structural context dimension;

[0160] Microarchitectural parameters of the target CPU are extracted from the LLVM target description file (ScheduleModel) to form a microarchitectural feature vector a, including: superscalar issue width (IssueWidth), number of each functional unit type, etc. Delay table entries for each functional unit The out-of-order execution window size (ROBSize), the rearrangement buffer depth, and the L1 instruction cache line size (these parameters are directly defined in the LLVM ScheduleModel object description file and can be read directly without additional calculations).

[0161] IssueWidth: Directly corresponds to the IssueWidth field in ScheduleModel.

[0162] Number of each functional unit type: The number of each type of execution unit is read directly from ProcResourceKinds;

[0163] Each functional unit's delay table entries: directly read the delay value corresponding to each opcode from the WriteLatency and ReadAdvance tables;

[0164] ROBSize: Directly corresponds to the MicroOpBufferSize field in ScheduleModel;

[0165] Reorder buffer depth: Same source as ROBSize, read directly;

[0166] L1 instruction cache line size: Read directly from the cache description of the target platform;

[0167] After reading, only normalization is needed, concatenating the data to form the original microarchitecture vector 'a', which is then projected through a fully connected network to form... (and after normalization, concatenate them to form the original microarchitecture vector) Then, it is projected as a context vector through a fully connected network. :

[0168] ,

[0169] in, LayerNorm represents the learnable parameters and is used for layer normalization.

[0170] By injecting cross-attention mechanism into the Mamba-IS model, conditional reasoning for different target microarchitectures can be achieved.

[0171] In a preferred embodiment of the present invention, such as Figure 3 As shown, the steps for inputting the instruction embedding matrix and global context vector into the Mamba selective state-space model to complete global sequence modeling with linear time complexity and output the scheduling priority score for each instruction are as follows:

[0172] Unlike the classic state-space model (S4) which uses a fixed parameter matrix, the state transition parameters of the S6 operator are dynamically determined by the current input, giving the model the ability to selectively filter information from the sequence.

[0173] like Figure 4 As shown, the continuous-time state-space equation of the S6 operator in the Mamba selective state-space model is:

[0174] ,

[0175] ,

[0176] in, It is a learnable diagonal matrix. = 512 is the latent space dimension of the model, corresponding to the feature dimension of each position in the input sequence; = 64 represents the hidden state dimension, i.e., the hidden variables in the state-space model. The dimensions; matrices B(·) and C(·) are both derived from the current input. The linear projection is dynamically generated, where D is the scalar jump connectivity coefficient. The instruction embedding at time t is x; This is the derivative of the hidden state with respect to time, i.e., the rate of change of the hidden state over continuous time. This represents the output vector at time t, corresponding to the encoded representation of the instruction at that time. This represents the hidden state vector at time t, with dimension . =64, carrying contextual information about the historical sequence;

[0177] After discretization with zero-order hold (ZOH), the step size Δ is determined by the input: ), thus obtaining the discrete recursive form:

[0178] ,

[0179] ,

[0180] ,

[0181] ,

[0182] Wherein, the state transition matrix Input matrix and output matrix All are from the current input Dynamically generated; The state transition matrix is ​​the discretized state transition matrix, which is the input when processing the τth instruction. Dynamically generated, controlling the hidden state The degree of preservation at the current time step, The input matrix is ​​the discretized form, derived from the current input. Dynamically generated, controlling the current input The degree of injection into the hidden state, The output matrix is ​​derived from the current input. Dynamically generated, controlling the hidden state How to map to output , Let be the hidden state vector at the τth instruction, with dimension . =64, carrying contextual information about the historical sequence. This is the output vector at the τ-th instruction, corresponding to the encoded representation of the current instruction; The discretization step size controls the granularity of discretization of the continuous-time state-space equations, determined by the current input. Dynamically generated, each instruction corresponds to a different step size, which is always positive through the softplus function; softplus is the activation function, softplus(x)=log(1+eˣ), ensuring that Δ is always positive; The input is a learnable linear projection matrix. Mapped to the pre-activation value of the step size; Let τ be the input vector at the τ-th instruction, which is the embedded representation of the current instruction; For learnable bias terms;

[0183] The selective mechanism enables the Mamba Selective State Space Model (Mamba-IS model) to adaptively and selectively retain or forget scheduling history information in previous hidden states based on the current instruction context, effectively modeling long-link data dependencies;

[0184] The Mamba selective state space consists of an input projection layer and L=6 stacked Mamba Blocks. The input projection layer linearly maps the instruction embedding matrix F to the latent space.

[0185] ,

[0186] in, The output of the input projection layer is the initial latent representation matrix after the instruction sequence is mapped to the latent space, which serves as the input to the first layer Mamba Block; F is the instruction embedding matrix with shape n×d, obtained from the feature extraction stage; The learnable linear projection weight matrix has a shape of d× Mapping the feature vector from the original dimension d to the latent space dimension ; The bias term is a learnable term with dimension . n is the total number of instructions in the current basic block; =512: Latent space dimension, that is, the representation dimension of each instruction in the latent space;

[0187] The calculation process for the l-th layer (l=1,…,6) Mamba Block is as follows:

[0188] ,

[0189] ,

[0190] ,

[0191] Where SSM(·) is the S6 selective scan operator, and σ is the Sigmoid function. , The LayerNorm is a learnable projection matrix, and its value is RMSNorm. This represents the element-wise product of the SSM output and the gate signal in the l-th Mamba Block, i.e., the intermediate representation after gated operation, with a shape of n× , For the first The output of the Mamba Block layer is obtained after residual concatenation, and has a shape of n× , This represents the intermediate projection result of the l-th layer, i.e. through The representation after linear projection, the transition amount before residual connection, has a shape of n× Each block contains a selective state-space layer (S6 operator), a gated linear unit (GLU), residual connections, and an RMSnorm.

[0192] After the outputs of the Mamba Blocks in layers 3 and 6, a lightweight microarchitecture cross-arch attention (CAA) layer is inserted respectively:

[0193] ,

[0194] in, This represents the instruction implicit representation matrix obtained after the output of the l-th layer (l=3 or l=6) Mamba Block is processed by the microarchitecture cross-attention layer, and serves as the input to the next layer of Mamba Block; For a learnable query projection matrix, Mapped to attention query; For the learnable key projection matrix, Mapped to attention key; For the learnable Value projection matrix, Mapped to attention value; For the target CPU microarchitecture context vector, dimension =64, broadcast as Key and Value to each position in the sequence, injecting microarchitectural information into the encoding of each instruction; Q (Query): Query matrix, output by the l-th layer Mamba Block. through The projection yields a vector representing each instruction "questioning" the microarchitecture information; K (Key): the key matrix, derived from the microarchitecture context vector. through The projection yields an "index" representing microarchitectural information; V (Value): a value matrix derived from the microarchitectural context vector. through The projection yields the actual content representing the microarchitectural information, which is then weighted by attention weights and injected into the encoding of each instruction.

[0195] Context vector The Key and Value for attention are broadcast to each position in the sequence, enabling instruction encoding to incorporate global information about the target microarchitecture and achieve cross-platform conditional reasoning.

[0196] After global average pooling and output MLP, the scheduling priority score vector of each instruction is obtained. :

[0197] ,

[0198] in, No. The output of the layer Mamba Block, P represents an n-dimensional real vector, where n is the number of instructions in the current basic block.

[0199] Mamba achieves linear computational complexity by introducing a selective scanning mechanism based on input dependencies. It also achieves high-quality global context capture for long sequences, compared to Transformer. Its complexity offers a significant advantage, making it suitable for real-time reasoning of basic blocks containing hundreds of instructions during compile time.

[0200] The parameterization mechanism of input dependencies enables the model to adaptively and selectively retain or forget historical state information based on the instruction context, thereby effectively capturing long-link data dependencies across multiple layers.

[0201] The priority score P output by the model is decoded into a valid scheduling sequence through a strategy: maintaining a dynamic ready set, activating the priority score only for instructions whose dependencies have been met, and masking the scores of unread instructions as negative infinity; employing a pointer network mechanism to iteratively select the highest priority instruction from the ready set for each scheduling slot to ensure that the output is a valid permutation; detecting resource conflicts of functional units, and if a conflict occurs, searching for a suboptimal solution through Beam Search (beam width k=4) to avoid greedy solutions getting trapped in local optima.

[0202] In a preferred embodiment of the present invention, a dependency-constrained pointer network is used to iteratively decode each scheduling slot to generate a valid instruction scheduling arrangement that satisfies all data dependency constraints. The specific method is as follows:

[0203] In each scheduling slot s (s=1,…,n), perform the following iterations:

[0204] Check if the predecessor of each instruction in the directed acyclic graph (DAG) G has been fully scheduled, and add the instructions that meet the conditions to the dynamic ready set. :

[0205] ,

[0206] in, These are candidate instructions that need to be determined whether to be added to the ready set. for The predecessor instruction, that is, there is an edge in the DAG ( → The instructions, For edge ( , The number of dependency latency cycles on ) Describes the set of directed edges of a DAG. This represents the set of instructions that have been scheduled up to the s-th scheduling slot. When all the front-ends are in it, Only then can it be added to the ready set. ;

[0207] Calculate attention score :

[0208] ,

[0209] ,

[0210] in, Encode the Mamba output of the i-th instruction. Let GRU be the hidden state of the decoder at step s. , , These are learnable parameters;

[0211] Select the instruction with the highest score when there are no resource conflicts. :

[0212] ,

[0213] like If a functional unit conflict exists between a scheduled instruction and another instruction (the demand on the same port exceeds capacity in the same cycle), a Beam Search with a beam width of k=4 is initiated. From the ready set, k candidate paths with the highest scores are selected, and the first instruction of the path with no conflict and the highest cumulative score is chosen as the first instruction. ;

[0214] Will Join the scheduled set Update the decoder hidden state. , into the scheduling slot ;

[0215] After decoding, output a valid scheduling sequence that satisfies all data dependencies and resource constraints. , For scheduling and arrangement, Let σ be the index of the instruction corresponding to the nth scheduling slot in the original instruction sequence, i.e., the original number of the last scheduled instruction. σ is a mapping from the scheduling slot number to the original instruction index, where σ(1), σ(2), …, σ(n) form a permutation from 1 to n.

[0216] By inheriting the LLVM MachineSchedStrategy interface, Mamba-IS is injected into the compiler as a custom scheduling strategy plugin. Model inference is invoked in the scheduleDAGMutations callback, with the output priority overriding the default list scheduling priority. To meet compile-time real-time requirements (target: single basic block scheduling latency <5ms), INT8 quantization (Post-Training Quantization) and ONNX Runtime deployment are employed.

[0217] The training data generation pipeline utilizes LLVM MCA (LLVM Machine Code Analyzer; LLVM itself is a technical term without a direct Chinese translation, but it is commonly referred to as LLVM in academic and engineering fields) and hardware PMU counters to dynamically analyze the compiled code, obtain measured IPC (instructions per cycle) as a monitoring signal, and construct training pairs (original instruction sequences and optimized scheduling sequences). The dataset covers benchmark programs such as SPEC CPU 2017, LLVM Test Suite, and cBench.

[0218] In a preferred embodiment of the present invention, such as Figure 6 As shown, it also includes a two-stage training step for the Mamba selective state-space model, specifically:

[0219] In the `scheduleDAGMutations()` callback of the `MachineScheduler::schedule()` process, the DAG information is serialized into model input, the ONNX Runtime is called to complete inference, and the output priority vector P overrides the default sorting decision of `MachineSchedStrategy::pickNode()`. The entire process requires zero modification to the LLVM core code and supports enabling this method on any target platform via the compilation option `-mllvm -enable-mamba-is`.

[0220] To meet the real-time requirements during compilation, the Mamba-IS model underwent Post-Training Quantization (PTQ) with INT8. While maintaining an accuracy loss of less than 0.5%, the model weights were quantized from float32 to int8, and all inference computations were performed using int8 integer operations. Experimental results show that the inference latency for a single basic block containing 200 instructions after quantization is 2.3ms, far below the target threshold of 5ms.

[0221] Machine basic blocks were extracted from benchmark programs such as SPEC CPU 2017, LLVM Test Suite, and cBench; throughput simulations were performed on different scheduling arrangements using LLVM MCA to predict IPC as a coarse-grained quality label; the actual operation of the SPEC CPU 2017 program was sampled using a hardware performance counter (PMU) to obtain accurate IPC as a fine-grained quality label; and the training set was organized into pairs of (original unscheduled sequences and high-quality scheduled sequences), totaling approximately 1.2 million training samples.

[0222] Phase 1 (Supervised Pre-training): Using LLVM MCA simulation IPC as the ranking quality label, and employing cross-entropy loss. Supervised learning of the pointer network decoding process:

[0223] ,

[0224] in, For the expert scheduling sequence in the 1st The instruction index of each scheduling slot, where n is the total number of scheduling slots, i.e., the total number of instructions; the optimizer uses... Train for 100 epochs with a batch size of 128;

[0225] Phase Two (PPO Reinforcement Learning Fine-tuning): The pre-trained model is fine-tuned using the IPC improvement measured by the measured PMU performance counter as the reward. The reward function... Defined as:

[0226] ,

[0227] Among them, IPC_Mamba-IS( ) represents the measured IPC of the program when using the scheduling arrangement π, and IPC_ListScheduler represents the measured IPC of the LLVM default list scheduling;

[0228] The PPO (Proximal Policy Optimization) algorithm is used to prune parameters. Value function coefficients Entropy regularity coefficient The training lasted for 20 epochs; the PPO objective function was... for:

[0229] ,

[0230] in, Importance sampling ratio, For generalized advantage estimation, For the value function loss, This is the policy entropy regularization term; Importance sampling ratio measures the ratio of the probability of the old and new strategies choosing the same action in the same state. For the current policy in state Select action The probability of; For the old strategy before the update, in the state Select action The probability of θ; θ is the learnable parameter of the current policy; The parameters for the old strategy before the update; The action selected for step s, i.e. the instruction selected for the s-th scheduling slot; The state at step s is a comprehensive representation of the current ready set, the scheduled set, and the decoder's hidden state. This represents the truncation function, which will... The value is limited to the range of [1-ε, 1+ε] to prevent the difference between the old and new strategies from being too large, which would lead to training instability. ε=0.2 is the pruning parameter.

[0231] like Figure 7 As shown, experimental results demonstrate that, compared to the default list scheduling strategy of LLVM, Mamba-IS achieves significant IPC improvements on both the SPEC CPU2017 integer subset (SPECint) and the floating-point subset (SPECfp), with particularly noticeable gains in scenarios involving complex control flows and long latency chains, thus verifying the effectiveness of the method of this invention.

[0232] The present invention also provides a compiler backend instruction scheduling optimization system based on the method described in the present invention, comprising:

[0233] like Figure 2 As shown, the program representation and serialization module is used to establish the scheduling DAG, extract instruction features and complete topology-related serialization processing, while generating a microarchitecture context representation;

[0234] The model computation module is used to output instruction priority scores based on the Mamba Block stacking structure and microarchitecture cross-attention mechanism;

[0235] like Figure 5 As shown, the scheduling decoding module is used to obtain a valid scheduling sequence by combining dynamic ready set constraints and iterative decoding process;

[0236] The LLVM integrated feedback module is used to apply the scheduling results to the LLVM backend via the MachineSchedStrategy interface and supports the training data generation process.

[0237] In a preferred embodiment of the present invention, the program characterization and serialization module includes a microarchitecture context injection unit. This unit reads parameters such as the target CPU's issue width, number of functional units, and delay table from the LLVM ScheduleModel to construct a microarchitecture descriptor vector. .

[0238] The specific embodiments described herein are merely illustrative examples of the present invention. Those skilled in the art can make various modifications or additions to the described embodiments or use similar methods to substitute them, without departing from the technology of the present invention or exceeding the scope defined by the appended claims.

[0239] In the embodiments of this application, terms such as "fixed," "fixed connection," and "fixed connection" refer to common fixing methods in the prior art, such as welding, riveting, and screws. "Rotary connection" refers to common rotary connection methods in the prior art, such as hinges and bearing rotation. If electrical components are provided, the functions, control, and power supply methods of all electrical components are common technical means in the prior art. This application has not improved them and they are not within the protection scope of this application. Therefore, this application will not elaborate on them.

[0240] Furthermore, the selection of materials and strength limitations for all components in this application can be made and arranged by those skilled in the art based on the site environment and the requirements of relevant national or industry standards, and are not within the scope of protection of this application. Therefore, this application will not elaborate on these points.

Claims

1. A compiler backend instruction scheduling optimization method based on the Mamba state-space model, characterized in that, Includes the following steps: Obtain the instruction sequence from the compiler backend machine basic block, analyze the dependencies between instructions, and construct a scheduling directed acyclic graph; The scheduling directed acyclic graph is topology-aware serialization is performed to extract the multi-dimensional feature vector and instruction embedding matrix of each instruction, and the target CPU microarchitecture descriptor is encoded into a global context vector. The instruction embedding matrix and global context vector are input into the Mamba selective state space model to complete global sequence modeling in linear time complexity and output the scheduling priority score of each instruction. By using a dependency-constraint-aware pointer network, iterative decoding is performed in each scheduling slot to generate a valid instruction scheduling arrangement that satisfies all data dependency constraints. The generated legal instruction scheduling arrangement is fed back to the compiler backend, and the priority selection process in the original list scheduling is adjusted accordingly to achieve instruction scheduling optimization.

2. The compiler backend instruction scheduling optimization method based on the Mamba state space model according to claim 1, characterized in that, The method for obtaining the instruction sequence from the compiler's backend machine basic blocks, analyzing the dependencies between instructions, and constructing a directed acyclic graph of the schedule is as follows: For the n MachineInstr instruction sequences in the LLVM MachineBasicBlock The analysis examines register read / write relationships and memory alias relationships between instruction sequences, including: true dependencies (RAW: read after write), anti-dependencies (WAR: write after read), output dependencies (WAW: write after write), and memory alias dependencies. Construct a Directed Acyclic Graph (DAG) for scheduling : DAY , Among them, vertex set Each edge in the directed edge set E corresponds one-to-one with the instruction. Indication of instructions right There is a delay. Dependence on each cycle.

3. The compiler backend instruction scheduling optimization method based on the Mamba state-space model according to claim 1, characterized in that, The steps for performing topology-aware serialization on the scheduled directed acyclic graph and extracting the multidimensional feature vector and instruction embedding matrix for each instruction are as follows: Using the longest path algorithm for directed acyclic graphs (DAGs), the critical path depth (CPD) of each instruction sequence is calculated: , in, It is the edge The number of dependency delay cycles, if If there is no successor node (out-degree is 0), then CPD( If ) = 0; There is a successor node, traversal All successor nodes Take the maximum value of "edge delay plus successor node CPD value" as the value. The critical path depth (CPD) can be calculated sequentially in reverse topological order (from the endpoint to the starting point) to obtain the CPD value for each instruction. Extract multidimensional feature vectors from each MachineInstr instruction sequence. eigenvectors This includes opcode embedding, operand type encoding, one-hot encoding of target functional units, normalized critical path depth (CPD) value, dependency graph structure context, ready time slots, and latency sensitivity coefficients, specifically: Opcode embedding: via a learnable embedding table Map machine opcodes to dense vectors, dimension ; Operand type encoding: One-hot concatenation of register type (integer / floating-point / vector), immediate flag, and memory operand flag; dimension ; One-hot encoding of the target functional unit refers to the encoding of the execution port required by the identifier instruction, and its dimensions are... Equal to the number of functional unit types of the target platform; Critical path depth normalized value The range is [0,1]; where CPD( ) represents the critical path depth of the k-th instruction, which is the longest weighted path length from this instruction to the endpoint of the DAG; It is the kth instruction, that is, the kth MachineInstr node in the instruction sequence; This represents the maximum CPD of all instructions in the current basic block, which is the total length of the critical path in the entire DAG; Dependency Graph Structure Context: Multi-hot encoding of the predecessor node set from the edge set E of the constructed scheduling DAG Multi-hot coding of successor node sets Binary markers for determining whether a path is on the critical path Specifically: judge Whether it is on the critical path is determined by the condition: IsCPD ( )=1, if CPD( ) + From the root node to Longest path = If an instruction satisfies both "longest path to destination" and "longest path from origin", it means that it is on the global critical path; the result is a binary label, 0 or 1, which is directly used as a scalar component and concatenated into the feature vector. Readiness time slot estimation: based on the earliest launchable period of the current partial schedule. Normalized to [0,1]; Delay sensitivity coefficient: When constructing the scheduling DAG, the latency sensitivity coefficient for each edge The latency value lat is read and marked from the instruction latency table (InstrItineraries / SchedReadWrite) of the LLVMScheduleModel, and then iterated through. Take the maximum delay value (lat_max) from all outgoing edges: , , For leaf nodes with an out-degree of 0 or nodes with a CPD of 0, LS( ) = 0, LS( ) is an instruction The delay sensitivity coefficient; Concatenate all components to obtain the original feature vector of each instruction sequence. ,in , Assign opcode embedding dimension; form instruction embedding matrix d is the total dimension of the feature vector, which is equal to the sum of the dimensions of each component. The dimension for embedding the opcode is set to 128; The dimension for encoding operand types is set to 16. The dimension of the one-hot encoding for the target functional unit is equal to the number of functional unit types on the target platform.

4. The compiler backend instruction scheduling optimization method based on the Mamba state space model according to claim 3, characterized in that, The method for encoding the target CPU microarchitecture descriptor into a global context vector is as follows: Perform an improved topological sort on the DAG to generate candidate sequences. The DAG structural context information (predecessor set, successor set, critical path marker) of each node is concatenated into a multi-dimensional feature vector to form a structure-aware input embedding matrix. ,in , For structural context dimension; Microarchitectural parameters of the target CPU are extracted from the LLVM target description file (ScheduleModel) to form a microarchitectural feature vector a, including: superscalar issue width (IssueWidth), number of each functional unit type, etc. Delay table entries for each functional unit The out-of-order execution window size (ROBSize), rearrangement buffer depth, and L1 instruction cache line size are normalized and then concatenated to form the original microarchitecture vector. Then, it is projected as a context vector through a fully connected network. : , in, LayerNorm represents the learnable parameters and is used for layer normalization. By injecting cross-attention mechanism into the Mamba-IS model, conditional reasoning for different target microarchitectures can be achieved.

5. The compiler backend instruction scheduling optimization method based on the Mamba state-space model according to claim 1, characterized in that, The steps for inputting the instruction embedding matrix and global context vector into the Mamba selective state-space model to complete global sequence modeling in linear time complexity and output the scheduling priority score for each instruction are as follows: The continuous-time state-space equation of the S6 operator in the Mamba selective state-space model is: , , in, It is a learnable diagonal matrix. = 512 is the latent space dimension of the model, corresponding to the feature dimension of each position in the input sequence; = 64 represents the hidden state dimension, i.e., the hidden variables in the state-space model. The dimensions; matrices B(·) and C(·) are both derived from the current input. The linear projection is dynamically generated, where D is the scalar jump connectivity coefficient. The instruction embedding at time t is x; This is the derivative of the hidden state with respect to time, i.e., the rate of change of the hidden state over continuous time. This represents the output vector at time t, corresponding to the encoded representation of the instruction at that time. This represents the hidden state vector at time t, with dimension . =64, carrying contextual information about the historical sequence; After zero-order hold (ZOH) discretization, the step size Δ is determined by the input: ), thus obtaining the discrete recurrence form: , , , , Wherein, the state transition matrix Input matrix and output matrix All are from the current input Dynamically generated; The state transition matrix is ​​the discretized state transition matrix, which is the input when processing the τth instruction. Dynamically generated, controlling the hidden state The degree of preservation at the current time step; The input matrix is ​​discretized from the current input. Dynamically generated, controlling the current input The degree of injection into the hidden state; The output matrix is ​​derived from the current input. Dynamically generated, controlling the hidden state How to map to output , Let be the hidden state vector at the τth instruction, with dimension . =64, carrying contextual information about the historical sequence; This is the output vector at the τ-th instruction, corresponding to the encoded representation of the current instruction; The discretization step size controls the granularity of discretization of the continuous-time state-space equations, determined by the current input. Dynamically generated, each instruction corresponds to a different step size, which is always positive through the softplus function; softplus is the activation function, softplus(x)=log(1+eˣ), ensuring that Δ is always positive; The input is a learnable linear projection matrix. Mapped to the pre-activation value of the step size; Let τ be the input vector at the τ-th instruction, which is the embedded representation of the current instruction; For learnable bias terms; The selective mechanism enables the Mamba selective state space model to adaptively and selectively retain or forget scheduling history information in the preceding hidden state according to the current instruction context, effectively modeling long-link data dependencies. The Mamba selective state space consists of an input projection layer and L=6 stacked Mamba Blocks. The input projection layer linearly maps the instruction embedding matrix F to the latent space. , in, The output of the input projection layer is the initial latent representation matrix after the instruction sequence is mapped to the latent space, which serves as the input to the first layer Mamba Block; F is the instruction embedding matrix with shape n×d, obtained from the feature extraction stage; The learnable linear projection weight matrix has a shape of d× Mapping the feature vector from the original dimension d to the latent space dimension ; The bias term is a learnable term with dimension 1. n is the total number of instructions in the current basic block; =512: Latent space dimension, that is, the representation dimension of each instruction in the latent space; The calculation process for the l-th layer (l=1,…,6) Mamba Block is as follows: , , , Where SSM(·) is the S6 selective scan operator, and σ is the Sigmoid function. , The LayerNorm is a learnable projection matrix, and its value is RMSNorm. This represents the element-wise product of the SSM output and the gate signal in the l-th Mamba Block, i.e., the intermediate representation after gated operation, with a shape of n× , For the first The output of the Mamba Block layer is obtained after residual connection, and has a shape of n× , This represents the intermediate projection result of the l-th layer, i.e. through The representation after linear projection, the transition amount before residual connection, has a shape of n× ; After the outputs of the Mamba Blocks in layers 3 and 6, a lightweight microarchitecture cross-attention layer (CAA) is inserted respectively: , in, This represents the instruction implicit representation matrix obtained after the output of the l-th layer (l=3 or l=6) Mamba Block is processed by the microarchitecture cross-attention layer, and serves as the input to the next layer of Mamba Block; For a learnable query projection matrix, Mapped to attention query; For the learnable key projection matrix, Mapped to attention key; For the learnable Value projection matrix, Mapped to attention value; For the target CPU microarchitecture context vector, dimension =64, broadcast as Key and Value to each position in the sequence, injecting microarchitectural information into the encoding of each instruction; Q (Query): Query matrix, output by the l-th layer Mamba Block. through The projection yields a vector representing each instruction "questioning" the microarchitecture information; K (Key): the key matrix, derived from the microarchitecture context vector. through The projection yields an "index" representing microarchitectural information; V (Value): a value matrix derived from the microarchitectural context vector. through The projection yields the actual content representing the microarchitecture information, which is then weighted by attention weights and injected into the encoding of each instruction. After global average pooling and output MLP, the scheduling priority score vector of each instruction is obtained. : , in, No. The output of the layer Mamba Block, P represents an n-dimensional real vector, where n is the number of instructions in the current basic block.

6. The compiler backend instruction scheduling optimization method based on the Mamba state space model according to claim 1, characterized in that, By using a dependency-constraint-aware pointer network, iterative decoding is performed at each scheduling slot to generate a valid instruction scheduling permutation that satisfies all data dependency constraints. The specific method is as follows: Check if the predecessor of each instruction in the DAG G has been scheduled, and add the instructions that meet the conditions to the dynamic ready set. : , in, These are candidate instructions that need to be determined whether to be added to the ready set. for The predecessor instruction, that is, there is an edge in the DAG ( → The instructions, For the edge ( , The number of dependency latency cycles on ) Describes the set of directed edges of a DAG. This represents the set of instructions that have been scheduled up to the s-th scheduling slot. When all the front-ends are in it, Only then can it be added to the ready set. ; Calculate attention score : , , in, Encode the Mamba output of the i-th instruction. Let GRU be the hidden state of the decoder at step s. , , These are learnable parameters; Select the instruction with the highest score when there are no resource conflicts. : , like If a functional unit conflict exists between a scheduled instruction and another instruction (the demand on the same port exceeds capacity in the same cycle), a Beam Search with a beam width of k=4 is initiated. From the ready set, k candidate paths with the highest scores are selected, and the first instruction of the path with no conflict and the highest cumulative score is chosen as the first instruction. ; Will Join the scheduled set Update the decoder hidden state. , into the scheduling slot ; After decoding, output a valid scheduling sequence that satisfies all data dependencies and resource constraints. , For scheduling and arrangement, This is the index of the instruction corresponding to the nth scheduling slot in the original instruction sequence, i.e., the original number of the last scheduled instruction.

7. The compiler backend instruction scheduling optimization method based on the Mamba state space model according to claim 1, characterized in that, It also includes a two-stage training step for the Mamba selective state-space model, specifically: Phase 1 (Supervised Pre-training): Using the LLVM machine code analyzer simulation IPC (instructions per cycle) as the ranking quality label, cross-entropy loss is employed. Supervised learning of the pointer network decoding process: , in, For the expert scheduling sequence in the 1st The instruction index of each scheduling slot, where n is the total number of scheduling slots, i.e., the total number of instructions; the optimizer uses Train for 100 epochs with a batch size of 128; Phase Two (PPO Reinforcement Learning Fine-tuning): The pre-trained model is fine-tuned using the IPC improvement measured by the measured PMU performance counter as the reward. The reward function... Defined as: , Among them, IPC_Mamba-IS( ) represents the measured IPC of the program when using the scheduling arrangement π, and IPC_ListScheduler represents the measured IPC of the LLVM default list scheduling; The PPO (Proximity Policy Optimization) algorithm is used to prune parameters. Value function coefficients Entropy regularity coefficient The training lasted for 20 epochs; the PPO objective function was... for: , in, The importance sampling ratio, For generalized advantage estimation, For the value function loss, This is the policy entropy regularization term; Importance sampling ratio measures the ratio of the probability of the old and new strategies choosing the same action in the same state. For the current policy in state Select action The probability of; For the old strategy before the update, in the state Select action The probability of θ; θ is the learnable parameter of the current policy; The parameters for the old strategy before the update; The action selected for step s, i.e. the instruction selected for the s-th scheduling slot; The state at step s is a comprehensive representation of the current ready set, the scheduled set, and the decoder's hidden state. This represents the truncation function, which will... The value is limited to the range of [1-ε, 1+ε] to prevent the difference between the old and new strategies from being too large, which would lead to training instability. ε=0.2 is the pruning parameter.

8. A compiler back-end instruction scheduling optimization system based on the method of any one of claims 1-7, characterized in that, include: The program representation and serialization module is used to build the scheduling DAG, extract instruction features and complete topology-related serialization processing, while generating a microarchitecture context representation; The model computation module is used to output instruction priority scores based on the Mamba Block stacking structure and microarchitecture cross-attention mechanism; The scheduling decoding module is used to combine dynamic ready set constraints and iterative decoding process to obtain a valid scheduling sequence; The LLVM integrated feedback module is used to apply the scheduling results to the LLVM backend via the MachineSchedStrategy interface and supports the training data generation process.

9. The compiler backend instruction scheduling optimization system according to claim 8, characterized in that, The program characterization and serialization module includes a microarchitecture context injection unit, which reads parameters such as the target CPU's launch width, number of functional units, and delay table from the LLVMScheduleModel to construct a microarchitecture descriptor vector. .