FPGA accelerator for state space model reasoning and calculation method thereof

By designing an FPGA accelerator for state space model inference, using parallel computing, pipeline design and quantitative acceleration technology, the problems of low computing efficiency, high resource occupation and large delay in Mamba model inference are solved, and efficient and low power consumption calculation effects are achieved.

CN120146185APending Publication Date: 2025-06-13PEKING UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510208388.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing Mamba model inference has problems such as low computing efficiency, high resource occupancy and large latency, especially the data dependence between the state space model and the input projection layer, resulting in low hardware utilization, large resource occupancy and high computational latency.

Method used

Design an FPGA accelerator for state space model inference, including ARM CPU, HBM/LPDDR module, DMA Read and Write module, Rmsnorm module, matrix multiplication unit (MMU), state space model unit (SSMU), and Hadamar Transform Unit (HTU). Through the collaborative work of these modules, parallel computing, pipeline design and quantization acceleration are achieved, reducing data dependence and calculation delay.

Benefits of technology

The computing efficiency has been significantly improved, the hardware utilization rate has increased from 58% to 96%, the computing throughput has increased by 40%, the overall computing time has been reduced by 32%, the resource occupation has been greatly reduced, and the latency has been reduced by 32%, which meets the real-time requirements and has good scalability and versatility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146185A_ABST
    Figure CN120146185A_ABST
Patent Text Reader

Abstract

The invention discloses an FPGA accelerator for state space model reasoning and a calculation method of the FPGA accelerator, and belongs to the field of artificial intelligence hardware acceleration. According to the method, the parallel computing capability and flexibility of the FPGA are deeply fused, the limitation of a traditional hardware architecture is broken through, and efficient acceleration for a large-scale deep learning task is achieved. Compared with a traditional accelerator, the hardware utilization rate is increased from 58% to 96%, the calculation throughput is remarkably improved, the real-time requirement is met, and the overall calculation time is further shortened through assembly line design and calculation reordering. The method can be adapted to Mama models of different scales, and has good expansibility and universality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence hardware acceleration, and particularly relates to an FPGA accelerator for state space model inference and its calculation method, which is especially suitable for low-latency and high-throughput inference scenarios of Mamba-like models. Background Art

[0002] In recent years, state space models (such as Mamba) have become an important alternative to the Transformer architecture due to their linear computational complexity characteristics. Compared with traditional large language models (LLMs), Mamba significantly reduces the computational complexity when processing long sequence tasks and shows superior performance in downstream tasks. However, the existing Mamba model inference mainly relies on the CPU / GPU platform, and there are the following bottlenecks:

[0003] 1. Low computational efficiency: There is a data dependency between the SSM and the input projection layer, resulting in sequential execution of the calculations and low hardware utilization.

[0004] 2. High resource occupancy: SSM calculations require storing a large number of intermediate activation values, occupying a large amount of on-chip storage resources.

[0005] 3. High latency: The Hadamard transform can assist the Mamba network for quantization, but the traditional Hadamard transform is implemented based on matrix multiplication, resulting in a high computational latency.

[0006] Although FPGA has significant potential in the field of hardware acceleration, the research on FPGA acceleration for SSM is still blank. Therefore, there is an urgent need for an efficient and low-power FPGA acceleration architecture to break through the above technical limitations. Summary of the Invention

[0007] The purpose of the present invention is to provide an FPGA accelerator for state space model inference and its calculation method to solve the problems of low computational efficiency, high resource occupancy, and large latency in the prior art.

[0008] The technical solution of the present invention is as follows:

[0009] An FPGA accelerator for state space model inference, characterized in that the following modules are provided on the FPGA accelerator: an ARM CPU, an HBM / LPDDR module, a DMA Read module, a DMA Read&Write module, an Rmsnorm module, an MMU module for performing all linear layer operations in the Mamba model, an SSMU module for performing SSM operations in the Mamba model, and an HTU module for performing hadamard transform operations. Among them, the ARM CPU communicates with the DMA Read module, the DMA Read&Write module, the Rmsnorm module, and the HTU module through the AXI-Lite bus, and is responsible for global scheduling, parameter configuration, and task start / stop control; all weight parameters of the Mamba model are efficiently stored in the HBM / LPDDR module; the HBM / LPDDR module is connected to the DMA Read module and the DMA Read&Write module through the high-bandwidth AXI bus to achieve fast access to the model weight parameters and the hidden state (h t );The DMA Read module sequentially reads the weight parameters (W in , W conv , W ouy ) of each layer from the HBM / LPDDR module and transmits them to the MMU module and the SSMU module through a dedicated data channel; the DMA Read&Write module is responsible for two-way transmission: reading the initial h t from the HBM / LPDDR module to the on-chip cache of the SSMU module and writing the updated one back to the HBM / LPDDR module; the MMU module and the SSMU module implement data interaction through the AXI protocol: the intermediate result output by the SSMU module is transmitted to the HTU module; after the HTU module completes the hadamard transform and quantization, the quantization result is transmitted back to the HBM / LPDDR module through the DMA channel; the Rmsnorm module is integrated at the data input / output end to normalize the input sequence and the final inference result, and is directly connected to the MMU module through an independent data path.

[0010] Furthermore, the matrix multiplication unit (MMU) adopts a tree-shaped multiplier-accumulator (MAC) structure, supporting parallel computing for the input / output projection layer. Among them, the matrix multiplication of the input projection layer (din×dout) is decomposed into multi-vector multiplications. Each vector multiplication is processed by a tree-shaped accumulator, which is composed of DSP units and adopts a hierarchical design: the parallel DSP units at the bottom layer perform multiply-accumulate operations on two input vectors; the parallel DSP units in the middle layer are local accumulators, summarizing partial sums; the top-layer DSP unit is a global accumulator, generating the final result and outputting the result of the multiply-accumulate of two vectors. In the MMU module of the present invention, multiple tree-shaped accumulators can be parallelized to achieve parallel computing of matrix multiplication. Through DSP packing technology, the operations are mapped to DSP units, improving resource utilization; the time multiplexing mechanism realizes alternating scheduling of linear layers, and the computing throughput is increased by 40%.

[0011] Furthermore, the state space model unit (SSMU) divides the input data into three parts, namely Δ, BCX, and Z, through a multiplexer. Δ is sent to the bias accumulation module through the AXI interface, and after accumulation, it is sent to the Softplus module through the AXI interface; BCX is sent to the Conv (convolution) module through the AXI interface; after the Conv module and Z pass through the multiplexer, they are sent to the Silu module through the AXI interface; the Silu module will divide the data into four parts, namely B, C, X, and Z, through the multiplexer. B and C are stored in the Buffer for subsequent reading; the Softplus module will send the Δ data to two EMU modules respectively, and perform multiplication operations with data A and data B respectively; after multiplying with A, it generates and sends it to an EMU module; after B and Δ pass through the EMU module, they get and X perform a multiplication operation through the EMU module to get and h t-1 After performing a multiplication operation through the EMU module, it is accumulated with to get h t ; h t One part is sent to the off-chip storage space for storage, and one part is sent to another EMU module to perform matrix multiplication with C and obtain h t C; another part of X will be sent to the EMU module to perform a multiplication operation with the off-chip stored D, and then accumulate with h t C to obtain Y, and then send it to the EMU module to perform a multiplication operation with Z to obtain the output sequence. In this way, the SSMU module can flexibly adjust the allocation of computing resources when processing complex computing tasks, thereby achieving maximized computing efficiency. The overall delay is reduced by 32%; the SSM operation operator is implemented using the element multiplication unit (EMU), and different parallel degrees are set for different operators to optimize the data flow balance.

[0012] Furthermore, the Hadamard Transform Unit (HTU) supports customized designs for rotation-assisted quantization, including a 128-point HTU based on the Fast Hadamard Transform (FHT) algorithm and a 40-point HTU based on a simple Matrix Multiplication Unit (MMU); the 128-point HTU adopts a multi-stage pipelined design with a Butterfly Core and a FIFO buffer, significantly reducing the computational latency; the 40-point HTU simplifies the computational logic and reduces the hardware resource occupancy by fixing the input to a Hadamard matrix (containing only 1 and -1).

[0013] A computational method for an FPGA accelerator for state space model inference, characterized in that its steps include:

[0014] 1) Model loading and initialization, specifically including:

[0015] 1-1) Store the weight parameters of the Mamba model in the HBM / LPDDR module of the FPGA accelerator layer by layer. Each layer of weights includes the input projection layer (W in ), SSM parameters (Wconv), and the output projection layer (W out );

[0016] 1-2) Load the initial hidden state (h t ) from the HBM module to the on-chip cache and dynamically update the output results of each layer;

[0017] 1-3) Quantize the SSM parameters of each layer and configure the quantization parameters to the HTU module;

[0018] 2) Input projection layer calculation, specifically including:

[0019] 2-1) After the input sequence (X in ) is normalized, it is transmitted to the MMU module;

[0020] 2-2) The MMU module alternately processes the input projection layer and the output projection layer calculations, where the input projection layer is preferentially executed, and the output Δ, B, and C parameters are stored in the on-chip double buffer;

[0021] 3) SSM inference pipeline execution, specifically including:

[0022] 3-1) The SSMU generates intermediate data in the order of Δ→B→C→X→Z to eliminate inter-layer dependencies;

[0023] 3-2) An Element Multiplication Unit (EMU), a State Update Unit (SUU), and an Activation Function Unit (AFU) are deployed inside the SSMU, and each unit is connected through a FIFO buffer;

[0024] 3-3) After the SSM calculation of each layer is completed, the updated ht Write back to the HBM module through the DMA module for the next layer to call;

[0025] 4) Hadamard transform and quantization acceleration, specifically including:

[0026] 4-1) Push the input data in batches to the input FIFO and complete the transform through the Butterfly Core operation;

[0027] 4-2) Quantize the data after the transform through the rotation-assisted PTQ algorithm, and dynamically adjust the quantization parameters through shift operations;

[0028] 4-3) The quantized result is for the output projection layer to call;

[0029] 5) Output projection layer and result generation, specifically including:

[0030] 5-1) The MMU switches to the output projection mode, loads the weights from the HBM module, and performs the final matrix multiplication;

[0031] 5-2) The inference result is normalized and output, and at the same time the global hidden state h is updated t , completing the single-layer calculation.

[0032] The beneficial effects of the present invention are as follows:

[0033] 1. The calculation efficiency is significantly improved:

[0034] The hardware utilization rate of the present invention is increased from 58% to 96%, significantly improving the calculation throughput; the pipeline design of the SSM layer reduces data dependence, and the overall calculation time is reduced by 32%.

[0035] 2. The resource occupation is greatly reduced:

[0036] The on-chip storage resource occupation of the SSMU of the present invention is reduced by 4 times; the optimized design of the Hadamard transform unit (HTU) reduces the calculation delay by 72%, while the hardware resource occupation remains unchanged.

[0037] 3. The delay is significantly reduced:

[0038] The present invention significantly reduces the calculation delay and meets the real-time requirements; through the pipeline design and calculation reordering, the overall calculation time is further reduced.

[0039] 4. Strong scalability:

[0040] The present invention can adapt to different scales of Mamba models and has good scalability and versatility. Description of the Drawings

[0041] Figure 1: Overall architecture diagram of how the Mamba2-2.7B model is deployed on the accelerator in the specific embodiments of the present invention, showing the interconnection relationship of the CPU, HBM, DMA, and core computing modules;

[0042] Figure 2 : Tree-shaped MAC structure of the matrix multiplication unit (MMU) and DSP packaging implementation in the specific embodiments of the present invention;

[0043] Figure 3 : Pipeline architecture of the SSMU and EMU parallelism configuration in the specific embodiments of the present invention;

[0044] Figure 4 : Butterfly core pipeline design of the HTU and simplified matrix multiplication implementation in the specific embodiments of the present invention. Specific implementation manners

[0045] Combined with the accompanying drawings below, the Mamba2-2.7B model is deployed on the accelerator to further elaborate the present invention.

[0046] The Mamba2-2.7B model has a total of 64 layers, and each layer has the same Mamba block structure, but the weight parameters of each layer are different. At the same time, the hidden state (ht) of the mamba model is passed between each layer. Each layer updates ht according to the weight parameters and input of the current layer.

[0047] The FPGA accelerator architecture described in the present invention is based on a partially expanded spatial architecture design. The single-layer Mamba block structure is precisely expanded on the FPGA platform, and the model output is obtained by circularly loading the weight parameters of 64 layers. This architecture deeply integrates the parallel computing ability and flexibility of the FPGA, breaks the limitations of traditional hardware architectures, and realizes efficient acceleration for large-scale deep learning tasks. Compared with traditional accelerators, this design significantly improves the utilization rate of hardware resources and achieves refined optimization in terms of operation efficiency, power consumption management, and hardware resource allocation.

[0048] The modules of this architecture are as follows: 1. The CPU module, which is an ARM CPU on the FPGA and controls the startup and termination of the accelerator; 2. The HBM / LPDDR (High Bandwidth Memory / Low Power DDR) module, which stores the weight parameters of 64 layers of the Mamba model and ht; 3. The DMARead module, which controls the accelerator to sequentially read the model weights of 64 layers from the HBM / LPDDR; 4. The DMA Read&Write module, which controls the accelerator to read and write back ht from the HBM / LPDDR; 5. The Rmsnorm module, which performs the rmsnorm normalization operation in the Mamba model; 6. The Matrix Multiplication Unit (MMU) module, which performs all linear layer operations in the Mamba model; 7. The SSM Unit (SSMU) module, which performs the SSM operation in the Mamba model; 9. The Hadamard Transform Unit (HTU) module, which supports the rotation-assisted quantization algorithm in the Mamba model and performs the hadamard transform operation.

[0049] As Figure 1 shown, the FPGA accelerator of the present invention adopts a hierarchical architecture design, and each module is interconnected through a high-speed data bus and a control signal. The specific connection relationship is as follows:

[0050] 1. The ARM CPU serves as the main controller and communicates with the DMA Read module, the DMA Read&Write module, the Rmsnorm module, and the HTU module through the AXI-Lite bus, and is responsible for global scheduling, parameter configuration, and task startup and stop control. Specifically, the ARM CPU is responsible for starting the accelerator, scheduling the data of each layer, transmitting the input data of each layer to the Rmsnorm module for processing, adding the projection layer data output by the MMU of each layer to the input data of that layer to generate the output data of that layer, and then transmitting it back to the CPU as the input data of the next layer.

[0051] 2. The HBM / LPDDR module is connected to the DMA Read module and the DMA Read&Write module through a high-bandwidth AXI bus to achieve fast access to the model weight parameters and the hidden state (ht).

[0052] 3. The DMA Read module sequentially reads the weight parameters (W in , W conv , W out ) of each layer from the HBM and transmits them to the MMU module and the SSMU module through a dedicated data channel.

[0053] 4. The DMA Read&Write module is responsible for two-way transmission: reading the initial ht from the HBM to the on-chip cache of the SSMU module and writing back the updated one to the HBM.

[0054] 5. The MMU module and the SSMU module implement data exchange through the AXI protocol: the Δ, B, and C parameters generated by the MMU in the input projection layer are transmitted to the SSMU for consumption by the SSMU pipeline; the intermediate results output by the SSMU are transmitted to the HTU module.

[0055] 6. The HTU module receives the intermediate data from the SSMU, completes the Hadamard transform and quantization, and then transmits the quantization result back to the HBM through the DMA channel for the output projection layer of the MMU module to call.

[0056] 7. The Rmsnorm module is integrated into the data input / output end to normalize the input sequence and the final inference result, and is directly connected to the MMU module through an independent data path.

[0057] All weight parameters of the Mamba model are efficiently stored in HBM / LPDDR, ensuring fast access to large-scale model parameters. This storage solution not only ensures high-speed calculations, but also minimizes the storage requirements on the FPGA chip, avoiding excessive consumption of FPGA internal storage resources, so that more computing tasks can be focused. This architecture consists of three core modules, each of which carries specific computing tasks and can fully utilize its hardware potential to ultimately achieve a global acceleration effect.

[0058] A. Matrix Multiplication Unit (MMU)

[0059] The matrix multiplication unit (MMU) is a key component in this architecture, responsible for executing all linear layer calculation tasks in the Mamba model, especially the matrix multiplication operations of the input projection layer and the output projection layer. In order to maximize the computing power, the MMU adopts a highly optimized time multiplexing structure, which supports the alternating scheduling of the input projection and output projection calculations, thereby reducing the idle time of computing resources and improving the computing throughput of the FPGA.

[0060] In terms of design, the MMU uses a tree accumulator structure to cleverly parallelize multiple accumulation processes in matrix multiplication, so that each accumulation operation can be completed in the shortest time. This design not only reduces redundant steps in the calculation process, but also improves memory access efficiency. In order to further improve the utilization of hardware resources, the MMU uses DSP packing technology, so that each DSP unit can perform two multiplication and accumulation operations at the same time, thereby achieving a high degree of computational parallelism. This parallel computing method enables the MMU to efficiently process large-scale matrix operations on the FPGA, which is particularly suitable for large-scale data streams and matrix computing requirements in deep neural networks.

[0061] Through these optimized designs, the MMU can significantly improve the utilization rate of FPGA hardware resources, increasing it from less than 60% in traditional designs to nearly 96%. This efficient use of resources not only improves computing performance but also better controls the power consumption of the chip.

[0062] Figure 2 Details of the tree-shaped multiplier-accumulator (MAC) structure and DSP optimization strategy of the MMU are shown as follows:

[0063] 1. Tree-shaped MAC architecture

[0064] The matrix multiplication of the input projection layer (din×dout) is decomposed into multi-vector multiplications. Each vector multiplication is processed by a tree-shaped accumulator, which is composed of DSP units and adopts a hierarchical design: the parallel DSP units at the bottom layer perform multiply-accumulate operations on two input vectors; the parallel DSP units in the middle layer are local accumulators that sum up partial sums; the top-layer DSP unit is the global accumulator that generates the final result and outputs the result of the multiply-accumulate of two vectors. Multiple tree-shaped accumulator units can be parallelized to achieve parallel computing of matrix multiplication.

[0065] 2. DSP packing technology

[0066] Through the packing technology, each DSP unit supports the multiply-accumulate operation of two 8-bit fixed-point numbers in a single cycle for data with the same input activation value (as shown in Figure 2 : A×a + A×b), reducing the consumption of DSP units by half. As shown in Figure 2 : For matrix-vector multiplication, the input activation value is a vector with a dimension of 1×d in and the input weight value is a matrix with a dimension of d in ×d out . If the DSP packing technology is not used, the number of parallel DSP units at the bottom layer of the tree-shaped multiply-accumulator should be d in ×d out , but with the DSP packing technology, the actual number of DSPs consumed is reduced to .

[0067] 3. Time multiplexing mechanism

[0068] The input projection layer and the output projection layer are alternately scheduled. The matrix multiplication operations of the two layers share the same MMU module. When the input projection layer is computing, the weights of the output projection layer are pre-loaded into the DSP cache, reducing idle cycles and increasing hardware utilization.

[0069] B. SSM unit (SSMU)

[0070] The SSM unit (SSMU) is specifically designed for the SSM layer in the Mamba model. It adopts a fully pipelined architecture to optimize the computing process. The main challenge of the SSM layer is the high-frequency external memory access of intermediate activation values, which often becomes a bottleneck and affects the computing efficiency. To solve this problem, dedicated hardware units are used in the SSMU design to execute all operators, and data is cached and transferred through FIFO buffers, thus avoiding frequent memory access and greatly improving the processing speed.

[0071] The design of the SSMU module focuses on optimizing parallelism. Each operator adopts a different parallelism configuration according to its computing characteristics to ensure the balanced distribution of the computing load on the hardware. At the same time, the pipelined design of the SSMU module reduces the dependencies between stages, enabling each operator to execute independently and complete the calculation in the shortest time. Through this optimization, the SSMU module can maximize the utilization of FPGA computing resources while ensuring high throughput, ensuring the efficient execution of the SSM layer in the Mamba model.

[0072] To further accelerate the calculation, the SSMU also adopts element-wise multiplication units (EMUs). These units are built based on the digital signal processor (DSP) units in the FPGA and support high-parallel element-level calculations. Each EMU adjusts the parallelism according to needs to adapt to the computing complexity and data flow characteristics of different operators. In this way, the SSMU module can flexibly adjust the allocation of computing resources when processing complex computing tasks, thereby achieving maximum computing efficiency.

[0073] The efficient design of the SSMU significantly reduces the on-chip storage requirements, reducing the storage occupancy to nearly 1%, while maintaining support for large-scale data parallel computing, ensuring high throughput of the system.

[0074] As Figure 3 shown, the SSMU divides the input data into three parts, Δ, BCX, and Z, through a multiplexer. Δ is sent to the bias accumulation module through the AXI interface, and after accumulation, it is sent to the Softplus module through the AXI interface. BCX is sent to the Conv (convolution) module through the AXI interface. After passing through the multiplexer, the Conv module and Z are sent to the Silu module through the AXI interface. The Silu module divides the data into four parts, B, C, X, and Z, through a multiplexer. B and C are stored in the buffer for subsequent reading. The Softplus module sends the Δ data to two 1×8 EMU modules respectively, and multiplies them with data A and data B respectively. After multiplying with A, it generates and is sent to a 2×8 EMU module. After passing through the 1×8 EMU module, B and Δ get The multiplication operation between X and [something] is performed through a 2×8 EMU module to obtain and h t-1 After performing the multiplication operation through a 2×8 EMU module, it is accumulated with to obtain h t 。h t Part of it is sent to off-chip storage space for storage, and part of it is sent to another 2×8 EMU module to perform matrix multiplication with C and obtain h through an accumulator t C. Another part of X will be sent to a 1×8 EMU module to perform a multiplication operation with D stored outside the chip, and then accumulated with h t C to obtain Y, and then sent to the last 1×8 EMU module to perform a multiplication operation with Z to obtain the output sequence.

[0075] C. Hadamard Transform Unit (HTU)

[0076] The Hadamard Transform Unit (HTU) is a customized hardware unit designed to support the complex rotation-assisted quantization algorithm in the Mamba model. There are two types of Hadamard transforms involved in the Mamba model: 128-point Hadamard transform and 40-point Hadamard transform. For the 128-point Hadamard transform, the HTU adopts an optimized scheme based on the Fast Hadamard Transform (FHT) algorithm. The FHT algorithm has a high degree of parallelism. It efficiently processes data through multiple stages of butterfly operations (Butterfly Core) and combines a FIFO queue for data caching and transmission. Each stage disassembles the computing tasks into multiple parallel operations through a pipeline design, enabling each data point to perform multiple transform calculations simultaneously, thus achieving extremely high computational density and throughput.

[0077] The design of the HTU takes into account the efficiency and performance of hardware implementation. In the 128-point Hadamard transform, by pushing the first 64 data elements into the input FIFO, the FIFO queue is used to achieve pipelined data transmission, ensuring a smooth transition in the data processing process. As subsequent data elements are received, they are processed in pairs through the Butterfly Core, and the output results are sent to the next processing stage or cache to ensure the efficient progress of subsequent calculations. This design not only greatly improves the data processing speed but also reduces the memory access latency and optimizes the computational efficiency of the hardware.

[0078] Such as Figure 4As shown, for the 128-point Hadamard, the entire calculation process is divided into 7 stages, and each stage has a Butterfly Core. The input data will be distributed to the FIFO or Butterfly Core through a multiplexer. The Butterfly Core will receive the directly fed data and the data fed through the FIFO with a delay, and perform an addition and a subtraction operation on the data simultaneously. The Butterfly Core will distribute the output data to the FIFO and the multiplexer respectively. The multiplexer will select the data output in sequence and send it to the next-stage Butterfly Core. The FIFO depth of the Butterfly Core will decrease stage by stage.

[0079] For the 40-point Hadamard transform, the HTU uses a simplified matrix multiplication unit (MMU) to implement it. Specifically, the input data of the Hadamard matrix is fixed to the values of 1 and -1, thereby reducing the complexity of hardware implementation. Through this design, the HTU can still maintain high computational efficiency and throughput when processing relatively small-scale Hadamard transforms.

[0080] Taking the deployment of the Mamba2-2.7B model as an example, the calculation process of the present invention is specifically implemented as follows:

[0081] 1. Model loading and initialization

[0082] Weight storage: The 64-layer weight parameters of the Mamba2-2.7B model are stored in the HBM / LPDDR module layer by layer. Each layer of weights includes the input projection layer (W in ), SSM parameters (Wconv), and the output projection layer (W out ).

[0083] Hidden state management: The initial hidden state (h t ) is loaded from the HBM to the on-chip cache through the DMA Read&Write module, and the output results of each layer are dynamically updated.

[0084] Quantization configuration: The rotation-assisted post-training quantization (PTQ) algorithm is adopted to perform 4-bit PoT (power of 2) quantization on the SSM layer parameters, and the quantization parameters are configured to the HTU module through the ARM CPU.

[0085] 2. Input projection layer calculation

[0086] Data input: The input sequence (X in ) is normalized by the Rmsnorm module and then transmitted to the MMU module.

[0087] Parallel matrix multiplication:

[0088] The MMU adopts a tree-shaped MAC structure, decomposing the din×dout matrix multiplication of the input projection layer into din×dout2din×2dout DSP units for parallel computing, and realizing single-cycle dual operations (multiplication and addition) through DSP packing technology.

[0089] Time-division multiplexing scheduling: The MMU alternately processes the input projection layer (generating Δ, B, C, X, Z) and the output projection layer calculation. The input projection layer is given priority, and the output Δ, B, and C parameters are stored in the on-chip dual buffer (Buffer A / B) to ensure continuous execution of the SSMU pipeline.

[0090] Data caching: The generated Δ, B, and C parameters are temporarily stored in Buffer A according to the block size (np×pp). At the same time, Buffer B receives the next batch of data, realizing the ping-pong operation between calculation and caching.

[0091] 3. Execution of the SSM inference pipeline

[0092] Computation reordering and chunking strategy:

[0093] Data generation order: The SSMU generates intermediate data in the order of Δ→B→C→X→Z to eliminate inter-layer dependencies.

[0094] Fine-grained chunking: The SSM calculation is split into chunks with np = 16 and pp = 64 along the hidden state dimension. Each chunk is processed independently, and the intermediate results are directly passed to the next stage through operation fusion, avoiding the storage of activation values.

[0095] Full pipeline architecture:

[0096] Dedicated computing units: The SSMU internally deploys an element multiplication unit (EMU), a state update unit (SUU), and an activation function unit (AFU), and each unit is connected through a FIFO buffer.

[0097] Dynamic parallelism adjustment: The EMU dynamically allocates parallelism (2 / 4 / 8-way) according to the operator complexity (such as and ) to balance the data stream throughput.

[0098] Hidden state update: After each layer of SSM calculation is completed, the updated ht is written back to the HBM through DMA for the next layer to call. Storage optimization: 80% of the intermediate activation values are reduced through operation fusion, and the on-chip storage occupancy is reduced from 16MB in the traditional scheme to 4MB.

[0099] 4. Hadamard transform and quantization acceleration

[0100] Execution of the 128-point HTU:

[0101] Butterfly Core Pipeline: Input data is pushed into the input FIFO in batches (64 points per batch), and the transformation is completed through 4 - level Butterfly Core operations. Each level of the core contains 32 parallel butterfly units, supporting the completion of 64 - point calculations in a single cycle.

[0102] Quantization Collaboration: The transformed data is quantized to 4 bits by the rotation - assisted PTQ algorithm. The quantization parameters (scaling factor, zero point) are dynamically adjusted through shift operations, reducing the off - chip communication bandwidth by 60%.

[0103] 40 - point HTU Execution:

[0104] Simplified Matrix Multiplication: The Hadamard matrix is fixed with values of ±1. The MMU performs bitwise operations (XOR instead of multiplication) on the 40 - point input vector and the pre - stored Hadamard kernel (40×40), reducing the calculation delay to 30% of the traditional scheme.

[0105] Data Return: The quantized results are called by the output projection layer.

[0106] 5. Output Projection Layer and Result Generation

[0107] Time - Multiplexing Scheduling: The MMU switches to the output projection mode, loads the Wout weights from the HBM, and performs the final matrix multiplication. Tree - shaped MAC Optimization: The d hidden ×d out operations are mapped to MAC units through DSP packing technology.

[0108] Result Output: The inference result is normalized by the Rmsnorm module and then output. At the same time, the global hidden state ht is updated to complete the single - layer calculation.

[0109] Finally, it should be noted that the purpose of publishing the embodiments is to help further understand the present invention. However, those skilled in the art can understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection claimed by the present invention is defined by the scope of the claims.

Claims

1. An FPGA accelerator for state-space model reasoning, characterized in that: An ARM CPU, an HBM / LPDDR module, a DMA Read module, a DMA Read&Write module, and an Rmsnorm module are arranged on the FPGA accelerator, as well as an MMU module for executing all linear layer operations in the Mamba model, an SSMU module for executing SSM operations in the Mamba model, and an HTU module for performing hadamard transform operations, wherein the ARM CPU communicates with the DMA Read module, the DMA Read&Write module, the Rmsnorm module, and the HTU module through an AXI-Lite bus, and is responsible for global scheduling, parameter configuration, and task start and stop control; all weight parameters of the Mamba model are efficiently stored in the HBM / LPDDR module; the HBM / LPDDR module is connected to the DMA Read module and the DMA Read&Write module through a high-bandwidth AXI bus to achieve fast access to model weight parameters and hidden states; the DMA Read module sequentially reads weight parameters of each layer from the HBM / LPDDR module, and transmits them to the MMU module and the SSMU module through a dedicated data channel; the DMA The Read&Write module is responsible for bidirectional transmission: reading the initial hidden state from the HBM / LPDDR module to the on-chip cache of the SSMU module, and writing the updated state back to the HBM / LPDDR module; the MMU module and the SSMU module implement data interaction through the AXI protocol: the intermediate results output by the SSMU are transmitted to the HTU module; after the HTU module completes the Hadamard transform and quantization, it transmits the quantization results back to the HBM / LPDDR module through the DMA channel; the Rmsnorm module is integrated in the data input / output end, normalizes the input sequence and the final inference result, and is directly connected to the MMU module through an independent data path.

2. The FPGA accelerator for state-space model reasoning according to claim 1, characterized in that: The MMU module adopts a tree-shaped multiplication-accumulation structure to support parallel computing of the input / output projection layer; wherein the matrix multiplication of the input projection layer is decomposed into multi-vector multiplications, each vector multiplication is processed by a tree-shaped accumulator, and the tree-shaped accumulator is composed of DSP units, and adopts a hierarchical design: the bottom parallel DSP unit performs multiplication and addition operations on two input vectors; The parallel DSP units in the middle layer are local accumulators that summarize the partial sums; the top-level DSP units are global accumulators that generate the final result and output the result of the multiplication and accumulation of two vectors.

3. The FPGA accelerator for state-space model reasoning according to claim 1, characterized in that: The SSMU module divides the input data into three parts: Δ, BCX, and Z through a multiplexer. Δ is sent to the bias accumulation module through the AXI interface, and after accumulation, it is sent to the Softplus module through the AXI interface; BCX is sent to the Conv module through the AXI interface; the Conv module and Z are sent to the Silu module through the AXI interface after passing through the multiplexer; the Silu module divides the data into four parts: B, C, X, and Z through the multiplexer, and B and C are stored in the Buffer for subsequent reading; the Softplus module sends the Δ data to two EMU modules respectively, and performs multiplication operations with data A and data B respectively; after multiplication with A, it generates Send to EMU module; B and Δ pass through EMU module to get and X are multiplied by the EMU module to obtain and the hidden state h t-1 After multiplication through the EMU module and Accumulate to get the hidden state h t ; Hidden state h t Part of it is sent to the off-chip storage space for storage, and part of it is sent to another EMU module for matrix multiplication with C and obtained through the accumulator h t C; the other part X is sent to the EMU module, multiplied with D stored outside the chip, and then with h t C performs accumulation operation to obtain Y, which is sent to the EMU module for multiplication operation with Z to obtain the output sequence.

4. The FPGA accelerator for state-space model reasoning according to claim 1, characterized in that: The HTU module includes a 128-point HTU based on a fast Hadamard transform FHT algorithm and a 40-point HTU based on a simple matrix multiplication unit MMU; the 128-point HTU adopts a multi-stage pipeline of a butterfly core and a FIFO buffer; the 40-point HTU uses a Hadamard matrix through a fixed input.

5. The computing method of the FPGA accelerator for state-space model reasoning as claimed in claim 1, wherein the steps include: 1) Model loading and initialization, including: 1-1) The weight parameters of the Mamba model are stored in the HBM / LPDDR module in layer order, and each layer of weight includes the input projection layer, SSM parameters and output projection layer; 1-2) The initial hidden state is loaded from the HBM / LPDDR module to the on-chip cache, and the output results of each layer are dynamically updated; 1-3) Quantize the SSM parameters of each layer and configure the quantized parameters to the HTU module; 2) Input projection layer calculation, including: 2-1) Input sequence X in After normalization, it is transmitted to the MMU module; 2-2) The MMU module alternately processes the input projection layer and the output projection layer calculations, where the input projection layer is executed first, and the output Δ, B, and C parameters are stored in the on-chip double buffer; 3) SSM inference pipeline execution, including: 3-1) The SSMU module generates intermediate data in the order of Δ→B→C→X→Z to eliminate inter-layer dependencies; 3-2) The SSMU module internally deploys an element multiplication unit, a state update unit, and an activation function unit, and each unit is connected through a FIFO buffer; 3-3) After each layer of SSM calculation is completed, the updated initial hidden state h t Write back to the HBM / LPDDR module via DMA for the next layer to call; 4) Hadamard transform and quantization acceleration, including: 4-1) Input data is pushed to the input FIFO in batches and transformed through butterfly operation; 4-2) The transformed data is quantized by the rotation-assisted PTQ algorithm, and the quantization parameters are dynamically adjusted through the shift operation; 4-3) The quantized results are used by the output projection layer; 5) Output projection layer and result generation, including: 5-1) The MMU switches to output projection mode, loads the output projection layer weights from the HBM / LPDDR module, and performs the final matrix multiplication; 5-2) The inference results are normalized and output, and the global hidden state is updated to complete the single-layer calculation.

6. The computing method of the FPGA accelerator for state-space model reasoning according to claim 5, characterized in that: In steps 1-3), the rotation-assisted post-training quantization PTQ algorithm is used to quantize the SSM parameters of each layer.

7. The computing method of the FPGA accelerator for state-space model reasoning according to claim 5, characterized in that: The HTU module implements pipeline transmission of data by pushing the first 64 data elements to the input FIFO in the 128-point Hadamard transform. As subsequent data elements are received, they are processed in pairs by the Butterfly Core, and the output results are sent to the next processing stage or cache.

Citation Information

Cited By

  • Adaptive optical matrix calculation acceleration method based on ARM architecture

    CN120724028A