FPGA superposition processor acceleration system and method based on state space duality
By designing a FPGA superimposed processor acceleration system based on state space duality, the Mamba2 model has solved the problems of high memory footprint, low element-by-element computing efficiency and serious sparse computing redundancy in deployment, and achieved significant acceleration effect.
Patent Information
- Application Number
- CN202510669681.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-23
AI Technical Summary
In actual deployment, the Mamba2 model has problems such as high memory footprint, low element-by-element computing efficiency and serious sparse computing redundancy, which is difficult to adapt to the existing hardware platform.
A FPGA overlay processor acceleration system based on state space duality is designed, including sparse predefined data acquirers, reconfigurable pulsating arrays, partial and buffers, element-by-element operation caches, function calculations, and on-chip memory management modules. The zero elements are eliminated by the sparse predefined data acquirer. The reconstructed pulsating array supports multiple calculation modes, and the partial and caches realize cross-period accumulation, element-by-element calculation cache integration results, the function calculation module completes nonlinear operations, and the on-chip memory management module optimizes the result cache.
It significantly reduces memory usage, improves element-by-element computing efficiency, and reduces sparse computing redundancy, thereby greatly accelerating the inference process of the Mamba2 model.
Smart Images

Figure CN120179608A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine learning, and particularly relates to an FPGA stacked processor acceleration system and method based on state space duality. Background Art
[0002] In today's digital age, the field of large language models has developed rapidly. Models represented by Transformer have been widely used in many scenarios such as intelligent assistants, content creation, and personalized recommendations, and they have shown excellent performance in dealing with long-range dependencies of sequential data. With the rise of edge computing, higher requirements are put forward for the inference performance of models in resource-constrained environments. Against this background, the Mamba model based on the state space model has emerged. It achieves nearly linear computational complexity through parallel scanning and loop calculation, and is suitable for resource-constrained deployment. The subsequent Mamba2 model introduces state space duality (SSD), further improving the computational efficiency on parallel hardware such as GPUs, attracting wide attention in the industry and bringing new solutions for long-sequence data processing.
[0003] However, a series of serious problems have emerged in the actual deployment of the Mamba2 model. Specifically, in terms of memory, a large number of broadcast operations in the SSD mechanism expand small tensors into large tensors for element-wise multiplication, resulting in a sharp increase in memory occupancy and bandwidth requirements during inference. The memory requirement in the double attention mode is as high as 41.55GB, and it also reaches 10.5GB in the loop mode, which poses a great pressure on memory resources. At the level of hardware computing performance, mainstream GPUs and traditional FPGA accelerators are designed for tensor compression and dense matrix calculations, and have insufficient support for element-wise calculations after broadcast expansion. In the SSD mechanism, element-wise operations account for 83.88% of the total inference process calculation delay, seriously slowing down the inference speed. The problem of computational redundancy is also prominent. There are a large number of sparse matrix multiplications in SSD calculations, and traditional computing units cannot effectively skip the zero-value regions in sparse matrices, resulting in a large number of redundant operations, significantly affecting the overall computational performance. Therefore, traditional sparse optimization and memory optimization methods have problems such as memory occupancy, computational efficiency, and redundant calculations, and it is difficult to adapt to the unique structure of the Mamba model.
[0004] It can be seen that the existing hardware platforms have problems of huge memory overhead, low element-wise calculation efficiency, and serious sparse calculation redundancy. Summary of the Invention
[0005] The present invention provides an FPGA stacked processor acceleration system and method based on state space duality. This system reduces the memory overhead in the state space duality calculation process, and improves the computational efficiency, inference speed, and energy efficiency.
[0006] To achieve the above object, the present invention adopts the following technical solutions: An FPGA stacked processor acceleration system based on state space duality, comprising: A sparse predefined data acquirer module, configured to generate a valid data access address sequence according to a predefined sparse pattern, and extract valid data according to the valid data access address sequence; wherein, the valid data includes input data and weights from which zero elements have been removed. A reconfigurable systolic array module, configured to receive the valid data transmitted by the sparse predefined data acquirer module, and perform dense matrix operations, sparse matrix operations, element-wise operations, and segmented multiplication operations on the valid data. A partial sum cache module, configured to store intermediate calculation results after the reconfigurable systolic array module performs dense matrix operations, sparse matrix operations, and segmented multiplication operations, and perform cross-cycle accumulation operations on the intermediate calculation results. An element-wise operation cache module, configured to receive the element-wise operation results after the reconfigurable systolic array module performs element-wise operations and the accumulated intermediate calculation results after the cross-cycle accumulation operations of the partial sum cache module. A function calculation module, configured to perform non-linear operations on the element-wise operation results and the accumulated intermediate calculation results transmitted by the element-wise operation cache module to obtain non-linear operation results. An on-chip memory management module, configured to directly cache the non-linear operation results output by the function calculation module into the FPGA internal cache for subsequent calculations or transmit them back to the DDR memory as the final results.
[0007] Further, it further includes an input cache and weight cache module, configured to provide access addresses corresponding to the input data and weights to be calculated. The sparse predefined data acquirer module extracts valid data from the access addresses of the input cache and weight cache module according to the valid data access address sequence.
[0008] Further, it further includes an instruction control unit module, configured to receive an instruction sequence transmitted by the host CPU, and control the sparse predefined data acquirer module, the reconfigurable systolic array module, the partial sum cache module, the element-wise operation cache module, the function calculation module, and the on-chip memory management module to work together according to the instruction sequence.
[0009] Further, the sparse predefined data acquirer module includes an address sequence generator, a sparse data register group, and a distributed queue unit. The address sequence generator generates a valid data access address sequence according to a predefined sparse pattern to eliminate zero-element data; extracts valid data according to the valid data access address sequence, and stores the valid data in the sparse data register bank; and transmits the valid data to the reconfigurable systolic array module through the distributed queue unit.
[0010] Further, the reconfigurable systolic array module includes a plurality of cascaded processing units; each processing unit includes an input selection multiplexer, a data input register, a control logic unit, a matrix multiply-accumulate unit, and an element-by-element calculation unit.
[0011] Further, a bi-directional connection is adopted both between the element-by-element operation cache module and the reconfigurable systolic array module and between the element-by-element operation cache module and the function calculation module.
[0012] A control method for an FPGA overlay processor acceleration system based on state space duality, based on the above-mentioned FPGA overlay processor acceleration system based on state space duality, includes: Through the sparse predefined data acquirer module, generate a valid data access address sequence according to a predefined sparse pattern, and extract valid data according to the valid data access address sequence; wherein, the valid data includes input data and weights with zero elements eliminated. Through the reconfigurable systolic array module, receive the valid data transmitted by the sparse predefined data acquirer module, and perform dense matrix operations, sparse matrix operations, element-by-element operations, and segmented multiplication operations on the valid data. Through the partial sum cache module, store the intermediate calculation results after the reconfigurable systolic array module performs dense matrix operations, sparse matrix operations, and segmented multiplication operations, and perform cross-cycle accumulation operations on the intermediate calculation results. Through the element-by-element operation cache module, receive the element-by-element operation results after the reconfigurable systolic array module performs element-by-element operations and the accumulated intermediate calculation results after the cross-cycle accumulation operations of the partial sum cache module. Through the function calculation module, perform non-linear operations on the element-by-element operation results and the accumulated intermediate calculation results transmitted by the element-by-element operation cache module to obtain non-linear operation results. Through the on-chip memory management module, directly cache the non-linear operation results output by the function calculation module into the FPGA internal cache for subsequent calculations or return them to the DDR memory as the final results.
[0013] Further, before the step of generating a valid data access address sequence according to a predefined sparse pattern and extracting valid data through the sparse predefined data acquirer module, it further includes: Based on the data flow graph formed by each operator in the state space duality calculation process, the state space duality calculation process is decomposed into multiple basic operator sequences that can be parallelized and reordered.
[0014] Furthermore, in the process of the reconfigurable systolic array module receiving the valid data transmitted by the sparse predefined data acquirer module and performing dense matrix operations, sparse matrix operations, element-wise operations, and segmented multiplication operations on the valid data, it further includes: Using an operation fusion framework, combining broadcast matrix multiplication and summation operations into vector matrix multiplication operations; moving the segmented multiplication in causal matrix generation to the element-wise operation stage; Selecting a calculation mode according to the operator type and generating a corresponding instruction sequence for the reconfigurable systolic array module to execute the corresponding operator operation.
[0015] Furthermore, When the operator type is dense matrix operation, general matrix multiplication is adopted to optimize the data block reading through the address mapping list; When the operator type is sparse matrix operation, sparse matrix multiplication is adopted to eliminate redundant calculations of zero elements by using the tensor rearrangement algorithm.
[0016] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides an FPGA overlay processor acceleration system based on state space duality, including modules such as a sparse predefined data acquirer, a reconfigurable systolic array, a partial sum cache, an element-wise operation cache, function calculation, and on-chip memory management. The sparse predefined data acquirer eliminates zero elements to reduce redundant data, significantly reducing redundant calculations; the reconfigurable systolic array flexibly supports multiple calculation modes, improving the utilization rate of hardware resources; the partial sum cache realizes cross-cycle accumulation, the element-wise operation cache integrates the results, the function calculation module completes non-linear operations, and the on-chip memory management module optimizes the result cache to ensure the on-chip implementation of SSD calculations, significantly reducing off-chip memory access and improving the inference efficiency and energy efficiency; this system addresses the problems of large memory occupancy, low element-wise calculation efficiency, and redundant sparse calculations in the Mamba2 model. By reducing redundant data transmission and calculations, optimizing memory access, and leveraging the reconfigurable characteristics of FPGA to efficiently execute various operations. Using this system effectively reduces memory occupancy, improves element-wise calculation efficiency, reduces redundant sparse calculations, and thus significantly accelerates the inference process of the Mamba2 model.
[0017] Preferably, in the present invention, by adding an input cache and a weight cache module, a unified source of input data and weight access addresses is provided for the system, enabling the sparse predefined data fetcher module to more efficiently extract valid data from fixed positions, avoiding confusion in data sources and additional address lookup operations, further improving the efficiency and accuracy of data extraction, and contributing to the more stable and rapid operation of the entire acceleration system.
[0018] Preferably, in the present invention, an instruction control unit module is introduced. By receiving the instruction sequence transmitted by the host CPU, the coordinated control of each module in the system is achieved. This enables the entire acceleration system to work in accordance with a predetermined order and logic, avoiding conflicts and confusion between modules, improving the overall operation efficiency and stability of the system, and ensuring the accurate and efficient completion of complex computing tasks.
[0019] Preferably, in the present invention, the sparse predefined data fetcher module specifically includes an address sequence generator, a sparse data register bank, and a distributed queue unit. The address sequence generator can accurately generate a valid data access address sequence according to the predefined sparse pattern, effectively eliminating zero-element data; the sparse data register bank provides a dedicated storage space for valid data, facilitating subsequent processing; the distributed queue unit realizes the efficient transmission of valid data. These components work together, enabling the sparse predefined data fetcher module to more quickly and accurately extract valid data, providing strong support for subsequent calculations.
[0020] Preferably, in the present invention, the reconfigurable systolic array module includes a number of cascaded processing units. The number of cascaded processing units provides powerful computing capabilities and can process multiple data simultaneously; the input selection multiplexer realizes flexible data selection and input; the data input register temporarily stores the input data, ensuring the stability of data processing; the control logic unit precisely controls the operation of the processing units; the matrix multiply-accumulate unit and the element-wise calculation unit respectively implement dense matrix operations and element-wise operations. These designs enable the reconfigurable systolic array module to efficiently and flexibly complete various computing tasks.
[0021] Preferably, in the present invention, the two-way connection mode of the element-wise operation cache module with the reconfigurable systolic array module and the function calculation module. This two-way connection enables the element-wise operation cache module to simultaneously receive the element-wise operation results from the reconfigurable systolic array module and the accumulated intermediate calculation results from the partial sum cache module, and timely transmit these results to the function calculation module for non-linear operations. At the same time, the function calculation module can also feedback the calculation results to the element-wise operation cache module, realizing flexible data interaction and sharing, and improving the overall computing efficiency and flexibility of the system.
[0022] The present invention also provides a control method for an FPGA stacked processor acceleration system based on state - space duality. Based on the above - mentioned acceleration system, this method uses a sparse predefined data acquirer to eliminate zero elements and reduce the data volume. The reconfigurable systolic array can flexibly execute various matrix operations. The partial - sum cache realizes cross - cycle accumulation of intermediate results. The element - by - element operation cache integrates element - by - element operations and accumulation results. The function calculation module completes non - linear operations, and the on - chip memory management module optimizes the result cache. Aiming at the problems of large memory occupation, low element - by - element calculation efficiency, and sparse calculation redundancy in the Mamba2 model, by reducing redundant data and calculations, optimizing memory access, and utilizing the reconfigurable characteristics of the FPGA, the calculation efficiency is improved. Using this method significantly reduces the memory occupation, improves the element - by - element calculation efficiency, reduces sparse calculation redundancy, thus greatly accelerating the inference process of the Mamba2 model and effectively solving the key problems faced by existing hardware platforms.
[0023] Preferably, in the present invention, before the sparse predefined data acquirer module extracts valid data, a step of decomposing the data - flow graph in the calculation process based on state - space duality is added. The beneficial effect is that the duality calculation process is decomposed into multiple basic operator sequences that can be parallelized and reordered, providing a basis for subsequent parallel calculation and operator reordering, helping to improve the calculation efficiency and flexibility of the system, enabling the system to more efficiently utilize hardware resources, and further enhancing the overall performance.
[0024] Preferably, in the present invention, an operation fusion framework is used to combine broadcast matrix multiplication and summation operations into vector matrix multiplication operations, reducing the calculation steps and latency; moving the segmented multiplication in causal matrix generation to the element - by - element operation stage to optimize the calculation process; selecting the calculation mode according to the operator type and generating the corresponding instruction sequence, enabling the reconfigurable systolic array module to more accurately execute the corresponding operator operations, improving the calculation accuracy and efficiency, enabling the system to better adapt to different calculation tasks, and giving full play to the hardware advantages of the FPGA.
[0025] Preferably, in the present invention, for dense matrix operations, general matrix multiplication is used and data - block reading is optimized through an address - mapping list, which can improve data - reading efficiency and reduce memory - access time; for sparse matrix operations, sparse matrix multiplication is used and tensor rearrangement algorithms are utilized to eliminate redundant zero - element calculations, reducing the calculation amount and memory occupation. These optimization methods further improve the calculation performance and resource utilization rate of the system, enabling the system to more efficiently process various types of calculation tasks and effectively solving the problems of the Mamba2 model in terms of memory and calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of the software stack and hardware micro - architecture of the FPGA stacked processor acceleration system based on state - space duality provided by an embodiment of the present invention; Figure 2 Schematic diagram of the sparse predefined data acquirer and processing unit provided by the present invention; Figure 3 Schematic diagram of the on-chip cache management and data interaction mechanism provided by the present invention; Figure 4 Detailed schematic diagram of the sparse matrix multiplication calculation process provided by the present invention; Figure 5 Detailed schematic diagram of the segmented multiplication calculation process provided by the present invention; Figure 6 Schematic diagram of the structure of an FPGA superposition processor acceleration system based on state space duality provided by the present invention. Specific implementation manners
[0027] To further understand the content of the present invention, the following describes the present invention in detail with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention and not for limiting it.
[0028] The following explains the technical terms involved in the present invention: FPGA (Field-Programmable Gate Array), field programmable gate array, is an integrated circuit that can configure hardware logic through programming and is commonly used in the fields of high-performance computing and hardware acceleration.
[0029] DDR (Double Data Rate), double data rate, is a random access memory (RAM) technology widely used in computers and electronic devices for high-speed data transmission.
[0030] SSD (State Space Duality), state space duality, is a key technology introduced in the Mamba2 model for improving the computing efficiency on parallel hardware such as GPUs.
[0031] GPU (Graphics Processing Unit), graphics processing unit, was originally used for graphics rendering and is now widely used in the fields of parallel computing, artificial intelligence, etc. to accelerate large-scale data processing.
[0032] Mamba is a machine learning model, specifically a sequence data processing model based on the state space model, with near-linear computational complexity and suitable for inference in resource-constrained environments.
[0033] Mamba2 is an improved version of the Mamba model, which introduces state space duality (SSD) to further improve the computing efficiency.
[0034] The CPU (Central Processing Unit) is the computing core and control core of a computer manufactured using very large scale integrated circuits.
[0035] The ASG (Address Sequence Generator) is an address sequence generator. Its function is to generate a valid data access address sequence according to a predefined sparse pattern, thereby eliminating zero-element data, and then extracting valid data according to this sequence and storing it in the sparse data register bank.
[0036] The BRAM (Block Random Access Memory), that is, the block random access memory, exists in units of blocks and has an independent storage structure and read / write ports.
[0037] The PCIe (Peripheral Component Interconnect Express) interface is a high-speed serial computer expansion bus standard used to connect various hardware devices inside a computer, such as graphics cards, solid-state drives, network cards, sound cards, etc.
[0038] As described in the background art, traditional sparse optimization strategies such as pruning, quantization, and memory optimization methods (model partitioning, computational graph fusion, and data flow optimization) cannot be directly and effectively applied to the unique structure of the Mamba model. Therefore, there is an urgent need to design a brand-new acceleration optimization scheme that can simultaneously solve the problems of high memory occupancy, low per-element calculation performance, and sparse calculation redundancy, and achieve efficient acceleration of the Mamba2 model; it can be seen that the existing hardware platforms have problems such as huge memory overhead, low per-element calculation efficiency, and serious sparse calculation redundancy.
[0039] To solve the above problems, as Figure 6As shown in the figure, this embodiment provides an FPGA overlay processor acceleration system based on state space duality, including a sparse predefined data acquirer module, which is used to generate a valid data access address sequence according to a predefined sparse pattern and extract valid data according to the valid data access address sequence; where the valid data includes input data and weights with zero elements removed; a reconfigurable systolic array module, which is used to receive the valid data transmitted by the sparse predefined data acquirer module and perform dense matrix operations, sparse matrix operations, element-wise operations, and segmented multiplication operations on the valid data; a partial sum cache module, which is used to store the intermediate calculation results after the reconfigurable systolic array module performs dense matrix operations, sparse matrix operations, and segmented multiplication operations, and perform cross-cycle accumulation operations on the intermediate calculation results; an element-wise operation cache module, which is used to receive the element-wise operation results after the reconfigurable systolic array module performs element-wise operations and the accumulated intermediate calculation results after the cross-cycle accumulation operations of the partial sum cache module; a function calculation module, which is used to perform non-linear operations on the element-wise operation results and the accumulated intermediate calculation results transmitted by the element-wise operation cache module to obtain non-linear operation results; an on-chip memory management module, which is used to directly cache the non-linear operation results output by the function calculation module into the FPGA internal cache for subsequent calculations or return them to the DDR memory as the final results.
[0040] Exemplarily, this embodiment provides an FPGA overlay processor acceleration system based on state space duality, which simultaneously improves two parts: the software stack and the hardware microarchitecture. The specific implementation of this acceleration system includes: The hardware microarchitecture of this FPGA overlay processor acceleration system includes: an instruction control unit module, an input cache and weight cache module, a sparse predefined data acquirer module, a reconfigurable systolic array module, a partial sum cache module, an element-wise operation cache module, a function calculation module, and an on-chip memory management module. The specific structure and functions are as follows: Instruction control unit module: Composition: It includes an instruction decoding unit, an instruction cache unit, and a timing control logic.
[0041] Function: Receive the instruction sequence compiled by the host CPU and control the coordinated work of each module of the entire processor system.
[0042] Input cache and weight cache module: Composition: The input cache and the weight cache are respectively composed of multiple high-speed on-chip storage units.
[0043] Connection: They are respectively connected to the sparse predefined data acquirer module to provide the data and weights to be calculated.
[0044] Function: Provide fast data and weight access to meet the high-bandwidth and low-latency computing requirements.
[0045] Sparse Predefined Data Fetcher Module: Composition: Address Sequence Generator (ASG), Sparse Data Register Bank, Distributed Queue Unit.
[0046] Connection: Input Cache and Weight Cache Modules to Reconfigurable Systolic Array Module.
[0047] Function: Generate an address sequence according to a predefined sparse pattern, exclude zero-element data loading in advance, thereby precisely controlling the transmission of sparse matrix data and minimizing redundant operations.
[0048] Reconfigurable Systolic Array Module: This module is composed of several Processing Elements (PEs) in cascade. The structure of each processing element is as follows: Composition of Processing Element (PE): Composed of an Input Selection Multiplexer, Data Input Register, Control Logic Unit, and Computational Unit (including Matrix Multiply-Accumulate Unit MAC and Element-Wise Computational Unit).
[0049] Connection Method: PEs are cascaded through horizontal and vertical data paths to support efficient cross-PE data transmission; at the same time, each PE unit is connected to the Sparse Predefined Data Fetcher Module to achieve precise data loading; the output of the PE is connected to the Partial Sum Cache Module and the Element-Wise Operation Cache Module to store the intermediate results of each stage of calculation.
[0050] Function and Computational Mode Support: The processing element supports General Matrix Multiplication (GEMM) to implement dense matrix operations, supports Sparse Matrix Multiplication (SpMM) to improve computational efficiency by dynamically skipping zero elements, supports element-wise operations to achieve efficient parallel point-wise calculations, and supports Segment-Multiplication to efficiently perform causal mask-related operations.
[0051] Partial Sum Cache Module: Composition: Multi-level Cache Register Array.
[0052] Connection: Connected to the output of the Reconfigurable Systolic Array Module and the Element-Wise Operation Cache Module.
[0053] Function: Store the intermediate results during the calculation process, support cross-cycle accumulation of multiple rounds of calculations, and ensure efficient and stable partial sum operations.
[0054] Element-Wise Operation Cache Module Composition: Two dedicated high-speed BRAM cache arrays (EleOps Buffer A, B).
[0055] Connection: This module is bidirectionally connected to the reconfigurable systolic array module and the function calculation module (SFU module), thus supporting efficient interaction for element-wise operations.
[0056] Function: Provide high-bandwidth caches required for element-wise calculations to ensure efficient element-wise data transfer.
[0057] Function calculation module: Composition: A pipelined calculation structure, including exponential function (EXP), activation function (SiLU), normalization calculation unit (RMSNorm), etc.; the above functions are all dedicated functions.
[0058] Connection: Connect to the element-wise operation cache module and the on-chip memory management module respectively.
[0059] Function: Efficiently complete the non-linear operations in the SSD algorithm, significantly reducing the calculation latency of non-linear functions.
[0060] On-chip memory management module: Composition: This module consists of a width-first buffer, a depth-first cache, and data interaction control logic.
[0061] Connection: Connect to the element-wise cache module, SFU module, and systolic array module.
[0062] Function: Through the intelligent allocation of caches, enable the SSD calculation process to be fully executed within the on-chip cache; reduce the number of off-chip memory accesses and improve the utilization rate of data transfer bandwidth.
[0063] Exemplarily, based on the hardware microarchitecture of the above FPGA stacked processor acceleration system, the overall data flow and signal transmission relationship of this system are as follows: Data input process: Data is read by the input cache and weight cache module, and then passed into the reconfigurable systolic array module through the sparse predefined data acquirer module for sparse calculation.
[0064] Calculation process: After the reconfigurable systolic array module completes the GEMM operation, SpMM operation, element-wise operation, and Segment-mul operation respectively, it outputs some results to the partial sum cache module or the element-wise operation cache module. The results of the element-wise operation cache module are further passed to the function calculation module for non-linear operation. Among them, the calculation results received by the partial sum cache module are the intermediate calculation results after the reconfigurable systolic array module executes the dense matrix operation, sparse matrix operation, and segmented multiplication operation, and perform cross-cycle accumulation operations on the intermediate calculation results; the element-wise operation cache module receives the element-wise operation results after the reconfigurable systolic array module executes the element-wise operation and the accumulated intermediate calculation results after the cross-cycle accumulation operation of the partial sum cache module.
[0065] Result output process: The final result processed by the function calculation module is directly cached on the chip through the on-chip memory management module, waiting for the next stage of calculation or transmitted back to the DDR memory through a dedicated path.
[0066] Exemplarily, in order to further improve the performance of the system, this embodiment also performs optimization and improvement on the software stack, that is, provides a control method for the FPGA overlay processor acceleration system based on the duality of the state space, as follows: The software stack design provided in this embodiment is committed to reducing the memory overhead in the SSD calculation process and improving the resource utilization efficiency of the hardware microarchitecture. The overall software stack runs on the host CPU and mainly consists of the following key steps: First, Data-flow Graph Unrolling: In order to efficiently map the Mamba model to the FPGA hardware platform, in this embodiment, first, the data flow graph (DFG) composed of each operator in the SSD calculation is analyzed and unfolded, and the complex SSD calculation is divided into several basic operator sequences that can be parallelized and reordered, so as to perform hardware mapping more efficiently later.
[0067] Second, adopt the Operator Fusion framework: In the SSD calculation, there are a large number of broadcast element-wise multiplication operations, causing huge memory pressure. To alleviate this problem, this embodiment proposes an operator fusion framework for software-hardware collaboration.
[0068] Exemplarily, the specific fusion strategy includes the following two important aspects: (1)Operator Merge: The broadcast matrix multiplication and the subsequent summation operation can essentially be uniformly represented as a simpler linear operator calculation. This embodiment proposes to fuse the two into a unified descriptor and express them uniformly through vector matrix multiplication. In this way, through operator merge, the original two-step operation (broadcast multiplication and summation) is unified into a single matrix multiplication operation, effectively shortening the calculation path and significantly reducing the memory overhead of the intermediate feature tensor.
[0069] (2)Operator Backward Shift: In the calculation of state space duality (SSD), the generation of the causal matrix M is usually achieved through segment-mul, which involves segment summation and exponentiation operations. In this method, by utilizing the properties of the SSD calculation itself, some operations are executed in advance, that is, the segment-mul is embedded into the subsequent element-wise operation process, thus shortening the overall calculation path. The implementation of operator backward shift effectively advances this segment calculation and reduces the memory burden and latency of the complex calculation stage.
[0070] Third, adopt the data allocation algorithm (Data Allocation): This embodiment also proposes a set of efficient data allocation algorithms to support the tile-granularity multiply-accumulate (MAC) calculation of matrix operations.
[0071] Specifically, it includes two calculation modes: (1)GEMM mode (General Matrix Multiplication): The GEMM mode supports linear matrix multiplication and the fused broadcast MAC calculation. By precisely defining the address mapping list before and after the broadcast through software, the required data blocks are efficiently read in each cycle and sent to the systolic array for vectorized MAC calculation after reordering, thus significantly reducing the memory occupancy.
[0072] (2)SpMM mode (Sparse Matrix Multiplication): To solve the problem of sparse calculation redundancy, this embodiment also designs a set of tensor-reorder-and-group transformation algorithms: First, apply a zero-element mask to the sparse matrix to eliminate the invalid calculations of zero elements, and then regroup and sort the non-zero element data in a compact manner for efficient loading into the systolic array calculation unit, greatly reducing the redundant calculations.
[0073] Fourth, Computation Mode Selection: The software stack of this embodiment integrates high-level instructions through the compiler to support flexible and diverse computation mode selections, including but not limited to: A, Load / Store mode: Responsible for data transfer between DDR and on-chip cache; B, Fetch mode: Schedules data to adapt to systolic array computing; C, GEMM mode: General matrix multiplication; D, SpMM mode: Sparse matrix multiplication; E, SegMul mode: Piecewise multiplication of causal masks; F, EleOps mode: Element-wise computation; G, SFU mode: Nonlinear function computation, such as EXP, SiLU, RMSNorm, etc.
[0074] By reasonably selecting the above different computing modes, this control method can be optimally configured according to the requirements of different computing stages, thereby further significantly improving the utilization rate of hardware computing resources.
[0075] In summary, the FPGA stacked processor acceleration system based on state-space duality provided in this embodiment has the following advantages: In terms of hardware microarchitecture: First, high efficiency of sparse computing: The sparse data predefined module effectively skips zero elements, significantly reducing redundant computations; Second, flexibility of multi-mode computing: The reconfigurable systolic array supports multiple computing modes through dynamic configuration, improving the utilization rate of hardware resources; Third, optimization of on-chip cache: The intelligent cache management strategy ensures the on-chip nature of SSD computing, significantly reducing off-chip memory access and improving inference efficiency and energy efficiency; Fourth, high efficiency of nonlinear computing: The efficient implementation of the function computing module further improves the overall computing performance.
[0076] In terms of software stack: First, significantly reduce memory overhead: Through the operation fusion framework and precise data allocation, the amount of intermediate tensor data is greatly reduced, reducing the memory bandwidth requirement; Second, improve computing efficiency: Data rearrangement and sparse computing mode selection significantly reduce zero-element redundant computations, improving the effective utilization rate of hardware resources; Third, improve inference speed and energy efficiency: The software scheduling strategy precisely controls the computing process, ensuring the efficient operation of FPGA hardware units and effectively improving the overall inference throughput and energy efficiency.
[0077] Next, for the FPGA stacked processor acceleration system based on state-space duality provided in this embodiment, specific implementation is carried out and described in conjunction with the accompanying drawings: As Figure 1As shown, it is a schematic diagram of the software stack and hardware microarchitecture of this system; Figure 1 On the left side, it shows the step composition of the software stack of the FPGA-overlay processor acceleration system (MambaOPU), including key steps such as data flow graph unfolding, operation fusion, data allocation, array mode selection, memory scheduling, and event generation. Data and instructions are sent from the host CPU to the acceleration system through the PCIe interface. On the right side, it shows the hardware microarchitecture composition of the FPGA-overlay processor acceleration system, including an instruction control unit module, an input cache and weight cache module, a sparse predefined data fetcher module, a reconfigurable systolic array module, a partial sum cache module, an element-wise operation cache module, a function calculation module, and an on-chip memory management module. Each module is interconnected through an efficient data path to achieve the efficient execution of SSD calculations.
[0078] As Figure 2 shown, Figure 2 On the left side of [], the sparse predefined data fetcher module is shown. Its core consists of an address sequence generator, a register bank, and a distributed queue, which precisely controls the data flow by generating a folded index to support different calculation modes (including sparse calculations in the Sa1, Sa3, and Sa5 stages). Figure 2 On the right side of [], it is a detailed structural schematic diagram of the processing element PE, showing the multiplexer switch, data register, and input / output signals of the calculation and control logic of the PE (including the current input CUR IN, weight input W IN, element-wise multiplication ELE MUL, cumulative multiplication CAS MUL, element-wise addition ELE ADD, cumulative calculation CAS MAC, bypass signal BYPASS, and the next input NEXT IN), reflecting the high flexibility and multi-calculation mode support ability of the processing element.
[0079] As Figure 3 shown, it clearly describes the data interaction process during the fusion calculation and element-wise calculation stages. The fusion calculation results are efficiently stored in the URAM-based cache and can further interact seamlessly with the element-wise cache. The left part of the figure shows the data interaction strategies between the fusion calculation and the element-wise calculation, including four cases: fusion to fusion, fusion to element-wise, element-wise to element-wise, and element-wise to fusion, to optimize the on-chip data flow and minimize off-chip memory access; Figure 3 The right part shows the access strategies and timelines of the on-chip cache and off-chip cache during different calculation stages (such as Sa1 - Sa6, Sr1 - Sr8, So1 - So4) in SSD calculations, reflecting the efficient data management scheme of this system and ensuring that the SSD calculation process is always efficiently completed in the on-chip cache.
[0080] As Figure 4As shown, the calculation process of sparse matrix multiplication (SpMM) is described, showing the specific operation process for tensor B ∈ R(1×3×3) and tensor C ∈ R(3×3×1). In each calculation cycle, each processing element (PE) loads two elements into register REG2 through the current input port and loads one element into register REG1 through the weight input port. Under the control of the multiplexer, the current computing unit PE[i,j] extracts a specific element from REG2 and performs a multiply-accumulate operation with the loaded element b_ij. The resulting partial product is passed vertically down to the next row of processing elements PE[i+1,j] in the subsequent cycle. At the same time, the input elements pre-stored in register REG2 are passed horizontally to the adjacent processing element PE[i,j+1]. Each processing element (PE) alternates the above calculations until all partial accumulations are completed. The final calculation result is output by the PE units in the last row.
[0081] As Figure 5 shown, each processing element loads the causal vector elements {1, ai} processed by exponentiation into register REG2 through the current input port. At the same time, the masked intermediate activation value {b} is loaded into register REG1 through the weight input port for performing element-wise multiplication. The partial product results generated inside the PE unit PE[i,j] are stored back into register REG2 or register REG3 according to specific requirements: If stored in register REG2, the result will be multiplied by the element b_ij newly loaded into the weight input port in the next calculation cycle; If stored in register REG3, the result is passed down to the next row of PE units PE[i+1,j] through the CAS_MUL path for cumulative multiplication with the element a_(i,j+1).
[0082] When each PE completes the calculation at a specific position, its calculation result is output to the external register through the bypass channel (BYPASS channel) controlled by control signal CTRL3 and control signal CTRL4.
[0083] The above embodiments are only one of the implementation manners capable of implementing the technical solution of the present invention. The scope of protection required by the present invention is not limited only by this embodiment, but also includes any changes, substitutions and other implementation manners that are easily conceivable by any person skilled in the art within the technical scope disclosed by the present invention.
[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. An FPGA overlay processor acceleration system based on state - space duality, characterized in that, including: A sparse predefined data fetcher module, configured to generate a valid data access address sequence according to a predefined sparse pattern, and extract valid data according to the valid data access address sequence; wherein, the valid data includes input data and weights with zero elements removed; A reconfigurable systolic array module, configured to receive the valid data transmitted by the sparse predefined data fetcher module, and perform dense matrix operations, sparse matrix operations, element-wise operations, and segmented multiplication operations on the valid data; A partial sum cache module, configured to store intermediate calculation results after the dense matrix operations, sparse matrix operations, and segmented multiplication operations are performed by the reconfigurable systolic array module, and perform cross-cycle accumulation operations on the intermediate calculation results; An element-wise operation cache module, configured to receive the element-wise operation results after the element-wise operations are performed by the reconfigurable systolic array module and the accumulated intermediate calculation results after the cross-cycle accumulation operations are performed by the partial sum cache module; A function calculation module, configured to perform non-linear operations on the element-wise operation results and the accumulated intermediate calculation results transmitted by the element-wise operation cache module to obtain non-linear operation results; An on-chip memory management module, configured to directly cache the non-linear operation results output by the function calculation module into the internal cache of the FPGA for subsequent calculations or return them to the DDR memory as the final results.
2. The FPGA overlay processor acceleration system based on state - space duality according to claim 1, characterized in that, It further includes an input cache and a weight cache module, configured to provide access addresses corresponding to the input data and weights to be calculated; The sparse predefined data fetcher module extracts valid data from the access addresses of the input cache and the weight cache module according to the valid data access address sequence.
3. The FPGA overlay processor acceleration system based on state - space duality according to claim 1, characterized in that, It further includes an instruction control unit module, configured to receive an instruction sequence transmitted by the host CPU, and control the sparse predefined data fetcher module, the reconfigurable systolic array module, the partial sum cache module, the element-wise operation cache module, the function calculation module, and the on-chip memory management module to work together according to the instruction sequence.
4. The FPGA overlay processor acceleration system based on state - space duality according to claim 1, characterized in that, The sparse predefined data fetcher module includes an address sequence generator, a sparse data register group, and a distributed queue unit; The address sequence generator generates a valid data access address sequence according to the predefined sparse pattern to remove zero-element data; extracts valid data according to the valid data access address sequence, and stores the valid data into the sparse data register group; The valid data is transmitted to the reconfigurable systolic array module through the distributed queue unit.
5. The FPGA overlay processor acceleration system based on state - space duality according to claim 1, characterized in that, The reconfigurable systolic array module includes a plurality of cascaded processing units; each processing unit includes an input selection multiplexer, a data input register, a control logic unit, a matrix multiply-accumulate unit, and an element-wise calculation unit.
6. The FPGA overlay processor acceleration system based on state - space duality according to claim 1, characterized in that, Both the element-wise operation cache module and the reconfigurable systolic array module and the element-wise operation cache module and the function calculation module are connected bidirectionally.
7. A control method for an FPGA overlay processor acceleration system based on state - space duality, characterized in that, The FPGA overlay processor acceleration system based on state space duality according to any one of claims 1-6, including: Through the sparse predefined data fetcher module, a valid data access address sequence is generated according to the predefined sparse pattern, and valid data is extracted according to the valid data access address sequence; wherein, the valid data includes input data and weights with zero elements removed; Receive valid data transmitted by the sparse predefined data acquirer module through the reconfigurable systolic array module, and perform dense matrix operations, sparse matrix operations, element-wise operations, and piecewise multiplication operations on the valid data; Store the intermediate calculation results after the reconfigurable systolic array module performs dense matrix operations, sparse matrix operations, and piecewise multiplication operations through the partial sum cache module, and perform cross-cycle accumulation operations on the intermediate calculation results; Receive the element-wise operation results after the reconfigurable systolic array module performs element-wise operations and the accumulated intermediate calculation results after the cross-cycle accumulation operations of the partial sum cache module through the element-wise operation cache module; Perform non-linear operations on the element-wise operation results and the accumulated intermediate calculation results transmitted by the element-wise operation cache module through the function calculation module to obtain non-linear operation results; Directly cache the non-linear operation results output by the function calculation module into the internal cache of the FPGA through the on-chip memory management module for subsequent calculations or return them to the DDR memory as the final results.
8. The control method for an FPGA overlay processor acceleration system based on state - space duality according to claim 7, characterized in that, Before the sparse predefined data acquirer module generates a valid data access address sequence according to the predefined sparse pattern and extracts valid data according to the valid data access address sequence, it further includes: Based on the data flow graph composed of various operators in the state space duality calculation process, decompose the state space duality calculation process into multiple parallel and reorderable basic operator sequences.
9. The control method of an FPGA superposition processor acceleration system based on state space duality according to claim 8, characterized in that, During the process that the reconfigurable systolic array module receives valid data transmitted by the sparse predefined data acquirer module and performs dense matrix operations, sparse matrix operations, element-wise operations, and piecewise multiplication operations on the valid data, it further includes: Apply the operation fusion framework to combine broadcast matrix multiplication and summation operations into vector matrix multiplication operations; move the piecewise multiplication in causal matrix generation to the element-wise operation stage; Select a calculation mode according to the operator type and generate a corresponding instruction sequence for the reconfigurable systolic array module to execute the corresponding operator operations.
10. The control method of an FPGA superposition processor acceleration system based on state space duality according to claim 9, characterized in that, When the operator type is dense matrix operation, use general matrix multiplication to optimize data block reading through the address mapping list; When the operator type is sparse matrix operation, use sparse matrix multiplication to eliminate redundant calculations of zero elements using the tensor rearrangement algorithm.
Citation Information
Patent Citations
Sparse neural network processor based on systolic array
CN110705703A
FPGA-based space-time diagram neural network accelerator structure
CN114548391A
Reducing systolic array power consumption using sparsity metadata
CN115526763A
Hyper-sampling for time domain amortization using hybrid precision convolutional neural networks
CN117859148A
Accelerator for accelerating reasoning process of unstructured sparse large language model
CN119443174A
Cited By
Redundant data parallel preprocessing optimization method based on heterogeneous computing platform
CN120508400A
A redundant data parallel preprocessing optimization method based on heterogeneous computing platform
CN120508400B
Field programmable gate array reasoning acceleration method based on data flow and related device
CN121503693A
Dataflow-based field programmable gate array inference acceleration method and related apparatus
CN121503693B