FPGA-based reconfigurable systolic array tensor ip core

CN119670829BActive Publication Date: 2026-09-29BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411864118.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2026-09-29
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

[0003]本发明的目的在于针对现有神经网络硬件加速技术中,运用统一的张量运算阵列造成生成式神经网络在Prefill、Decoding两阶段不适配导致的算法部署并行度差、硬件利用率低的问题,提出了一种基于FPGA的可重构脉动阵列张量IP核,此IP在原有二维脉动阵列的基础上引入阵列拆分重构、模式配置和激活广播机制,使边缘端单用户场景下的生成式模型推理过程中的两阶段兼容,实现较高的硬件利用率

Benefits of technology

[0016]通过脉动阵列重构与控制流、数据流平行流动,实现了Transformer Decoder两阶段(Prefill、Decoding)张量运算在边缘端单用户设备上统一的高效部署,提升了硬件利用率并降低了生成式模型的推理延时。通过两级控制信号双向流动,降低了核心控制器控制信号产生逻辑复杂度,减小了控制器控制信号扇出规模并精准控制每个运算单元在不同模式下的数据来源和运算行为。通过在Decoding阶段进行激活广播,减少二级缓存的数据加载和访存次数。通过张量指令解耦规划,解耦出的每组分块张量指令优先得到输出张量的完整行,利于张量核心与生成式语言模型中其它行向量操作融合、流水执行。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119670829B_ABST
    Figure CN119670829B_ABST
Patent Text Reader

Abstract

The application discloses a reconfigurable systolic array tensor IP core based on FPGA, and the architecture comprises a core controller, a shift distributor, a data processing array and four-stage data cache.The tensor core can simultaneously adapt to the general tensor multiplication and addition operation of the Prefill and Decode stages of the generated neural network through reconfigurable design.The core controller receives tensor multiplication instructions and configuration information, decouples the data flow, and generates multiple block tensor multiplication instructions.The shift distributor decodes the block tensor multiplication instructions, controls the data communication between the first and second level caches (Distributed RAM and Bram), generates the control signals of the data processing array, and enters the array flow together with the data.The data processing array connects different second level BRAM caches according to the received configuration information, and processes the data in a bubble-free systolic flow.The four-stage data cache realizes the multi-stage pingpong operation of the tensor operation core operation data loading, realizes the multi-stage pipeline design of the data from the off-chip HBM to the on-chip URAM, the URAM to the BRAM and the BRAM to the operation unit.For the edge generated network inference task of a single user, the core controller generates the data processing unit mode configuration signal in a fine-grained (tensor element level) manner, which is transmitted through the shift distributor and flows in parallel with the operation data, greatly improves the mode configuration efficiency and hardware resource utilization, reduces the high fan-out problem caused by the broadcast control, and designs the control flow by means of the systolic array structure characteristics, so that the complex control algorithm is realized with low resource occupation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of digital integrated circuits, electronic information and artificial intelligence, and in particular to a reconfigurable systolic array tensor IP core based on FPGA. Background Technology

[0002] Since the advent of Chat-GPT, generative large language models have gained widespread public attention, with an explosive increase in demand across various fields, including daily life, work, scientific research, and the military. As investment in the research and development of large language models has gradually increased and model training hardware has been upgraded, better foundational models have emerged from academia and industry, becoming increasingly sophisticated in applications such as machine translation, text summarization, dialogue systems, and question-answering systems. In the field of generative large language models, foundational models are generally composed of Transformer Decoder structures or their variations. The acceleration achieved when deployed on edge single-user devices is ultimately limited by the significant differences in activation sizes across different stages of the Decoder structure. The two main stages in the generative language model inference process—Prefill and Decoding—involve activation operations of two-dimensional tensors and one-dimensional tensor multiplication and addition, respectively. Using a uniform tensor operation array on edge single-user devices makes it difficult to efficiently deploy these two types of tensor operations, posing a significant challenge to the model's inference process. Therefore, we consider reconfigurable design for systolic tensor operation arrays with high data reusability, while supporting tensor operations of different scales and dimensions in two stages during inference, achieving higher hardware utilization on the same array size, thereby improving the unified inference speed of generative models. Summary of the Invention

[0003] The purpose of this invention is to address the problem of poor algorithm deployment parallelism and low hardware utilization caused by the mismatch between the prefill and decoding stages of generative neural networks due to the use of a unified tensor operation array in existing neural network hardware acceleration technologies. This invention proposes a reconfigurable systolic array tensor IP core based on FPGA. This IP introduces array splitting and reconstruction, mode configuration and activation broadcasting mechanisms on the basis of the original two-dimensional systolic array, so as to make the two stages of generative model inference process compatible in single-user scenarios at the edge and achieve higher hardware utilization.

[0004] To achieve the above-mentioned technical objectives, the technical solution implemented by this invention is as follows:

[0005] A reconfigurable systolic array tensor IP core based on FPGA is characterized by comprising: a core controller, a shift distributor, a data processing array, and a four-level data cache.

[0006] The core controller is used to receive tensor multiply-accumulate instructions and decouple them according to the data flow to generate multiple block tensor multiply-accumulate instructions.

[0007] The shift distributor is used to decode block tensor multiplication instructions. The decoded mode information controls the data communication between the first and second level caches (Distributed RAM and Bram), generates control signals for the data processing array, and flows into the array along with the data.

[0008] The data processing array is used to connect to different secondary BRAM caches according to the configuration information, process data in a bubble-free pulsating streaming manner, and return the output results to the secondary cache in different formats.

[0009] The four-level data cache enables multi-level ping-pong operations for loading data from tensor operations and core operations, and realizes a multi-level pipeline design for data from off-chip HBM to on-chip URAM, URAM to BRAM, and BRAM to the arithmetic unit.

[0010] Optionally, the core controller includes a two-level instruction stack. The first level caches the IP peripheral tensor operation instructions, and the second level caches each operation instruction after tensor decoupling. The decoupled instructions include the second-level instruction tensor operation dimension information, the input and output addresses of the second-level cache BRAM, the effective dimension of array call, output transpose enable, P / D mode configuration, and the start and end flag information of the first and second level instructions.

[0011] Optionally, the shift distributor performs periodic shifting of the control signal and the start / end flag. The shifting mode resonates parallel to the input mode of the pulsation array, reducing the complexity of generating control signals for the core controller. The shift distributor configures different connections between the secondary cache (BRAM) and the primary cache (Distributed RAM) of the arithmetic unit in different modes (Prefill or Decoding). During the Prefill phase, it controls the interaction between the left column and top row of the array and their respective secondary caches; during the Decoding phase, it controls the broadcast interaction between the reconfigurable primary cache of the arithmetic unit and its corresponding secondary cache. Synchronous control signals and data are transmitted in parallel to the arithmetic array using a common streaming pulsation mode.

[0012] Optionally, the data processing array performs tensor multiplication, bias accumulation, and residual accumulation in a bidirectional pulsating manner, and is equipped with a post-multiplication-accumulation quantization module to quantize intermediate results. It supports horizontal and vertical flow of the quantized accumulation results and performs matrix transposition without buffering and reshaping. The output direction of the multiplication-accumulation results is set to the left and top, reducing the delay in starting the output of the computation results.

[0013] Optionally, the four-level data cache caches four dimensions of data from bottom to top. The fourth level HBM caches the entire model parameters, the third level URAM caches single-layer model parameters and vocabulary and logit weights, the second level BRAM caches operator weights and fusion operator weights, and the first level Distributed RAM temporarily caches a set of activations, weights and intermediate results.

[0014] A bubble-free tensor operation data stream based on FPGA is characterized by using the block-based splitting results of complete tensor operations to guide the decoupling of complete tensor operation instructions, generating block tensor operation instructions. The corresponding control signals generated by these split instructions are transmitted in a synchronous data pulsating form to a reconfigurable pulsating array and flow within each operation unit. The control signals employ a two-stage flow: first, from the core controller to the shift distributor; second, from the shift distributor to the corresponding row or column of operation units. The control signals in the shift distributor control the secondary buffers corresponding to the array rows or columns, activating and weighted unicast and broadcast combinations of inputs at different stages of model inference. The control signals, accompanying the data flow, determine the computational behavior of the operation units and the flow paths and timing of input and output data.

[0015] The advantages and beneficial effects of the technical solution adopted in this invention are as follows:

[0016] By reconstructing the systolic array and allowing parallel flow of control and data, a unified and efficient deployment of the Transformer Decoder's two-stage (Prefill, Decoding) tensor operations on a single-user edge device is achieved, improving hardware utilization and reducing inference latency of generative models. Through bidirectional flow of two-level control signals, the logical complexity of the core controller's control signal generation is reduced, the controller's control signal fan-out scale is decreased, and precise control over the data source and computational behavior of each computational unit in different modes is achieved. Activation broadcasting during the Decoding stage reduces the number of data loading and memory accesses in the secondary cache. Through tensor instruction decoupling planning, each decoupled block of tensor instructions preferentially obtains the complete line of the output tensor, facilitating the fusion and pipelined execution of the tensor core with other row vector operations in the generative language model. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the overall structure of the present invention;

[0018] Figure 2 This is a schematic diagram of the core controller structure;

[0019] Figure 3 This is a schematic diagram of the shift distributor structure;

[0020] Figure 4This is a schematic diagram of the arithmetic unit structure; Detailed Implementation

[0021] like Figure 1 The overall structure shown is illustrated in this example, which relates to a reconfigurable systolic array tensor IP core based on FPGA. Its architecture includes: a core controller, a shift distributor, a data processing array, and a four-level data cache.

[0022] The core controller is as follows Figure 2 The core controller receives tensor multiplication and addition instructions and decouples them according to the data flow, generating multiple block tensor multiplication and addition instructions. Simultaneously, it generates corresponding instruction status control signals, which are transmitted in parallel with the instructions to the horizontal and vertical shift allocators. Specifically, the core controller architecture consists of one instruction flow and one control flow. The instruction flow decouples the IP call instructions cached in the first-level instruction cache stack, splitting large-scale matrix multiplication and addition into multiple block matrix multiplication and addition instructions according to the data flow format required by the output matrix. The second-level instruction information includes the tensor operation dimension information of the second-level instructions, the input and output addresses of the second-level cache BRAM, the valid dimension of the array call, output transpose enable, and P / D mode configuration. The decoupled second-level instructions are cached in the second-level instruction stack, and the multiple decoupled instructions are decoded and transmitted to the corresponding horizontal and vertical shift allocators. While the instruction flow executes, the control flow generates instruction status flags corresponding to the first-level instructions and the decoupled second-level instructions in the instruction status controller based on the instruction validity flag bits, controlling the issuance and submission of both levels of instructions.

[0023] The shift distributor is as follows Figure 3 The system uses decoded block tensor multiplication instructions to generate tensor operation control signals. Various mode information within these control signals governs data communication between the primary and secondary caches (Distributed RAM and BRAM), generating control signals for the data processing array. These signals flow pulsatingly into the array along with the data. Specifically, the decoded control signals are sent from the cache connected to the upper-level core controller to the control signal generation module, which generates: tensor end flag, operation flag, transpose flag, inference stage flag, and array start-up scale control signal. Simultaneously, the decoded control signals are transmitted from the secondary cache interface to the secondary cache, conveying the input data address and data length information of the decoupled instructions. The secondary cache output data and the generated control signals are packaged and merged within the data control synchronizer before being sent to the data processing array for pulsating flow.

[0024] Each computing unit in the data processing array is as follows: Figure 4The data processing array connects to different secondary BRAM caches based on configuration information, enabling bubble-free, pulsating streaming data processing. Specifically, the data processing array comprises multiple two-dimensional computation units, the number of which is configurable. Each computation unit includes two bidirectional buses: a horizontal inflow / outflow bus and a vertical inflow / outflow bus, supporting data pulsation between different units in the array. Each unit contains four horizontal and vertical buffers (top, bottom, left, and right) serving as intermediate buffers for bus transmission. Each computation unit receives flag signals assigned by its shift divider: a stage flag, a bias / residual flag, an enable flag, a last bit flag, and a transpose flag, thereby determining the corresponding computational form and the corresponding input / output flow direction. Each unit performs bidirectional pulsating tensor multiplication, bias accumulation, and residual accumulation, equipped with a post-multiply-accumulate quantization module, supporting horizontal and vertical output flow directions. Matrix transposition is performed without buffer reshaping, and the output direction of the multiply-accumulate result is set to the left and top, reducing the output startup delay of the computation result.

[0025] The four-level data cache implements multi-level ping-pong operations for loading data from tensor operations and core computations. The hardware supports a multi-level pipeline design for data transfer from off-chip HBM to on-chip URAM, URAM to BRAM, and BRAM to the computation unit. Specifically, the four levels of data cache cache data of different dimensions: the fourth level HBM caches the entire model parameters; the third level URAM caches single-layer model parameters and vocabulary / logit weights; the second level BRAM caches operator weights and fusion operator weights; and the first level Distributed RAM temporarily caches a set of activations, weights, and intermediate results. This four-level data cache enables a pipeline design for data loading, storage, and computation at different granularities, achieving a unified parameter caching and transfer mechanism.

Claims

1. A reconfigurable systolic array tensor IP core circuit structure based on FPGA, characterized in that, include: The core controller, shift distributor, data processing array, and four-level data cache; The core controller is used to receive tensor multiply-accumulate instructions and decouple them according to the data flow to generate multiple block tensor multiply-accumulate instructions; The shifter is used to decode block tensor multiplication instructions, decompose a tensor instruction into multiple multiplication and accumulation instructions with partial rows and partial columns, and use partial row and column multiplication and accumulation instructions as pattern information to control data flow; the decoded pattern information controls the data communication between the first-level cache Distributed RAM and the second-level cache Bram, generates control signals for the data processing array, and enters the array along with the data. The data processing array is used to connect to different secondary BRAM caches according to the configuration information, process data in a bubble-free pulsating streaming manner, and return the output results to the secondary cache in different formats. The four-level data cache is used to implement multi-level pingpong operations for loading data from tensor operation kernels, and to realize a multi-level pipeline design for data from off-chip HBM to on-chip URAM, URAM to BRAM, and BRAM to arithmetic units. The shift distributor performs periodic shifting of the control signal and the start and end flags. The shifting method resonates in parallel with the pulse array input method, reducing the complexity of generating control signals for the core controller. The shift distributor uses different connection configurations for the secondary cache BRAM and the primary cache Distributed RAM of the arithmetic unit according to different operation stages. In the Prefill stage, the left column and the top row of the control array interact with the secondary caches of the corresponding rows and columns. In the Decoding stage, the control array broadcasts and interacts with the reconfigurable arithmetic unit primary cache and its corresponding secondary cache. Synchronous control signals and data are transmitted to the arithmetic array in parallel in a common streaming pulsation mode.

2. The FPGA-based reconfigurable systolic array tensor IP core circuit structure according to claim 1, characterized in that, The core controller contains a two-level instruction stack. The first level caches the IP peripheral tensor operation instructions, and the second level caches each operation instruction after tensor decoupling. The decoupled instructions include the second-level instruction tensor operation dimension information, the input and output addresses of the second-level cache BRAM, the effective dimension of array call, output transpose enable, P / D stage configuration, and the start and end flag information of the first and second level instructions.

3. The FPGA-based reconfigurable systolic array tensor IP core circuit structure according to claim 1, characterized in that, The shift distributor performs periodic shifting of the control signal and the start and end flags. The shifting method resonates in parallel with the pulse array input method, reducing the complexity of generating control signals for the core controller. The shift distributor uses different connection configurations for the secondary cache (BRAM) and the primary cache (Distributed RAM) of the arithmetic unit according to the operation stage. In the Prefill stage, the left column and the top row of the control array interact with the secondary caches of the corresponding rows and columns. In the Decoding stage, the control array broadcasts and interacts with the reconfigurable arithmetic unit primary cache and its corresponding secondary cache. Synchronous control signals and data are transmitted to the arithmetic array in parallel in a common streaming pulsation mode.

4. The FPGA-based reconfigurable systolic array tensor IP core circuit structure according to claim 1, characterized in that, The data processing array performs tensor multiplication, bias accumulation, and residual accumulation in a bidirectional pulsating manner. It is equipped with a quantization module after multiplication and accumulation to quantize intermediate results. It supports the horizontal and vertical flow of the quantized accumulation results and performs matrix transposition without buffering and reshaping. The output direction of the multiplication and accumulation results is set to the left and top to reduce the output startup delay of the operation results.

5. The FPGA-based reconfigurable systolic array tensor IP core circuit structure according to claim 1, characterized in that, The four-level data cache caches data in four dimensions from bottom to top. The fourth level HBM caches the entire model parameters, the third level URAM caches single-layer model parameters and vocabulary and logit weights, the second level BRAM caches operator weights and fusion operator weights, and the first level Distributed RAM temporarily caches a set of activations, weights and intermediate results.

6. A bubble-free tensor computation data scheduling and computation control method based on FPGA, characterized in that, Using the circuit structure as described in claim 3, the complete tensor operation generated by the circuit structure in claim 1 is divided into blocks. The decomposition result guides the decoupling of the complete tensor operation instructions, generating block tensor operation instructions. The decomposition result guides the decoupling of the complete tensor operation instructions, generating block tensor operation instructions. The control signal generated during the decoupling process maintains the same pulsating form as the data and is transmitted in a streaming pulsating manner to the reconfigurable pulsating array as described in claim 4, and flows in each operation unit. The control signal adopts a two-stage flow method: the first stage flows from the core controller to the shift distributor, and the second stage flows from the shift distributor to the operation unit of the corresponding row or column. The control signals in the shift divider control the secondary buffers corresponding to the rows or columns of the array, and activate and weightedly combine unicast and broadcast inputs at different stages of model inference; the control signals that flow with the data determine the computational behavior of the arithmetic unit, the flow path and timing of input and output data.

Citation Information

Patent Citations

  • High-order separation serial shift complement multiplication and addition operation circuit and systolic array system

    CN117850738A

  • Memory efficient dropout, with reordering of dropout mask elements

    US11256987B1