Transform-oriented special accelerator system
By designing a dedicated accelerator system for Transformer, using multiple reconfigurable computing cores and distributed storage architectures, the problem of large computing and memory requirements for Transformer models is solved, and low latency and high performance computing effects are achieved.
Patent Information
- Application Number
- CN202510221189.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-06
AI Technical Summary
The Transformer model has huge demands in computing and memory, which has hindered the application of high-performance Transformer systems.
Design a dedicated accelerator system for Transformer, using multiple reconfigurable computing cores to interconnect through an on-chip network, and combining a distributed storage architecture to achieve efficient parallel computing.
Through multi-threaded parallel computing, efficiently utilize data parallelism characteristics to achieve low latency and high performance computing effects, reducing energy consumption and latency.
Smart Images

Figure CN120106158A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a dedicated accelerator system for Transformer. Background Art
[0002] The visual Transformer model has recently become a milestone in artificial intelligence. The algorithm has improved the performance of tasks such as machine translation and computer vision to a level that was previously unattainable. However, the Transformer model has strong performance but also requires a large amount of memory overhead and huge computing power. This has significantly hindered the application of high-performance Transformer systems.
[0003] Quantization is one of the important methods for deep learning model compression. It reduces the model size and computational complexity by reducing the data type of the parameters (such as using 8-bit or smaller integers in FP32 floating-point parameters). When the quantization bit width is large (such as 8 bits and above), the impact on the model accuracy is small because a larger range of values can be represented, and the loss of accuracy can be ignored. However, in order to further improve the model acceleration effect, you can consider using a smaller quantization bit width (such as less than 8 bits), which will bring a certain loss of accuracy. The choice of quantization bit width can be fixed, using a uniform quantization bit width for the entire model or a certain layer of the network. It can also be variable, using different quantization bit widths for different layers, which can balance between accuracy and performance. Quantization is an important means to achieve model compression and acceleration. Understanding the principles and methods of quantization is crucial for learning model compression and hardware acceleration.
[0004] The matrix parameters processed by the pruning algorithm will become a sparse matrix. That is, the number of non-zero elements (NZ) is much smaller than the number of zero elements. In order to skip the calculation of a large number of zero elements in the sparse matrix, various sparse matrix storage formats are proposed to compress the sparse matrix to reduce computing and storage resources. Typical formats include COO, CSR, BCSR and MBR. Among them, COO means storing non-zero elements in (row index, column index, value) tuples; CSR means storing in row pointers and column indices to facilitate row access; BCSR means the block version of CSR, which is suitable for vectorization and GPU operations; MBR means storing the degree of rows and columns to exploit the matrix repetitive patterns.
[0005] Sparse matrix multiplication operations account for most of the computational effort during Transformer inference. The accelerator by Hongwu Peng et al. has conducted in-depth research on sparse matrix multiplication accelerators. The sparse matrix multiplication accelerator combines the column-balanced block pruning algorithm and the CSCB compression method, and designs a dedicated pipeline by exploiting the parallelism between and within PEs. The scheme stores the compressed sparse weight matrix and its index matrix in BRAM, and extracts an element in the vector and a data block in the compressed sparse matrix according to the index value of the index matrix. Then, ordinary dense matrix multiplication is performed within a PE, and the result is sent to the accumulation module for accumulation. However, the scheme will traverse all block columns of the sparse matrix to obtain the final dot product output, which is inefficient. Summary of the invention
[0006] In view of this, the present invention proposes a dedicated accelerator system for Transformer. The present invention can realize efficient parallel computing functions, efficiently utilize the data parallel characteristics in the algorithm through multi-threaded parallelism, and realize low-latency computing.
[0007] To achieve the above object, the technical solution adopted by the present invention is:
[0008] A Transformer-oriented dedicated accelerator system includes a plurality of reconfigurable computing cores, which are interconnected to form an on-chip network; each reconfigurable computing core includes a memory, a computing unit, and a reconfiguration controller; the memory is used to store parameters, data, and instructions required for calculation; the reconfiguration controller is used to reconfigure the computing unit so that the computing unit has a specific operator function during calculation; the computing unit realizes calculation of a specific function after being configured by the reconfiguration control unit;
[0009] The dedicated accelerator system divides and packages the data into fixed-size data packets and attaches address routing information to each data packet so that each data packet has complete routing information; the data packet jumps between the reconfigurable computing cores according to the routing information and is finally sent to the target reconfigurable computing core.
[0010] Furthermore, the multiple reconfigurable computing cores are interconnected in a rectangular array, and each reconfigurable computing core is connected to at most four other reconfigurable computing cores.
[0011] Furthermore, the reconstruction controller configures the registers of the computing unit to be configured as a specific operator function, specifically in the following manner:
[0012] The reconfiguration controller reads instructions from the memory to obtain function and performance configuration information;
[0013] The reconfiguration controller performs a write operation on the function register of the computing unit according to the function and performance configuration information;
[0014] The reconfiguration controller reads the status register of the computing unit and determines whether the function register is written correctly;
[0015] The reconfiguration controller starts the computing unit, and the computing unit starts computing a specific function.
[0016] Compared with the prior art, the present invention has the following advantages:
[0017] (1) The algorithm mapping of the present invention is simple. Generally speaking, in the same Transformer dedicated accelerator system, the functions and performance of multiple reconfigurable computing cores are the same. After the compiler completes the algorithm analysis and optimization, there is no need to distinguish the specific physical core details, so that the algorithm mapping and matching work can be completed conveniently at the logical core level.
[0018] (2) The present invention realizes a high-computing computing system through a distributed computing architecture. The distributed storage architecture can shorten the physical distance between storage and computing units to the maximum extent, realize near-storage computing, thereby eliminating the "storage wall" bottleneck and realizing high-performance computing with low energy consumption and low latency through on-chip ultra-high storage bandwidth. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 Diagram of performing parallel computations on time-dependent data.
[0020] Figure 2 A schematic diagram of the structure of a dedicated accelerator system for Transformer.
[0021] Figure 3 Schematic diagram of distributed data storage in a dedicated accelerator system.
[0022] Figure 4 Schematic diagram of distributed computing in a dedicated accelerator system.
[0023] Figure 5 This is a schematic diagram of serializing and cutting the input image according to the requirements of the Transformer algorithm. DETAILED DESCRIPTION
[0024] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0025] In order to process data in parallel sequence, it is necessary to cut spatial data into N sequence blocks and perform parallel computing, and expand time series data into N sequence blocks according to the time dimension and perform parallel computing. Figure 1As shown in the figure, the Transformer neural network processes data in a parallel sequence. To perform parallel computing on time series related data, spatial data is cut into N sequence blocks and then parallel computing is performed, and time series data is expanded into N sequence blocks according to the time dimension and then parallel computing is performed. Therefore, a dedicated accelerator system is required to achieve efficient parallel computing functions.
[0026] A dedicated accelerator system for Transformer, such as Figure 2 As shown in the figure, through the NoC on-chip network technology, multiple reconfigurable computing cores are interconnected in a rectangular form according to the MESH topology, thereby constructing a large-scale on-chip distributed computing and storage architecture. It includes multiple reconfigurable computing cores, and the reconfigurable computing cores are interconnected through the NoC on-chip network structure inside the chip. The characteristic of the on-chip network interconnection structure is that the data bus transmission data is divided and packaged into fixed-size data packets, and address routing information is attached to each data packet, so that each data packet has complete routing information, can jump between reconfigurable computing cores, and finally sent to the target reconfigurable computing core. Based on the interconnection architecture, each reconfigurable computing core only needs to be connected to a maximum of four cores, which greatly simplifies the interconnection problem of a large number of reconfigurable computing cores. Through the design of a reasonable routing algorithm, the bandwidth resources of the Internet can be maximized.
[0027] The reconfigurable computing core adopts an AI-specific processor architecture that integrates storage and computing. Each reconfigurable computing core includes a memory, a computing unit, and a reconfiguration controller; among which:
[0028] The memory is used to store the parameters, data and instructions required for calculation;
[0029] The reconfiguration controller is used to reconfigure the computing unit so that the computing unit has a specific operator function during calculation;
[0030] The computing unit adopts reconfigurable computing technology and has multiple operator functions. It can be configured into a certain operator function by the reconfigurable control unit to achieve certain AI calculations.
[0031] The Transformer dedicated accelerator system has multi-core parallel computing features, and efficiently utilizes the data parallel characteristics in the algorithm through multi-threaded parallelism to achieve low-latency computing.
[0032] The reconfiguration is the process of configuring the registers of the computing unit to configure it for a certain computing function. The reconfiguration mainly includes the steps of writing function registers, reading status registers, and enabling. The specific method is as follows:
[0033] The reconfiguration controller reads instructions from the memory to obtain function and performance configuration information;
[0034] The reconfiguration controller performs a write operation on the function register of the computing unit according to the function and performance configuration information;
[0035] The reconfiguration controller reads the status register of the computing unit and determines whether the function register is written correctly;
[0036] The reconfiguration controller starts the computing unit, and the computing unit starts computing a specific function.
[0037] When the dedicated accelerator system is working, the input data to be calculated is processed in a parallel sequence manner. The parallel sequence manner refers to performing parallel calculation after cutting the spatial data into N sequence blocks by block, and performing parallel calculation after expanding the time series data into N sequence blocks by time dimension. Each data packet is attached with address routing information, so that each data packet has complete routing information. The data packet jumps between the reconfigurable computing cores according to the routing information and is finally sent to the target reconfigurable computing core.
[0038] The compiler completes the algorithm analysis and optimization, and completes the algorithm mapping and matching work. Among them, the compiler compiles the neural network algorithm model into algorithm binary data. The mapping and matching work does not need to distinguish the specific physical core details, and the functions and performance configured on multiple reconfigurable computing cores are the same.
[0039] like Figure 3 As shown in the figure, in the Transformer dedicated accelerator system, data (including input data, temporary data, and output data) is stored in a distributed storage architecture. The distributed storage architecture can shorten the physical distance between storage and computing units to the maximum extent, realize near-storage computing, thereby eliminating the "storage wall" bottleneck and achieving low-energy and low-latency high-performance computing through the ultra-high storage bandwidth on the chip.
[0040] The reconfigurable computing core can have multiple configurations, including functional configuration and performance configuration. The functional and performance configurations come from the instruction information stored in the memory, and the reconfigurable computing core can obtain the functional and performance configuration information by reading the instruction information. In terms of functional configuration, the operator types supported by the reconfigurable computing core are configurable, and certain operator functions can be added or deleted. In terms of performance configuration, both computing power and storage capacity are configurable. The typical computing power configuration is 4TOPS, and the typical storage capacity configuration is 4MB.
[0041] like Figure 4 As shown, in the Transformer dedicated accelerator system, a high-computing-power computing system is achieved through a distributed computing architecture.
[0042] like Figure 5As shown in the figure, an input image to be calculated is serialized and cut according to the requirements of the Transformer algorithm. The serialized cut image naturally has the characteristics of parallel computing. Each cut sub-image can input different reconfigurable computing cores, thereby realizing parallel computing with high computing power.
[0043] The present invention can realize Transformer multi-core parallel computing, and efficiently utilize the data parallel characteristics in the algorithm through multi-threaded parallelism to achieve low-latency computing.
Claims
1. A dedicated accelerator system for Transformer, characterized in that: The invention comprises a plurality of reconfigurable computing cores, wherein the plurality of reconfigurable computing cores are interconnected to form an on-chip network; each reconfigurable computing core comprises a memory, a computing unit and a reconfiguration controller; the memory is used to store parameters, data and instructions required for calculation; the reconfiguration controller is used to reconfigure the computing unit so that the computing unit has a specific operator function during calculation; the computing unit realizes calculation of a specific function after being configured by the reconfiguration control unit; The dedicated accelerator system divides and packages the data into fixed-size data packets and attaches address routing information to each data packet so that each data packet has complete routing information; the data packet jumps between the reconfigurable computing cores according to the routing information and is finally sent to the target reconfigurable computing core.
2. A dedicated accelerator system for Transformer according to claim 1, characterized in that: The multiple reconfigurable computing cores are interconnected in a rectangular array, and each reconfigurable computing core is connected to at most four other reconfigurable computing cores.
3. A Transformer-oriented dedicated accelerator system according to claim 1, characterized in that: The reconfiguration controller configures the registers of the computing unit to be configured as a specific operator function, specifically in the following manner: The reconfiguration controller reads instructions from the memory to obtain function and performance configuration information; The reconfiguration controller performs a write operation on the function register of the computing unit according to the function and performance configuration information; The reconfiguration controller reads the status register of the computing unit and determines whether the function register is written correctly; The reconfiguration controller starts the computing unit, and the computing unit starts computing a specific function.