Transformer accelerator based on mixed precision calculation array
By designing an architecture based on hybrid precision computing array in Transformer accelerator, the problems of traditional accelerators in accuracy and performance balance are solved, efficient Transformer network computing is achieved, and energy consumption and latency are reduced.
Patent Information
- Application Number
- CN202510063829.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
When traditional Transformer accelerators handle large-scale data and complex model training, they face the problem of balancing accuracy and performance, resulting in increased energy consumption and computing latency. Traditional hardware cannot be efficiently parallelized during large-scale computing, resulting in inefficient utilization of computing resources and storage bandwidth.
A Transformer accelerator based on a hybrid precision computing array is designed, including input memory, a top-level controller, a post-processing unit and a multiple hybrid precision computing array, each array contains multiple processing units, process the multiplication accumulation results through an addition tree and a cutoff module, and optimize data interaction through on-chip storage.
Optimize network accuracy and performance through hybrid precision calculation arrays, reduce off-chip on-chip data interaction, reduce power consumption, and complete end-to-end Transformer network calculations.
Smart Images

Figure CN119990209A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of integrated circuit technology, and in particular to a Transformer accelerator based on a mixed precision computing array. Background Art
[0002] Since its introduction, the Transformer model has become the mainstream model in fields such as natural language processing and computer vision. The Transformer relies on the self-attention mechanism and has significant advantages in processing long sequence data, but its computational complexity is relatively high, especially in the process of large-scale data and complex model training, the demand for hardware resources increases sharply.
[0003] In traditional technology, although hardware accelerators can improve their computing efficiency, they still face many problems in practical applications. The first is the balance between accuracy and performance. For accelerators such as central processing units (CPUs) or graphics processing units (GPUs), the operation process needs to handle a large amount of data transmission and calculations under standard floating-point calculation precision (such as FP32), resulting in significant energy consumption and computing delays. In addition, matrix operations in Transformers usually involve large-scale data access and frequent memory interactions, and traditional hardware often cannot achieve efficient parallelism when processing these large-scale calculations, resulting in inefficient use of computing resources and storage bandwidth. Summary of the invention
[0004] In view of the defects in the prior art, an object of the present invention is to provide a Transformer accelerator based on a mixed precision computing array.
[0005] In the first aspect, the present application implements a Transformer accelerator based on a mixed precision computing array, characterized in that it includes: an input memory, a top-level controller, a post-processing unit, and a plurality of mixed precision computing arrays, each of which contains a plurality of processing units, and the processing units are used to calculate the multiplication and accumulation results; the multiplication and accumulation results of the mixed precision computing array are transmitted to the output memory after being processed by the addition tree and the truncation module, and the output memory transmits the calculation results to the post-processing unit for Softmax operation, activation function and element multiplication and addition operation; wherein:
[0006] The input memory is used to store the data output by the post-processing unit and the output memory, and provide the mixed precision computing array with data for computing;
[0007] The top-level controller is used to control the interaction between the data in the off-chip dynamic random access memory and the data on the chip.
[0008] Optionally, the processing unit of the mixed precision computing array includes: a controller, a multiplier-accumulator of mixed computing, a weight memory, a part and a memory; wherein:
[0009] The controller is used to configure a calculation mode, wherein the calculation mode includes any one of floating point 8 bits, floating point 4 bits, integer 8 bits, and integer 4 bits;
[0010] The multiplication and accumulation device of the mixed calculation is used to perform multiplication and accumulation calculation according to the configured calculation mode;
[0011] The weight memory is used to configure the calculation weight;
[0012] The part and memory are used to store a part of the multiplication and accumulation result.
[0013] Optionally, the multiplier-accumulator of the mixed calculation is specifically used to perform the following steps:
[0014] Determine the calculation mode;
[0015] Perform floating point calculation and / or integer calculation according to the determined calculation mode;
[0016] When the floating point mode is used for calculation, the output result is obtained through mantissa calculation and error checking;
[0017] When integer mode is used for calculation, the output result is calculated by the integer multiplier;
[0018] Configure the multiplication related registers to delay one beat;
[0019] Use a carry-lookahead adder to perform addition operations to obtain multiplication and accumulation results;
[0020] Configure the addition-related registers to delay one beat;
[0021] Output the multiplication and accumulation results according to the determined amount calculation mode.
[0022] Optionally, the mantissa calculated in floating point mode and the addition calculated in integer mode share the same register.
[0023] Optionally, the addition tree is used to accumulate results output by the mixed precision calculation array, and the truncation module is used to truncate the output bit width of the accumulation result obtained by the addition tree.
[0024] Optionally, the input of the post-processing unit is passed from the output memory or the input memory and performs multiple operations to support the Transformer network operator.
[0025] Optionally, a transposition module is also included, which is used to transpose the output results of the mixed precision calculation array and transfer the output results to the mixed precision calculation array to be used as weights for the next calculation.
[0026] Optionally, when the number of the input memories is greater than 1, it is used to support parallel data storage and retrieval.
[0027] Optionally, when the number of the input memories is 2, the two input memories interact with the off-chip dynamic random access memory, the mixed precision computing array, and the output memory according to the instructions of the top-level controller to complete parallel data transmission and storage.
[0028] In a second aspect, an embodiment of the present application provides a computing device, comprising a Transformer accelerator based on a mixed-precision computing array as described in any one of the first aspects.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] The present invention is provided with a Transformer accelerator including: an input memory, a top-level controller, a post-processing unit, and several mixed precision computing arrays, each of which contains multiple processing units, and the processing units are used to calculate the multiplication and accumulation results; the multiplication and accumulation results of the mixed precision computing array are processed by the addition tree and the truncation module and then transmitted to the output memory, and the output memory transmits the calculation results to the post-processing unit for Softmax operation, activation function and element multiplication and addition operation; the input memory is used to store the data output by the post-processing unit and the output memory, and provide the mixed precision computing array with data for calculation; the top-level controller is used to control the interaction between the data in the off-chip dynamic random access memory and the on-chip data. Thus, the network accuracy and performance can be optimized; the interaction of off-chip and on-chip data can be effectively reduced through on-chip storage distribution, and the power consumption can be reduced; and the end-to-end Transformer network calculation can be completed through the analysis of the neural network and hardware configuration. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings in the following descriptions are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without creative work. By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes and advantages of the present invention will become more obvious:
[0032] Figure 1 A schematic diagram of the structure of a Transformer accelerator in an embodiment of the present invention;
[0033] Figure 2 This is a flowchart of the accelerated transformer neural network calculation in an embodiment of the present invention;
[0034] Figure 3 Schematic diagram of the processing flow of a mixed precision computing array in an embodiment of the present invention;
[0035] Figure 4 Schematic diagram of the circuit structure of the mixed precision computing array in an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0037] It should be noted that when a component is referred to as being "fixed to" or "disposed on" another component, it can be directly on the other component or indirectly on the other component. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element. In addition, connection can be used for fixing or for circuit connection.
[0038] It should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the embodiments of the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0039] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0040] The present invention aims to provide a Transformer accelerator with a multi-precision computing array, which balances accuracy and performance through multiple computing modes and optimizes energy consumption and latency through on-chip storage.
[0041] Exemplarily, a Transformer accelerator based on a mixed precision computing array provided in the present application may include: an input memory, a top-level controller, a post-processing unit, and several mixed precision computing arrays, each mixed precision computing array containing multiple processing units, and the processing units are used to calculate the multiplication and accumulation results; the multiplication and accumulation results of the mixed precision computing array are transmitted to the output memory after being processed by the addition tree and the truncation module, and the output memory transmits the calculation results to the post-processing unit for Softmax operation, activation function and element multiplication and addition operation; wherein: the input memory is used to store the data output by the post-processing unit and the output memory, and provide the mixed precision computing array with data for calculation; the top-level controller is used to control the interaction between the data in the off-chip dynamic random access memory (DRAM) and the data on the chip.
[0042] In an optional implementation, the addition tree is used to accumulate the results output by the mixed precision computing array, and the truncation module is used to truncate the output bit width of the accumulated results obtained by the addition tree.
[0043] In an optional embodiment, the input of the post-processing unit is transferred from the output memory or the input memory and multiple operations are performed to support the Transformer network operator.
[0044] In an optional embodiment, the processing unit of the mixed precision computing array includes: a controller, a multiplier-accumulator for mixed computing, a weight memory, and a partial and memory; wherein: the controller is used to configure the computing mode, and the computing mode includes: any one of floating point 8 bits, floating point 4 bits, integer 8 bits, and integer 4 bits; the multiplier-accumulator for mixed computing is used to perform multiplication and accumulation calculations according to the configured computing mode; the weight memory is used to configure the calculation weights; and the partial and memory is used to store a part of the multiplication and accumulation results.
[0045] In this embodiment, the processing unit can flexibly control the calculation mode by configuring the controller to meet the accuracy and performance requirements of the Transformer network.
[0046] In an optional implementation, when the number of the input memories is greater than 1, it is used to support parallel data storage and retrieval, thereby achieving efficient computing.
[0047] Exemplarily, when the number of the input memories is 2, the two input memories interact with the off-chip dynamic random access memory, the mixed precision computing array, and the output memory according to the instructions of the top-level controller to complete parallel data transmission and storage.
[0048] For example, Figure 1 FIG. 1 is a schematic diagram of the structure of the Transformer accelerator in an embodiment of the present invention; Figure 1 As shown, the accelerator includes: four mixed-precision multiplication-accumulation calculation arrays, each processing unit of which includes a mixed-precision multiplication-accumulation unit, a weight memory, a partial sum memory and a controller; an addition tree and truncation module; an output memory; a post-processing unit; a transposition module; two input memories; a top-level controller and input-output interface.
[0049] In an optional implementation, the multiplier-accumulator of the mixed calculation is specifically used to perform the following steps:
[0050] Determine the calculation mode;
[0051] Perform floating point calculation and / or integer calculation according to the determined calculation mode;
[0052] When the floating point mode is used for calculation, the output result is obtained through mantissa calculation and error checking;
[0053] When integer mode is used for calculation, the output result is calculated by the integer multiplier;
[0054] Configure the multiplication related registers to delay one beat;
[0055] Use a carry-lookahead adder to perform addition operations to obtain multiplication and accumulation results;
[0056] Configure the addition-related registers to delay one beat;
[0057] Output the multiplication and accumulation results according to the determined amount calculation mode.
[0058] In this embodiment, the mantissa calculated in floating-point mode and the addition calculated in integer mode share a register.
[0059] For example, Figure 2 is a flowchart of the accelerated transformer neural network calculation in an embodiment of the present invention; Figure 2As shown in the figure, it mainly includes three parts: the calculation of the linear layer network, the calculation of the attention layer network and the calculation of the feedforward network layer. The linear layer network mainly performs matrix multiplication calculations on the input and the three-layer weights to obtain the Q, K, and V matrices as the attention layer calculation input. In a single calculation head, the matrix K is transposed and then matrix multiplied with Q, and then the Softmax operation is performed to obtain the matrix P. Finally, the matrix multiplication of P and V is obtained as the output of the feedforward layer network input.
[0060] Combination Figure 1 , Figure 2 , Figure 1 The accelerator stores the input of the network through an input memory, and the weight values are stored in the weight memory of the computing array. After the multiplication and accumulation values are obtained through the mixed precision computing array, the value of Q is obtained through the addition tree and the truncation module and stored in another input memory through the output memory. When the matrix Q is transmitted, the matrix K calculation is performed in parallel, and then K is stored in the weight memory of the computing array through the transposition module. Next, the multiplication and accumulation of Q and the transposed matrix K can be calculated, and the result will be used for Softmax calculation by the post-processing unit. In this process, the result of the matrix V is obtained and stored in the weight memory for the multiplication and accumulation calculation of P and V. The remaining feedforward network calculation layer is mainly composed of multiplication and accumulation calculations, which are similar to the above calculations.
[0061] In an optional embodiment, a transposition module is further included, wherein the transposition module is used to transpose the output result of the mixed precision calculation array and transfer the output result to the mixed precision calculation array to be used as the weight for the next calculation.
[0062] For example, Figure 3 Schematic diagram of the processing flow of a mixed precision computing array in an embodiment of the present invention; Figure 4 Schematic diagram of the circuit structure of the mixed precision computing array in an embodiment of the present invention.
[0063] Combination Figure 3 , Figure 4 As shown in FIG. 1 , when performing multiplication and accumulation calculations through a mixed precision calculation array, the multiplication and accumulation calculation steps of a single processing unit are as follows:
[0064] Step 1: Both multiplication inputs are 16 bits wide. Each input can be regarded as two 8-bit data or four 4-bit data. The preprocessing module will close other calculation mode paths according to the calculation mode selection. Floating-point calculation will process the exponent, mantissa and sign bit; integer calculation will be directly sent to the multiplication module.
[0065] Step 2: The multiplication calculation will be divided into two paths to calculate floating point and integer.
[0066] The integer path is calculated and output by the integer multiplier, and the exponential module obtains the output through mantissa calculation and error checking;
[0067] Step 3: The pipeline will make the multiplication-related registers beat once to meet the frequency requirements.
[0068] Step 4: The addition preprocessing also includes gating logic to reduce power consumption. The mantissa in floating-point mode and the addition in integer mode share a register to reduce resources.
[0069] Step 5: The addition module obtains output through the carry-lookahead adder, and integer and floating-point outputs are obtained according to different modes.
[0070] Step 6: The pipeline will make the addition-related registers beat to meet the frequency requirements.
[0071] Step 7: Select the final multiply-accumulate output according to the mode.
[0072] In this embodiment, a Transformer accelerator is provided, which includes: an input memory, a top-level controller, a post-processing unit, and a plurality of mixed precision computing arrays. Each mixed precision computing array includes a plurality of processing units, and the processing units are used to calculate the multiplication and accumulation results; the multiplication and accumulation results of the mixed precision computing array are processed by the addition tree and the truncation module and then transmitted to the output memory, and the output memory transmits the calculation results to the post-processing unit for Softmax operation, activation function and element multiplication and addition operation; the input memory is used to store the data output by the post-processing unit and the output memory, and provide the mixed precision computing array with the data for calculation; the top-level controller is used to control the interaction between the data in the off-chip dynamic random access memory and the data on the chip. Thus, the network accuracy and performance can be optimized; the interaction between the off-chip and on-chip data can be effectively reduced through the on-chip storage distribution, and the power consumption can be reduced; and the end-to-end Transformer network calculation can be completed through the analysis of the neural network and the hardware configuration.
[0073] An embodiment of the present application also provides a computing device, including the Transformer accelerator based on the mixed precision computing array described in the above embodiments.
[0074] In this embodiment, the processing unit can support four computing modes: floating point 8-bit, floating point 4-bit, integer 8-bit, integer 4-bit. In order to complete high-throughput computing, a two-stage pipeline is inserted, and multiplication preprocessing, multiplication, pipeline, addition preprocessing, addition, pipeline and output are completed according to the mode selection.
[0075] In this embodiment, a mixed-precision computing array is designed to optimize network accuracy and performance based on the relationship between quantization calculation and accuracy of the transformer network. On-chip storage distribution is used to effectively reduce the interaction of off-chip and on-chip data and reduce power consumption. In addition, end-to-end transformer network calculations can be completed through analysis of neural networks and hardware configuration.
[0076] The above is the core idea of the present invention. In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in combination with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0077] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same and similar parts between the embodiments can be referred to each other. The above description of the disclosed embodiments enables professionals and technicians in this field to implement or use the present invention. Various modifications to these embodiments will be obvious to professionals and technicians in this field, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in this article, but will comply with the widest range consistent with the principles and novel features disclosed herein.
[0078] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art may make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.
Claims
1. A Transformer accelerator based on a mixed precision computing array, characterized in that: include: An input memory, a top-level controller, a post-processing unit, and a plurality of mixed precision computing arrays, each of which includes a plurality of processing units, and the processing units are used to calculate multiplication and accumulation results; the multiplication and accumulation results of the mixed precision computing arrays are processed by an addition tree and a truncation module and then transmitted to an output memory, and the output memory transmits the calculation results to the post-processing unit for Softmax operation, activation function and element multiplication and addition operation; wherein: The input memory is used to store the data output by the post-processing unit and the output memory, and provide the mixed precision computing array with data for computing; The top-level controller is used to control the interaction between the data in the off-chip dynamic random access memory and the data on the chip.
2. The Transformer accelerator based on mixed precision computing array according to claim 1, characterized in that: The processing unit of the mixed precision computing array includes: a controller, a multiplier-accumulator of mixed computing, a weight memory, a part and a memory; wherein: The controller is used to configure a calculation mode, wherein the calculation mode includes any one of floating point 8 bits, floating point 4 bits, integer 8 bits, and integer 4 bits; The multiplication and accumulation device of the mixed calculation is used to perform multiplication and accumulation calculation according to the configured calculation mode; The weight memory is used to configure the calculation weight; The part and memory are used to store a part of the multiplication and accumulation result.
3. The Transformer accelerator based on mixed precision computing array according to claim 2, characterized in that: The multiplier-accumulator of the mixed calculation is specifically used to perform the following steps: Determine the calculation mode; Perform floating point calculation and / or integer calculation according to the determined calculation mode; When the floating point mode is used for calculation, the output result is obtained through mantissa calculation and error checking; When integer mode is used for calculation, the output result is calculated by the integer multiplier; Configure the multiplication related registers to delay one beat; Use a carry-lookahead adder to perform addition operations to obtain multiplication and accumulation results; Configure the addition-related registers to delay one beat; Output the multiplication and accumulation results according to the determined amount calculation mode.
4. The Transformer accelerator based on mixed precision computing array according to claim 3, characterized in that: The mantissa calculated in floating-point mode and the addition calculated in integer mode share the same register.
5. The Transformer accelerator based on mixed precision computing array according to any one of claims 1 to 4, characterized in that: The addition tree is used to accumulate the results output by the mixed precision calculation array, and the truncation module is used to truncate the output bit width of the accumulation result obtained by the addition tree.
6. The Transformer accelerator based on a mixed precision computing array according to any one of claims 1 to 4, characterized in that: The input of the post-processing unit is passed from the output memory or the input memory and performs various operations to support the Transformer network operator.
7. The Transformer accelerator based on a mixed precision computing array according to any one of claims 1 to 4, characterized in that: It also includes a transposition module, which is used to transpose the output result of the mixed precision calculation array and transfer the output result to the mixed precision calculation array to be used as the weight for the next calculation.
8. The Transformer accelerator based on a mixed precision computing array according to any one of claims 1 to 4, characterized in that: When the number of the input memories is greater than 1, it is used to support parallel data storage and retrieval.
9. The Transformer accelerator based on mixed precision computing array according to claim 8, characterized in that: When the number of the input memories is 2, the two input memories interact with the off-chip dynamic random access memory, the mixed precision computing array, and the output memory according to the instruction of the top-level controller to complete parallel data transmission and storage.
10. A computing device, characterized in that: A Transformer accelerator based on a mixed precision computing array comprising any one of claims 1 to 9.