An implementation system and method based on a pulsating array self-attention mechanism

By adopting PE calculation unit with feedback mechanism and parallel input on the ZYNQ development platform, a self-attention calculation unit based on pulsating arrays is constructed, which solves the problems of slow computing speed and poor parallelism in the prior art, and achieves high calculation accuracy and high efficiency self-attention calculation.

CN116187407BActive Publication Date: 2025-06-13XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310215066.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-08
Publication Date
2025-06-13
Estimated Expiration
2043-03-08

AI Technical Summary

Technical Problem

When the prior art realizes the self-attention mechanism on the ZYNQ development platform, the calculation speed is slow and the parallelism is poor, and the feedback signal is lacking, resulting in low calculation accuracy.

Method used

The self-attention calculation unit based on the pulsating array is adopted, and the PE calculation unit with a feedback mechanism is used to realize the hardware computing design of the self-attention mechanism through parallel input.

Benefits of technology

The calculation bandwidth of the pulsating array is improved, high computing accuracy and high working efficiency are achieved, and high efficiency calculations can be completed while high computing bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116187407B_ABST
    Figure CN116187407B_ABST
Patent Text Reader

Abstract

An implementation system and method based on a systolic array self-attention mechanism. The system includes: a preprocessing unit that processes and transmits an input matrix; a PE computing unit that caches, performs multiplication operations, and accumulates operations on the processed data, and outputs the multiplication and addition results; a systolic array that horizontally and vertically transmits the multiplication and addition results, and repeats the above operations until the multiplication and addition calculations are completed; a matrix multiplication computing unit that obtains the product matrix of matrix and matrix K; a self-attention mechanism module that obtains the module output Attention(Q, K, V); the method is an application of an implementation system based on a systolic array self-attention mechanism. By using a PE computing unit with a feedback mechanism to form a systolic array module and adopting a parallel input method, the present invention is applied to the self-attention mechanism algorithm, realizing the hardware computing design of the self-attention mechanism, and having the characteristics of improving the computing bandwidth of the systolic array, achieving higher computing accuracy while having a high computing bandwidth, and improving work efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural networks, and particularly relates to an implementation system and method based on a systolic array self-attention mechanism. Background Art

[0002] Hardware-accelerated computing is a technology that uses dedicated hardware, such as FPGAs, ASICs, etc., to complete accelerated computing tasks to improve computing efficiency and performance. It is widely used in fields such as artificial intelligence, big data analysis, cryptography, and communication, among which deep learning is one of the important applications. Currently, hardware-accelerated computing architectures include FPGAs, ASICs, and GPUs, etc. Each architecture has its own characteristics, advantages, and disadvantages, and needs to be selected according to the application scenario. To facilitate the use of hardware-accelerated computing, many high-level programming models and tools have also been developed. The design and optimization of hardware-accelerated computing algorithms are one of the important aspects of hardware-accelerated computing research. Some optimization algorithms include distributed computing, parallel computing, pipelined computing, etc. Recently, cloud computing platforms have also started to support hardware-accelerated computing. In the future, with the continuous development and progress of hardware technology, hardware-accelerated computing is expected to play a role in a wider range of application scenarios.

[0003] Systolic array computing is a special parallel computing method implemented through a hardware circuit. The following are the general steps: First, a pulse signal needs to be defined to trigger the calculation. This pulse signal can be generated by a timer or triggered by an external signal; according to the algorithm of systolic array computing, the corresponding circuit is designed, including a data input circuit, an arithmetic circuit, a control circuit, etc. Among them, the arithmetic circuit can include adders, multipliers, logic gates, etc.; since systolic array computing is a parallel computing method, the systolic array circuit needs to be replicated multiple times to support the parallel computing of multiple data; multiple systolic array circuits are connected together to form a systolic array computing system. When connecting, issues such as control signals and data transmission need to be considered; after the design and connection of the hardware circuit are completed, debugging and testing are required to ensure the correctness and reliability of the system.

[0004] ZYNQ is a series of SoC (System on Chip) products launched by Xilinx. It integrates an ARM Cortex-A9 processor and an FPGA, featuring high performance, low power consumption, programmability, and scalability in processing capabilities. Among them, the Cortex-A9 processor provides high-performance application processing capabilities, and the programmable logic (FPGA) offers high flexibility and scalability, which can be optimized and adjusted according to different application scenarios. In addition, the ZYNQ chip is also equipped with a variety of interfaces and peripherals, such as PCIe, USB, Ethernet, SD card, etc., facilitating the connection of various peripheral devices. ZYNQ is mainly used in the fields of high performance, low power consumption, real-time computing, and signal processing, such as machine vision, drones, industrial control, network communication, etc. In addition, the ZYNQ series products also offer different levels of models to meet the needs of different application scenarios. In short, the integration advantages and rich peripheral interfaces of the ZYNQ series chips enable it to have broad application prospects in various fields.

[0005] In the prior art, implementing the self-attention mechanism on the ZYNQ development platform is achieved by calling a large number of DSP computing units to perform multiplication and addition operations to complete matrix multiplication operations, thereby implementing the attention algorithm. The data input of this method is serial input, so the calculation speed is slow and the parallelism is poor.

[0006] Compared with the implementation scheme of parallel expansion for implementing multiplication operations and the implementation scheme of systolic matrix for implementing multiplication operations, when the number of channels and the size of the input and output matrices increase, the parallel expansion scheme will cause long data packets and increase the situation of fan-in and fan-out data paths, resulting in a lower speed of matrix multiplication operations.

[0007] For the patent application with the application number 【202211216188.4】 and the title "Systolic Array, Systolic Array System and Its Operation Method, Device, Storage Medium", by determining the working mode indicated by the received working instruction according to the received working instruction, when the working mode is the sorting mode, after different configuration values are assigned to the control registers of the first basic operation unit and the second basic operation unit in the systolic array through the sorting control signal sent by the array controller, the feature data of the feature buffer and its corresponding label data are gradually input into the systolic array in groups for synchronous sorting operations, and after the sorting is completed, the sorted feature data and its corresponding synchronous label data are output via the output buffer and returned to the system bus; however, this patent application lacks a feedback signal, cannot determine whether the operation of the previous computing unit is completed, and cannot achieve the effect of real-time verification. Therefore, it has the disadvantages of poor parallelism and low calculation accuracy. Summary of the Invention

[0008] To overcome the shortcomings of the above-mentioned existing technologies, the purpose of the present invention is to provide a system and method for implementing a self-attention mechanism based on a systolic array. A systolic array module is composed of PE computing units with a feedback mechanism, and at the same time, a parallel input method is adopted and applied to the self-attention mechanism algorithm, realizing the hardware computing design of the self-attention mechanism, which has the characteristics of improving the computing bandwidth of the systolic array, achieving a relatively high computing accuracy while having a high computing bandwidth, and improving work efficiency.

[0009] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0010] A system for implementing a self-attention mechanism based on a systolic array, comprising:

[0011] A preprocessing unit: processes the data input row-direction matrix data data1 and the input column-direction matrix data data2 to obtain the misaligned row vector data data3 and column vector data data4, and transmits the row vector data data3 and the column vector data data4 to the PE computing unit;

[0012] A PE computing unit: caches the data data3 obtained in the preprocessing unit in the cache cache1, caches the data data4 obtained in the preprocessing unit in the cache cache2, multiplies the data data3 cached in the cache cache1 and the data data4 cached in the cache cache2 through a multiplier to obtain the data data5, and then caches the data data5 in the cache cache3. At the same time, the cache cache3 sends a feedback signal to the cache cache1 and the cache cache2 to judge whether the transmission of the data data3 and the data data4 is completed. If completed, the data data5 is transmitted to the adder to be added to the data in the adder of the previous PE computing unit to obtain the data data6. If not completed, wait for the transmission of the data data3 and the data data4 to be completed and then perform multiplication calculation to obtain the data data5, and transmit the calculation result data data5 to the adder to be added to the data in the adder of the previous PE computing unit to obtain the data data6. The data data6 is transmitted to the adder in the next PE computing unit, and when the next clock cycle arrives, the data data3 and the data data4 are transmitted to the row input end dim1 and the column input end dim2 of the next PE computing unit in the row direction and the column direction, and repeat the above operations;

[0013] A systolic array: composed of multiple PE computing units, horizontally and vertically transmits the multiplication and addition results output by the PE computing units, and continuously repeats the above operations until the multiplication and addition calculations of all elements in the matrix are completed;

[0014] Matrix multiplication calculation unit: The matrices Q and K are calculated using a systolic array through a data shaping module. After passing the calculation result through the data shaping unit, it is cached in cache3. The cached result is output through the data output module, and the product matrix of the Q matrix and the K matrix is obtained.

[0015] Self-attention mechanism module: The input matrices Q and K are used to obtain the product matrix through the preloading unit and the matrix multiplication calculation unit. Then, the product matrix is normalized through the SoftMax calculation unit to obtain the data SoftMax<Q, K>. The obtained data SoftMax<Q, K> is multiplied with the input matrix V through the matrix multiplication calculation unit to obtain the data Attention(Q, K, V), and the data Attention(Q, K, V) is output through the data output module.

[0016] The described preprocessing unit: includes a data cache array and a data shaping and loading module;

[0017] Data cache array: used to cache the input row-direction matrix data data1 and the input column-direction matrix data data2;

[0018] Data shaping and loading module: used to divide the input row matrix and input column matrix saved in the data cache array into horizontal row vector forms and vertical column vector forms according to the row slicing or column slicing method. Then, the sliced horizontal row vectors and vertical column vectors are sent into the shift register, and by using the method of delaying beats, each row or each column of the sliced row vectors and column vectors is misaligned to obtain new row vectors or column vectors after misalignment. The misaligned new row vectors or column vectors are transmitted to the PE calculation unit.

[0019] The adder in the PE calculation unit will simultaneously send a feedback signal to the multiplier to determine whether the multiplication calculation is completed. If completed, data5 is added to data6 in the adder of the previous PE calculation unit to obtain data6 of this PE calculation unit. If not completed, wait for the multiplication calculation to complete to obtain data5, and transmit the calculation result data5 to the adder to add it to the data6 in the adder of the previous PE calculation unit to obtain data6 of this PE calculation unit.

[0020] The described SoftMax calculation unit: used to normalize the data output by the matrix multiplication calculation unit.

[0021] A method for implementing a self-attention mechanism based on a systolic array includes the following steps:

[0022] Step 1: The Q matrix and the input K matrix of the input data are multiplied by a matrix multiplication calculation unit to obtain the product matrix of the Q matrix and the K matrix;

[0023] Step 2: According to the product matrix obtained in Step 1, a normalization calculation is performed through a SoftMax calculation unit to obtain the data SoftMax<Q, K>;

[0024] Step 3: According to the data SoftMax<Q, K> obtained in Step 2 and the input matrix V, repeat Step 1 to obtain the data Attention(Q, K, V).

[0025] The specific process of Step 1 is as follows:

[0026] Step 1.1: Cache the input row matrix Q and the input column matrix K through a data cache array;

[0027] Step 1.2: Through a data shaping and loading module, the input row matrix Q and the input column matrix K cached in Step 1.1 are arranged out of position in a row or column manner to obtain the vector data data3 arranged out of position in a row and the vector data data4 arranged out of position in a column, and then the data data3 and the data data4 are transmitted to the PE calculation unit;

[0028] Step 1.3: Cache the data data3 obtained in the preprocessing unit in the cache cache1 through the row input port dim1, cache the data data4 obtained in the preprocessing unit in the cache cache2 through the column input port dim1, multiply the data data3 cached in the cache cache1 and the data data4 cached in the cache cache2 through a multiplier to obtain the data data5. At the same time, the multiplier transmits the data data5 to the adder through the cache cache3 for an addition operation with the data data5 in the previous PE calculation unit to obtain the data data6 of this PE calculation unit, and transmits the obtained multiplication and addition result data data6 to the adder of the next PE unit; Through multiple PE calculation units, the multiplication and addition results output by the PE calculation unit are transmitted horizontally and vertically, and the above operations are continuously repeated until all elements in the matrix Q matrix K are multiplied and added; After passing the matrix Q and the matrix K through the data shaping module, the product matrix of the Q matrix and the K matrix is obtained.

[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0030] 1. Based on the multi-head self-attention module in the Transformer neural network, the present invention proposes a self-attention calculation unit based on a systolic array, and finally realizes self-attention calculation, which has the characteristics of high calculation accuracy, fast calculation speed and large throughput.

[0031] 2. The present invention proposes a PE computing unit with a feedback mechanism. Compared with the multiplication and addition operations of traditional PE computing units, the accuracy of the calculation results has been greatly improved.

[0032] 3. The present invention proposes a PE computing unit used in conjunction with a Transformer neural network hardware accelerator, which has high parallelism, accurate calculation results, and fast calculation speed. Therefore, it can perfectly replace the operation of using a cyclic array and a DSP for multiplication and addition calculations in traditional matrix operations to implement matrix multiplication operations.

[0033] 4. The present invention, through a data shaping and loading module and a matrix multiplication operation module, inputs each row or each column of the matrix simultaneously to achieve the characteristics of improving parallel input and output, thereby increasing the data throughput, and has the characteristics of fast calculation speed, low latency, and simple structure.

[0034] 5. The present invention realizes matrix multiplication operations based on a systolic array and finally applies it to the self-attention algorithm. Therefore, it can accurately and quickly complete self-attention calculation operations. Compared with traditional calculation methods, it is faster and has higher accuracy, improving the research efficiency of self-attention calculations.

[0035] 6. The matrix multiplication computing unit based on matrices in the present invention improves the parallel input of the circuit and the calculation accuracy compared with existing matrix multiplications.

[0036] 7. The present invention selects a systolic matrix to implement matrix multiplication operations, so that the data path of long data packets and multi-fan-in and fan-out is transformed into data communication between each PE. Therefore, it has the characteristics of high parallelism, simple structure, high operating frequency, and high calculation accuracy. Brief Description of the Drawings

[0037] Figure 1 is a schematic diagram of the preprocessing of the input matrix data of the present invention.

[0038] Figure 2 is the internal structure diagram of the PE computing unit of the present invention.

[0039] Figure 3 is a schematic diagram of the internal data flow direction of the systolic array of the present invention.

[0040] Figure 4 is the flowchart of the matrix multiplication unit of the systolic array of the present invention.

[0041] Figure 5 is a schematic diagram of matrix multiplication based on the systolic array of the present invention.

[0042] Figure 6It is a schematic diagram of the hardware algorithm of the self-attention mechanism based on the present invention.

[0043] Figure 7 It is a flowchart of the method of the present invention. Specific implementation manner

[0044] The working principle of the present invention will be described in detail below in conjunction with embodiments.

[0045] A method for implementing a self-attention mechanism based on a systolic array, comprising:

[0046] See Figure 1 , a preprocessing unit: processes the input row-direction matrix data data1 and the input column-direction matrix data data2 to obtain the misaligned row vector data data3 and column vector data data4, and transmits the row vector data data3 and column vector data data4 to the PE calculation unit;

[0047] The preprocessing unit described above: includes a data cache array and a data shaping and loading module;

[0048] The data cache array: is used for caching the input row-direction matrix data data1 and the input column-direction matrix data data2;

[0049] The data shaping and loading module: is used to divide the input row matrix and input column matrix saved in the data cache array into horizontal row vector forms and vertical column vector forms according to the row slicing or column slicing method, and then send the sliced horizontal row vectors and vertical column vectors into the shift register. By using the method of delaying beats, each row or each column of the sliced row vectors and column vectors is misaligned to obtain new misaligned row vectors or column vectors, and the misaligned new row vectors or column vectors are transmitted to the PE calculation unit; inputting each row or each column of the matrix simultaneously to achieve the characteristics of improving parallel input and output, thereby improving the data throughput, and having the characteristics of fast calculation speed, low latency, and simple structure.

[0050] See Figure 2, PE calculation unit: Cache the data data3 obtained in the preprocessing unit in cache cache1, and cache the obtained data data4 in cache cache2. Multiply the data data3 cached in cache cache1 and the data data4 cached in cache cache2 through a multiplier to obtain data data5. Then, cache the data data5 in cache cache3. At the same time, cache cache3 sends a feedback signal to cache cache1 and cache cache2 to determine whether the transmission of data data3 and data4 is completed. If it is completed, transmit the data data5 to the adder to add it to the data in the adder of the previous PE calculation unit to obtain data data6. If it is not completed, wait for the transmission of data data3 and data4 to be completed and then perform multiplication calculation to obtain data data5, and transmit the calculation result data data5 to the adder to add it to the data in the adder of the previous PE calculation unit to obtain data data6. At the same time, the adder sends a feedback signal to the multiplier to determine whether the multiplication calculation is completed. If it is completed, add the data data5 to the data data6 in the adder of the previous PE calculation unit to obtain the data data6 of this PE calculation unit. If it is not completed, wait for the multiplication calculation to be completed to obtain data data5, and transmit the calculation result data data5 to the adder to add it to the data data6 in the adder of the previous PE calculation unit to obtain the data data6 of this PE calculation unit. After that, transmit the data data6 to the adder in the next PE calculation unit. When the next clock cycle arrives, transmit the data data3 and data data4 to the row input port dim1 and the column input port dim2 of the next PE calculation unit in the row direction and the column direction, and repeat the above operations. Compared with the multiplication and addition calculation operations of traditional PE calculation units, the accuracy of the calculation results has been greatly improved;

[0051] The specific calculation process of the PE calculation unit is divided into three stages: input, calculation, and accumulation output. In the input stage, each PE calculation unit reads data from the output of the previous PE calculation unit and stores it in its own cache. If a PE calculation unit is at the starting column position, it writes the matrix data input in the column direction into the cache array. If a PE calculation unit is at the starting row position, it determines whether it needs to read the matrix data in the row direction from the cache array in the row direction according to the operating state of the circuit. When all PE calculation units have completed caching, the systolic array enters the calculation stage. In the calculation stage, the PE calculation units at non-starting positions perform the operation of multiplying and adding the row input data and the column input data, and then transfer the result to the subsequent PE calculation units. Since the PE calculation unit at the starting column position does not need to obtain data from the previous PE calculation unit, it only needs to multiply the matrix in the row direction cache unit and the matrix in the column direction cache unit, and transfer the calculation result to the intermediate cache unit. The specific calculation process is described as follows: A feedback signal path is led out from cache3 and connected to cache1 and cache2 respectively to determine whether the data input is completed. If it is completed, it enters the multiplication stage; otherwise, it waits until the input is completed before performing the multiplication calculation. An adder leads out a feedback signal path and connects it to the multiplier to determine whether the multiplication operation is completed. If it has been completed, the multiplication calculation result is sent to the adder through cache3 and added to the calculation result of the adder in the previous PE calculation unit. If it has not been completed, the completed calculation result is cached in cache3, and after waiting for the multiplication operation to end, the multiplication calculation result is input to the adder through cache3 and added to the calculation result of the adder in the previous PE calculation unit. When the calculation is completed, it enters the next stage. In the accumulation output stage, the PE calculation units at non-end positions transfer the results in the cache to the next adjacent PE calculation unit. The PE calculation unit at the end position outputs the result to the accumulation buffer for operation. After the PE calculation unit at the end position completes the operation, it marks the completion of a matrix multiplication operation. The result in the accumulation register is output to the output cache, and then the accumulator is cleared. It has high parallelism, accurate calculation results, and fast calculation speed, so it can perfectly replace the operation of using a cyclic array and DSP for multiplication and addition calculation to implement matrix multiplication in traditional matrix operations.

[0052] See Figure 3 , Figure 4, systolic array: composed of multiple PE computing units, which horizontally and vertically transfer the multiplication and addition results output by the PE computing units, and continuously repeat the above operations until the multiplication and addition calculations of all elements in the matrix are completed; a systolic matrix is selected to implement matrix multiplication operations, so that the data path of long data packets and multi-fan-in and multi-fan-out is transformed into data communication between each PE. Therefore, it has the characteristics of high parallelism, simple structure, high operating frequency, and high calculation accuracy.

[0053] See Figure 5 , matrix multiplication computing unit: The matrix Q and matrix K are calculated by using the systolic array through the data shaping module. After passing the calculation result through the data shaping unit, it is cached in the cache 3, and the cached result is output through the data output module to obtain the product matrix of the Q matrix and the K matrix;

[0054] See Figure 6 , self-attention mechanism module: The input matrix Q and the input matrix K pass through the preloading unit and the matrix multiplication computing unit to obtain the product matrix, and then the product matrix is normalized by the SoftMax computing unit to obtain the data SoftMax<Q, K>. The obtained data SoftMax<Q, K> and the input matrix V are subjected to matrix multiplication through the matrix multiplication computing unit to obtain the data Attention(Q, K, V), and the data Attention(Q, K, V) is output through the data output module.

[0055] The described SoftMax computing unit is used to normalize the product matrix of matrix Q and matrix K.

[0056] See Figure 7 , an implementation method based on the self-attention mechanism of the systolic array, including the following steps:

[0057] Step 1: The input data Q matrix and the input K matrix pass through the matrix multiplication computing unit to obtain the product matrix of the Q matrix and the K matrix;

[0058] Step 2: According to the product matrix obtained in Step 1, perform normalization calculation through the SoftMax computing unit to obtain the data SoftMax<Q, K>;

[0059] Step 3: Repeat Step 1 with the data SoftMax<Q, K> obtained in Step 2 and the input matrix V to obtain the data Attention(Q, K, V).

[0060] The specific process of the described Step 1 is:

[0061] Step 1.1: Cache the input row matrix Q and the input column matrix K through the data cache array;

[0062] Step 1.2: The input row matrix Q and input column matrix K cached in Step 1.1 are misaligned row - by - row or column - by - column through the data shaping and loading module to obtain the vector data data3 with row - misaligned arrangement and the vector data data4 with column - misaligned arrangement. Then, the data data3 and data4 are transmitted to the PE computing unit;

[0063] Step 1.3: The data data3 obtained in the pre - processing unit is cached in the cache cache1 through the row input port dim1 of the row input end, and the data data4 obtained in the pre - processing unit is cached in the cache cache2 through the column input port dim1 of the column input end. The data data3 cached in the cache cache1 and the data data4 cached in the cache cache2 are multiplied by a multiplier to obtain the data data5. At the same time, the multiplier transmits the data data5 through the cache cache3 to the adder for an addition operation with the data data5 in the previous PE computing unit to obtain the data data6 of this PE computing unit. The obtained multiplication - addition result data data6 is passed to the adder of the next PE unit; Through multiple PE computing units, the multiplication - addition results output by the PE computing units are transmitted horizontally and vertically, and the above operations are continuously repeated until all the multiplication - addition calculations of the elements in the matrix Q and matrix K are completed; After the matrix Q and matrix K pass through the data shaping module, the product matrix of the Q matrix and the K matrix is obtained.

[0064] The calculation formula of the self - attention mechanism is as follows:

[0065]

[0066] Where: Q represents the input Q matrix, K represents the input K matrix, V represents the input V matrix, d Q,K represents the dimension of the Q matrix and the K matrix, SoftMax represents the normalization process, and T represents the transpose;

[0067] The self - attention computing unit based on the systolic array described in the present invention can be applied to the hardware acceleration design of the Transformer neural network, or other places where self - attention hardware design is required.

Claims

1. An implementation system based on a systolic array self-attention mechanism, characterized in that, it includes: Preprocessing unit: Processes the input row-direction matrix data data1 and the input column-direction matrix data data2 to obtain the misaligned row vector data data3 and column vector data data4, and transmits the row vector data data3 and column vector data data4 to the PE calculation unit; PE calculation unit: Caches the data data3 obtained in the preprocessing unit in the cache cache1, caches the data data4 obtained in the preprocessing unit in the cache cache2, multiplies the data data3 cached in the cache cache1 and the data data4 cached in the cache cache2 through a multiplier to obtain the data data5, and then caches the data data5 in the cache cache3. At the same time, the cache cache3 sends a feedback signal to the cache cache1 and the cache cache2 to determine whether the transmission of the data data3 and the data data4 is completed. If completed, the data data5 is transmitted to the adder to be added to the data in the adder of the previous PE calculation unit to obtain the data data6. If not completed, wait until the transmission of the data data3 and the data data4 is completed, then perform multiplication calculation to obtain the data data5, and transmit the calculated result data data5 to the adder to be added to the data in the adder of the previous PE calculation unit to obtain the data data6. The data data6 is transmitted to the adder in the next PE calculation unit, and when the next clock cycle arrives, the data data3 and the data data4 are transmitted to the row input end dim1 and the column input end dim2 of the next PE calculation unit in the row direction and the column direction, and repeat the above operations; Systolic array: Consists of multiple PE calculation units, performs horizontal and vertical transmission on the multiplication and addition results output by the PE calculation units, and continuously repeats the above operations until the multiplication and addition calculations of all elements in the matrix are completed; Matrix multiplication calculation unit: Calculates the matrices Q and K through the data shaping module using the systolic array. After passing the calculation result through the data shaping unit, it is cached in the cache cache3, and the cached result is output through the data output module to obtain the product matrix of the Q matrix and the K matrix; Self-attention mechanism module: Obtains the product matrix by passing the input matrix Q and the input matrix K through the preloading unit and the matrix multiplication calculation unit, then normalizes the product matrix through the SoftMax calculation unit to obtain the data SoftMax<Q, K>, multiplies the obtained data SoftMax<Q, K> with the input matrix V through the matrix multiplication calculation unit to obtain the data Attention(Q, K, V), and outputs the data Attention(Q, K, V) through the data output module.

2. The implementation system based on a systolic array self-attention mechanism according to claim 1, characterized in that, The preprocessing unit described above: includes a data cache array and a data shaping and loading module; The data cache array: is used to cache the input row-direction matrix data data1 and the input column-direction matrix data data2; The data shaping and loading module: is used to divide the input row matrix and the input column matrix saved in the data cache array into a horizontal row vector form and a vertical column vector form in a row-splitting or column-splitting manner, and then send the split horizontal row vectors and vertical column vectors into the shift register. Using the method of delaying beats, each row or each column of the split row vectors and column vectors is misaligned to obtain a new row vector or column vector after misalignment, and the new row vector or column vector after misalignment is transmitted to the PE calculation unit.

3. An implementation system based on the systolic array self-attention mechanism according to claim 1, characterized in that, The adder in the PE calculation unit will simultaneously send a feedback signal to the multiplier to determine whether the multiplication calculation is completed. If it is completed, the data data5 is added to the data data6 in the adder of the previous PE calculation unit to obtain the data data6 of this PE calculation unit. If it is not completed, wait for the multiplication calculation to be completed to obtain the data data5, and transmit the calculation result data data5 to the adder to be added to the data data6 in the adder of the previous PE calculation unit to obtain the data data6 of this PE calculation unit.

4. An implementation system based on the systolic array self-attention mechanism according to claim 1, characterized in that, The SoftMax calculation unit: is used to perform normalization calculation on the data output by the matrix multiplication calculation unit.

5. An implementation method based on the systolic array self-attention mechanism for any one of the systems according to claims 1-4, characterized in that, includes the following steps: Step 1: The Q matrix and the input K matrix of the input data pass through the matrix multiplication calculation unit to obtain the product matrix of the Q matrix and the K matrix Step 2: According to the product matrix obtained in Step 1, perform normalization calculation through the SoftMax calculation unit to obtain the data SoftMax<Q, K>; Step 3: According to the data SoftMax<Q, K> obtained in Step 2 and the input matrix V, repeat Step 1 to obtain the data Attention(Q, K, V).

6. An implementation method based on the systolic array self-attention mechanism according to claim 5, characterized in that, The specific process of Step 1 is: Step 1.1: Cache the input row matrix Q and the input column matrix K through the data cache array; Step 1.2: Misalign the input row matrix Q and the input column matrix K cached in Step 1.1 in a row or column manner to obtain the vector data data3 with row misalignment and the vector data data4 with column misalignment, and then transmit the data data3 and the data data4 to the PE calculation unit; Step 1.3: Cache the data data3 obtained in the preprocessing unit in cache cache1 through the row input port dim1, and cache the data data4 obtained in the preprocessing unit in cache cache2 through the column input port dim1. Multiply the data data3 cached in cache cache1 and the data data4 cached in cache cache2 through a multiplier to obtain the data data5. At the same time, the multiplier transmits the data data5 through cache cache3 to the adder to perform an addition operation with the data data5 in the previous PE computing unit to obtain the data data6 of this PE computing unit. Pass the obtained multiplication and addition result data data6 to the adder of the next PE unit; perform horizontal and vertical transmission of the multiplication and addition results output by the PE computing unit through multiple PE computing units, and continuously repeat the above operations until the multiplication and addition calculations of all elements in matrix Q and matrix K are completed; pass matrix Q and matrix K through the data shaping module to obtain the product matrix of matrix Q and matrix K.

Citation Information

Patent Citations

  • Systolic array, systolic array system, operation method and device of systolic array system, and storage medium

    CN115423084A

  • Parallel matrix multiplier based on single field programmable gate array (FPGA) and implementation method for parallel matrix multiplier

    CN102662623A

  • Hardware accelerator capable of configuring sparse attention mechanism

    CN113901747A