Matrix multiplication apparatus and method based on systolic array, and electronic device
By designing a matrix multiplication device based on pulsating arrays, using basic computing units and control signal channels, sparse and dense matrix multiplication operations are realized, which solves the problem of high computational complexity in the prior art and improves the calculation efficiency.
Patent Information
- Application Number
- PCT/CN2024/130454
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-11-07
- Publication Date
- 2025-07-03
AI Technical Summary
The existing computing architecture has high computational complexity when dealing with sparse matrix multiplication operations and is not good at dealing with sparse matrix multiplication operations with irregular patterns, resulting in low computational efficiency of the Transformer model.
A matrix multiplication device based on a pulsating array is designed, and a pulsating array arranged by a plurality of basic computing units is combined with vector displacement units and multiple control signal channels to realize sparse and dense matrix multiplication operations, reducing the computational complexity.
It reduces the computational complexity of matrix multiplication operations, improves the computational efficiency, and can take into account both sparse and dense matrix multiplication operations, and is suitable for scenarios such as Transformer models and sparse transformer models.
Smart Images

Figure CN2024130454_03072025_PF_FP_ABST
Abstract
Description
Matrix multiplication device, method and electronic device based on systolic array Technical Field
[0001] The present application belongs to the field of integrated circuit and neural network technology, and in particular relates to a matrix multiplication device, method and electronic device based on a systolic array.
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 29, 2023, with application number 202311869723.0 and invention name “Matrix multiplication device, method and electronic device based on systolic array”, the entire contents of which are incorporated by reference into this application. Background Art
[0003] The Transformer model is an attention-based neural network model used to process sequential data. Compared to traditional recurrent neural network models, the Transformer model has better parallel performance and shorter training time. Therefore, it has been widely used in fields such as natural language processing (NLP) and computer vision.
[0004] The Transformer model uses self-attention to capture contextual information from the entire sequence and obtain predictions. The self-attention module uses full attention and requires three matrix inputs: Q (query matrix), K (key matrix), and V (value matrix). The calculation involves matrix multiplication. Because of this matrix multiplication, the model's computational complexity is proportional to the square of the input sequence length. The longer the input sequence, the more complex the computation. Furthermore, as the Transformer model's application scenarios expand, the length of the input sequence increases, placing a heavy burden on the Transformer model's computational and memory requirements.
[0005] Currently, sparse transformer models can be used to reduce the computational complexity of Transformer models. Sparse transformer models can use global attention, local attention, and random attention to replace the original full attention, involving sparse matrix multiplication. However, in the current computing architecture, the computing units are better at processing dense matrix operations and are not good at processing sparse matrix multiplication operations with irregular patterns. The computational complexity of sparse transformer models is still high, and the computational efficiency is poor. Technical issues
[0006] The embodiments of the present application provide a matrix multiplication device, method, and electronic device based on a systolic array, which reduces the computational complexity of matrix multiplication operations and improves computational efficiency.
[0007] In a first aspect, an embodiment of the present application provides a matrix multiplication device based on a systolic array, comprising:
[0008] A systolic array formed by an arrangement of multiple basic computing units, wherein adjacent two basic computing units in each column of the systolic array are connected by a vector displacement unit, wherein the basic computing unit is used to perform vector dot multiplication operations, and the vector displacement unit is used to perform vector displacement operations; in the systolic array, a first transmission channel is established between adjacent two basic computing units in the row direction, a second transmission channel and / or a third transmission channel is established between adjacent two basic computing units in the column direction, and a fourth transmission channel is established between adjacent two basic computing units in the diagonal direction, wherein the first transmission channel is used to input row vectors of a first matrix, and at least one of the second, third, and fourth transmission channels is used to input column vectors of a second matrix;
[0009] a plurality of first control signal channels connected to the second transmission channels and / or the third transmission channels of the basic computing units in a plurality of column directions of the systolic array; the first control signal channels being used to transmit first control signals, the first control signals being used to control the working state of the vector displacement units;
[0010] Multiple second control signal channels are connected to the fourth transmission channels of the basic computing units in multiple diagonal directions of the systolic array; the second control signal channels are used to transmit second control signals, and the second control signals are used to control whether the column vectors of the second matrix input to the basic computing units are input in the column direction or the diagonal direction.
[0011] In a second aspect, embodiments of the present application provide a matrix multiplication method based on a systolic array, which is applied to a matrix multiplication device based on a systolic array provided in embodiments of the present application. The method includes:
[0012] Acquire multiple row vectors of the first matrix, multiple column vectors of the second matrix, the first control signal, and the second control signal; the first control signal is used to control the working state of the vector displacement unit, and the second control signal is used to control whether the column vectors of the second matrix input to the basic computing unit are input in the column direction or the diagonal direction;
[0013] Inputting a plurality of row vectors of the first matrix into a plurality of rows of basic computing units in a one-to-one correspondence along the row direction of the systolic array; and inputting a plurality of column vectors of the second matrix into a plurality of columns of basic computing units in a one-to-one correspondence along the column direction of the systolic array, or inputting a plurality of diagonal directions of basic computing units in a one-to-one correspondence along the diagonal direction of the systolic array, under the control of the first control signal and the second control signal;
[0014] The input vector is calculated by the basic calculation unit based on the systolic array to obtain a matrix multiplication result of the first matrix and the second matrix.
[0015] In a third aspect, an embodiment of the present application provides an electronic device, comprising the matrix multiplication device based on the systolic array provided in the first aspect above.
[0016] In a fourth aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method provided in the second aspect above when executing the computer program.
[0017] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the second aspect above is implemented.
[0018] This application provides
[0019] Matrix multiplication of systolic arrays
[0020] The device, through the design of a systolic array architecture with reconfigurable basic computing units and hardware circuit design, can achieve both sparse and dense matrix multiplication operations, supporting a variety of matrix multiplication scenarios and reducing the computational complexity of matrix multiplication. This reduces computational complexity and improves computational efficiency in scenarios requiring sparse and / or dense matrix multiplication, such as Transformer models and sparse Transformer models. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] FIG1 is a schematic diagram of a structure of a self-attention module in a Transformer model provided in an embodiment of the present application;
[0022] FIG2 is a set of schematic diagrams of complete attention, global attention, local attention, and random attention provided by an embodiment of the present application;
[0023] FIG3 is a schematic diagram of a calculation principle of a basic calculation unit provided in an embodiment of the present application;
[0024] FIG4 is a schematic diagram of a calculation principle of a vector displacement unit provided in an embodiment of the present application;
[0025] FIG5 is a schematic diagram of a systolic array provided in an embodiment of the present application;
[0026] FIG6 is a schematic structural diagram of a matrix multiplication device based on a systolic array according to an embodiment of the present application;
[0027] FIG7A is a schematic diagram of a first matrix and a second matrix in a general mode provided by an embodiment of the present application;
[0028] FIG7B is a schematic diagram showing the principle of data flow in a matrix multiplication device based on a systolic array corresponding to FIG7A provided in an embodiment of the present application;
[0029] FIG8A is a schematic diagram of a first matrix and a second matrix in a first sparse mode provided by an embodiment of the present application;
[0030] FIG8B is a schematic diagram showing the principle of data flow in a matrix multiplication device based on a systolic array corresponding to FIG8A provided in an embodiment of the present application;
[0031] FIG9A is a schematic diagram of a first matrix and a second matrix in a second sparse mode provided by an embodiment of the present application;
[0032] FIG9B is a schematic diagram showing the principle of data flow in a matrix multiplication device based on a systolic array corresponding to FIG8A provided in an embodiment of the present application;
[0033] FIG10 is a flowchart of a matrix multiplication method based on a systolic array according to an embodiment of the present application;
[0034] FIG11 is a schematic diagram of a matrix storage provided in an embodiment of the present application;
[0035] FIG12 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. Modes for Carrying Out the Invention
[0036] The matrix multiplication device and method based on the systolic array provided in this application can realize both sparse matrix multiplication operations and dense matrix multiplication operations, and can be applied to scenarios such as Transformer models and sparse transformer models that require sparse matrix multiplication operations, thereby reducing the computational complexity of matrix multiplication and improving computational efficiency.
[0037] First, the concepts involved in this application are explained.
[0038] 1. Structure of the Self-Attention Module in the Transformer Model
[0039] Figure 1 is a schematic diagram of the structure of the self-attention module in the Transformer model provided in an embodiment of the present application. As shown in Figure 1, the self-attention module requires three input matrices, including: Q (query matrix), K (key matrix), and V (value matrix). The self-attention module includes two matrix multiplication (MatMul) operations and one matrix normalization (Softmax) operation. Specifically, Q and KT are MatMuled to generate a similarity matrix S. S is then Softmaxed and then MatMuled with V to obtain the result.
[0040] The matrix multiplication device and method based on the systolic array provided in this application can realize matrix multiplication in the self-attention module, and the attention mechanisms used include but are not limited to complete attention, global attention, local attention and random attention.
[0041] 2. Full attention, global attention, local attention, and random attention.
[0042] Figure 2 is a set of schematic diagrams of complete attention, global attention, local attention, and random attention provided in an embodiment of the present application, showing the matrix calculation results under different attention mechanisms, involving dense matrices and sparse matrices. The computational complexity of matrix multiplication operations in different attention mechanisms is different, and the computational complexity of complete attention is more complex. Specifically, as shown in Figure 2, the complete attention mechanism obtains dense matrix calculation results, and global attention, local attention, and random attention obtain sparse matrix calculation results, wherein the patterns of the sparse matrices are different. Exemplarily, the dark part in Figure 2 is valid data. For example, the pattern of the sparse matrix obtained by local attention is in the shape of a diagonal line, that is, the row vectors in the sparse matrix are arranged in the form of a diagonal line; the pattern of the sparse matrix obtained by random attention is a random shape.
[0043] 3. Dense matrix, sparse matrix
[0044] Dense matrices and sparse matrices are two contrasting matrix types. Generally, a dense matrix is one in which most elements are non-zero or valid, while a sparse matrix is one in which most elements are zero or invalid. In both dense and sparse matrices, non-zero or valid elements occupy storage space, so dense matrices require more storage space. Furthermore, in matrix operations, since every element must be calculated, the greater the number of non-zero or valid elements, the greater the computational complexity and the longer the calculation time.
[0045] 4. Basic computing unit (PE)
[0046] The basic computing unit is used for dot multiplication of vectors. This application does not limit the number of vectors. For ease of explanation, this application uses the basic computing unit used for dot multiplication of two vectors as an example. This application does not limit the name of the basic computing unit. For example, the basic computing unit is also called a dot multiplication unit or expressed as a PE (Processing Element).
[0047] For example, Figure 3 is a schematic diagram of a calculation principle of a basic computing unit provided in an embodiment of the present application. As shown in Figure 3, the two input vectors input to the basic computing unit (PE) are a 1×m-dimensional vector and an n×1-dimensional vector, respectively. m and n can be equal or unequal. If m and n are unequal, the vector with the lower dimension needs to be padded with zeros so that the two input vectors have the same dimension. Optionally, the zero-padding operation can be performed in either the high-order or low-order bits.
[0048] 5. Vector displacement unit (shift)
[0049] The vector displacement unit is used for vector displacement operations. This application does not limit the name of the vector displacement unit. For example, the vector displacement unit is also called a shift unit or expressed as shift.
[0050] For example, FIG4 is a schematic diagram of a calculation principle of a vector shift unit provided in an embodiment of the present application. As shown in FIG4 , the input vector of the input vector shift unit (shift) is an (n+1)×1-dimensional vector, which is represented as 0, 1, 2, ..., n from high to low. After passing through the vector shift unit (shift), the input (n+1)×1-dimensional vector is shifted by one data unit as a whole, and the vacant positions are padded with zeros. The output vector is (n+1)×1-dimensional, which is represented as 1, 2, ..., n, 0 (zero-padded position) from high to low.
[0051] It should be noted that the embodiment of the present application does not limit the direction of vector displacement. The vector as a whole can be shifted one data unit toward the lower bit, or the vector as a whole can be shifted one data unit toward the higher bit.
[0052] 6. Pulsation, Pulsation Array
[0053] Systolic refers to the way valid data is transmitted in the array like a wave, producing a very regular data flow. The flow of data is regular and periodic.
[0054] A systolic array is a network composed of tightly coupled basic units governed by data flow. These units, also called nodes, implement specific functions, such as vector dot products. Each node interacts with one or more surrounding nodes, and computational results can be stored within the unit or passed to other nodes.
[0055] For example, Figure 5 is a schematic diagram of a systolic array provided in an embodiment of the present application. As shown in Figure 5, the systolic array includes four basic units, labeled basic unit 1 through basic unit 4. Vector X is sequentially systolic-input to basic units 1 through 4. Basic unit 1 performs a calculation based on the input vector X to obtain result 1. Similarly, basic unit 2 performs a calculation based on the input vector X to obtain result 2, basic unit 3 performs a calculation based on the input vector X to obtain result 3, and basic unit 4 performs a calculation based on the input vector X to obtain result 4.
[0056] The matrix multiplication device and method based on the systolic array provided in this application are described in detail below with reference to the accompanying drawings.
[0057] FIG6 is a schematic diagram of a structure of a matrix multiplication device based on a systolic array provided in an embodiment of the present application. As shown in FIG6 , the matrix multiplication device based on a systolic array provided in this embodiment may include:
[0058] A systolic array is composed of a plurality of basic computing units 61. Adjacent basic computing units 61 in each column of the systolic array are connected by a vector displacement unit 62. The basic computing units 61 are used to perform vector dot multiplication operations, while the vector displacement unit 62 is used for vector displacement operations. In the systolic array, a first transmission channel is established between adjacent basic computing units 61 in the row direction, a second transmission channel and / or a third transmission channel are established between adjacent basic computing units 61 in the column direction, and a fourth transmission channel is established between adjacent basic computing units 61 in the diagonal direction. The first transmission channel is used to input row vectors of the first matrix, and at least one of the second, third, and fourth transmission channels is used to input column vectors of the second matrix.
[0059] Multiple first control signal channels communicate with the second transmission channels and / or third transmission channels of the basic computing units 61 in multiple column directions of the systolic array. The first control signal channels are used to transmit first control signals, which are used to control the operating state of the vector shift unit 62.
[0060] Multiple second control signal channels are connected to the fourth transmission channels of the basic computing units 61 in multiple diagonal directions of the systolic array. The second control signal channels are used to transmit second control signals, which are used to control whether the column vectors of the second matrix input to the basic computing units 61 are input in the column direction or the diagonal direction.
[0061] For example, for ease of explanation, in FIG6 , the basic computing unit 61 is represented as PE, the vector shift unit 62 is represented as shift, and the systolic array includes PEs of a 4×4 architecture. However, FIG6 does not limit this.
[0062] This embodiment provides a systolic array-based matrix multiplication device, comprising a systolic array formed by multiple rows and columns of basic computing units. This embodiment does not limit the number of rows and columns in the systolic array; for example, the number of rows is M and the number of columns is N. M and N can be equal or unequal. For example, FIG6 illustrates M=N=4. The systolic array can implement systolic input in the row, column, and diagonal directions. For each row, each basic computing unit in the same row receives the same row vector of a first matrix in the row direction as an input vector to the basic computing unit. For example, in FIG6, PE11, PE12, PE13, and PE14 are located in the same row and receive the same row vector. The basic computing units can also receive column vectors of a second matrix in the column or diagonal direction as another input vector. For each column of the systolic array, any two adjacent basic computing units are connected via a vector shift unit. The vector shift unit is activated or deactivated (or bypassed) under the control of a first control signal. Depending on whether the vector shift unit is active (or bypassed), different basic computing units in the same column may receive the same or different column-wise input vectors. For example, in Figure 6, PE12, PE22, PE32, and PE42 are on the same column-wise transmission channel. If the vector shift unit in that column is inactive, PE12, PE22, PE32, and PE42 can receive the same column-wise input vector. If the vector shift unit in that column is active, PE12, PE22, PE32, and PE42 can receive different column-wise input vectors, with the column vectors differing by one data unit per transmission level. For each diagonal line of the systolic array, the diagonal transmission channel is controlled by a second control signal. For example, in Figure 6, PE13, PE22, and PE31 are on the same diagonal transmission channel. On this transmission channel, PE13, PE22, and PE31 can receive the same column-wise input vector.
[0063] The matrix multiplication device based on a systolic array provided in this embodiment operates in principle in that a basic computing unit has a first transmission channel in the row direction, through which the row vectors of the first matrix can be transmitted sequentially along the first transmission channel in the row direction. The basic computing unit also has a second transmission channel and / or a third transmission channel in the column direction, as well as a fourth transmission channel in the diagonal direction. Under the control of a first control signal and a second control signal, the column vectors of the second matrix can be transmitted sequentially along the column or diagonal transmission channels. Thus, for each basic computing unit, the row vectors of the first matrix and the column vectors of the second matrix can be obtained, performing a vector dot multiplication operation. Consequently, the entire matrix multiplication device based on a systolic array can perform a matrix multiplication operation between the first and second matrices. Furthermore, under the control of the first control signal, depending on whether the vector shift unit is active, the column vectors of the second matrix obtained by different basic computing units in the same column may or may not be the same. Under the control of the second control signal, the basic computing unit can obtain the column vectors of the second matrix in the column direction or in the diagonal direction. Therefore, different forms of matrix multiplication operations can be performed, including dense matrix multiplication and sparse matrix multiplication, resulting in either a dense matrix or a sparse matrix.
[0064] It can be seen that the matrix multiplication device based on the systolic array provided in this embodiment can realize sparse matrix multiplication operations through the design of a systolic array architecture of reconfigurable basic computing units and hardware circuit design, thereby reducing the computational complexity of matrix multiplication and improving computational efficiency. Moreover, it can also realize both sparse matrix multiplication operations and dense matrix multiplication operations, supporting a variety of matrix multiplication operation scenarios. In scenarios such as Transformer models and sparse transformer models that require sparse matrix multiplication operations and / or dense matrix multiplication operations, the computational complexity is reduced and the computational efficiency is improved.
[0065] It should be noted that when calculating the multiplication of a first matrix and a second matrix, the number of row vectors of the first matrix may be the same as or different from the number of rows of the systolic array in the matrix multiplication device based on the systolic array, and the number of column vectors of the second matrix may be the same as or different from the number of columns of the systolic array in the matrix multiplication device based on the systolic array. If the dimensions of the first matrix, the dimensions of the second matrix, and the dimensions of the basic computing units in the matrix multiplication device based on the systolic array are different, and a single input calculation by the matrix multiplication device based on the systolic array cannot complete the multiplication operation of the first matrix and the second matrix, the first matrix and the second matrix can be split into matrices, and the matrix multiplication device based on the systolic array can perform multiple iterative calculations using the input, and then concatenate the results of each calculation to obtain the multiplication result of the first matrix and the second matrix. This embodiment does not limit the matrix splitting method.
[0066] Optionally, the first control signal and the second control signal are related to a matrix multiplication mode.
[0067] Among them, the matrix multiplication mode includes:
[0068] A general mode is used to perform a matrix multiplication operation on a first matrix and a second matrix to obtain a first result matrix.
[0069] The first sparse mode is used to perform a matrix multiplication operation on the first matrix and the second matrix to obtain a second result matrix. In the general mode and the first sparse mode, the second matrix is input into the basic computing unit in different ways, and the way the second matrix is input into the basic computing unit is controlled by a second control signal.
[0070] A second sparse mode is configured to perform a matrix multiplication operation on the first matrix and the second matrix to obtain a third result matrix, wherein the row vectors in the first matrix are arranged in a diagonal form, and the second matrix is a dense matrix. The operating state of the vector shift unit is different in the general mode and the second sparse mode, and the operating state of the vector shift unit is controlled by the first control signal.
[0071] Optionally, the first result matrix is a dense matrix, the second result matrix is a sparse matrix, the row vectors in the second result matrix are arranged in a diagonal form, and the third result matrix is a dense matrix.
[0072] The following takes the structure shown in FIG6 as an example, where the first matrix is A and the second matrix is B, and describes in detail the first control signal and the second control signal in different matrix multiplication modes.
[0073] Optionally, in one implementation, the matrix multiplication mode is a general mode. In the general mode, the first control signal is used to control the vector displacement unit to be inoperative, and the second control signal is used to control the second matrix input to the basic computing unit via the second transmission channel to be input in the column direction, with each column of the basic computing unit corresponding to a column in the second matrix.
[0074] The general mode can be used for dense matrix multiplication or dense convolution operations, and the result of the calculation is a dense matrix.
[0075] For example, FIG7A is a schematic diagram of a first matrix and a second matrix in a general mode provided in an embodiment of the present application, and FIG7B is a schematic diagram of the principle of data flow in a matrix multiplication device based on a systolic array corresponding to FIG7A provided in an embodiment of the present application.
[0076] As shown in Figure 7A, the first matrix A is a 4×16-dimensional matrix, and the second matrix B is a 16×4-dimensional matrix. The multiplication of the first matrix A and the second matrix B can obtain a 4×4-dimensional dense matrix C. The first matrix A includes four row vectors, labeled a0 to a3. The second matrix B includes four column vectors, labeled b0 to b3.
[0077] As shown in Figure 7B , in general mode, the first control signal is used to control the vector shift unit Shift to be inoperative and bypassed. The second control signal is used to control the column vectors of the second matrix input to the basic computing unit PE to be input in the column direction. The input vectors of each basic computing unit PE are transmitted sequentially in the row and column directions. Through the first transmission channel in the row direction, each basic computing unit PE in the same row obtains the same row vector; through the second transmission channel in the column direction, each basic computing unit PE in the same column obtains the same column vector.
[0078] Exemplarily, in FIG7B , the data flow is represented by a solid line. For example, for the basic computing unit PE11, the input vector includes the row vector a0 of the first matrix input in the row direction and the column vector b0 of the second matrix input in the column direction. For the basic computing unit PE12, the input vector includes the row vector a0 of the first matrix input in the row direction and the column vector b1 of the second matrix input in the column direction. For the basic computing unit PE21, since the vector shift unit shift in the column direction is bypassed, the input vector includes the row vector a1 of the first matrix input in the row direction and the column vector b0 of the second matrix input in the column direction. For the basic computing unit PE22, since the vector shift unit shift in the column direction is bypassed, the input vector includes the row vector a1 of the first matrix input in the row direction and the column vector b1 of the second matrix input in the column direction.
[0079] As shown in Figures 7A and 7B , using a systolic array-based matrix multiplication device, a0 is dot-multiplied with b0, b1, b2, and b3, respectively. a1 is dot-multiplied with b0, b1, b2, and b3, respectively. The same applies to a2 and a3. The final result of the multiplication of the first matrix A and the second matrix B is a 4×4 dense matrix C.
[0080] Note that in this example, the window size in the row and column directions of the matrix calculation results is 4. In actual applications, if the window size is larger than the number of PEs per row or column of the basic computing unit in the systolic array (in this example, the number of PEs per row and column is 4), it is necessary to split the first and second matrices, perform multiple iterative calculations, and concatenate the results in the row and / or column directions to obtain the final result.
[0081] Optionally, in another implementation, the matrix multiplication mode is a first sparse mode. In the first sparse mode, the first control signal is used to control the vector displacement unit to be inoperative, and the second control signal is used to control the second matrix input to the basic computing unit through the fourth transmission channel to be input in a diagonal direction, with each basic computing unit in the diagonal direction corresponding to a column in the second matrix.
[0082] For example, FIG8A is a schematic diagram of a first matrix and a second matrix in a first sparse mode provided in an embodiment of the present application, and FIG8B is a schematic diagram of the principle of data flow in a matrix multiplication device based on a systolic array corresponding to FIG8A provided in an embodiment of the present application.
[0083] As shown in Figure 8A, the first matrix A is a 4×16-dimensional matrix, and the second matrix B is a 16×7-dimensional matrix. The multiplication of the first matrix A and the second matrix B can generate a 4×7-dimensional sparse matrix C with a diagonal pattern. The row vectors in matrix C are arranged in a diagonal pattern. The first matrix A includes four row vectors, labeled a0 to a3. The second matrix B includes seven column vectors, labeled b0 to b6.
[0084] As shown in Figure 8B, in the first sparse mode, the first control signal is used to control the vector shift unit Shift to be inoperative and bypassed. The second control signal is used to control the column vectors of the second matrix input to the basic computing unit to be input in the diagonal direction. The input vectors of each basic computing unit PE are transmitted sequentially in the row direction and the diagonal direction. Through the first transmission channel in the row direction, each basic computing unit PE in the same row obtains the same row vector; through the fourth transmission channel in the diagonal direction, each basic computing unit PE in the same diagonal line obtains the same column vector.
[0085] For example, in FIG8B , the data flow is represented by a solid line. For example, for the basic computing unit PE11, the input vector includes the row vector a0 of the first matrix input in the row direction and the column vector b0 of the second matrix input in the diagonal direction. For the basic computing unit PE12, the input vector includes the row vector a0 of the first matrix input in the row direction and the column vector b1 of the second matrix input in the diagonal direction. It can be understood that for the basic computing units in the first row, the column direction and the diagonal direction can be considered to be the same direction. For the basic computing unit PE21, the input vector includes the row vector a1 of the first matrix input in the row direction and the column vector b1 of the second matrix input in the diagonal direction. For the basic computing unit PE22, the input vector includes the row vector a1 of the first matrix input in the row direction and the column vector b2 of the second matrix input in the diagonal direction. For the basic computing unit PE24, the input vector includes the row vector a1 of the first matrix input in the row direction and the column vector b4 of the second matrix input in the diagonal direction.
[0086] As shown in Figures 8A and 8B , through the matrix multiplication device based on a systolic array, a0 is dot-multiplied with b0, b1, b2, and b3, respectively. A1 is dot-multiplied with b1, b2, b3, and b4, respectively. The same applies to a2 and a3. The final multiplication result of the first matrix A and the second matrix B is a 4×7 sparse matrix C with a diagonal pattern. The row vectors in matrix C are arranged in a diagonal pattern.
[0087] Note that in this example, the diagonal window size in the matrix calculation result is 4. In practice, if the window size is larger than the number of PEs per row in the systolic array (4 in this example), multiple iterations of the calculation are required, and the results are concatenated row-wise to obtain the final result.
[0088] Optionally, in another implementation, the matrix multiplication mode is a second sparse mode. In the second sparse mode, the first control signal is used to control the operation of the vector shift unit, and the second control signal is used to control the input of column vectors of the second matrix into the basic computing unit via the third transmission channel, where each column of the basic computing unit corresponds to a column in the second matrix.
[0089] For example, FIG9A is a schematic diagram of a first matrix and a second matrix in a second sparse mode provided in an embodiment of the present application, and FIG9B is a schematic diagram of the principle of data flow in a matrix multiplication device based on a systolic array corresponding to FIG8A provided in an embodiment of the present application.
[0090] As shown in Figure 9A, the first matrix A is a 4×19-dimensional matrix, and the second matrix B is a 19×4-dimensional matrix. The multiplication of the first matrix A and the second matrix B produces a 4×4-dimensional dense matrix C. The first matrix A includes four row vectors, labeled a0 through a3. The first matrix A is a sparse matrix, in which the row vectors are arranged diagonally. The second matrix B is a dense matrix, consisting of four column vectors, labeled b0 through b3.
[0091] As shown in Figure 9B, in the second sparse mode, the first control signal is used to control the vector shift unit to operate, and the second control signal is used to control the column vectors of the second matrix input to the basic computing unit to be input in the column direction. The input vectors of each basic computing unit PE are transmitted sequentially in the row direction and the column direction respectively. Through the first transmission channel in the row direction, each basic computing unit PE in the same row obtains the same row vector. Through the third transmission channel in the column direction, due to the operation of the vector shift unit, different basic computing units PE in the same column obtain different column vectors. Moreover, with each level of transmission down in the column direction, the column vector is shifted by one data unit through the vector shift unit.
[0092] For example, in FIG9B , the data flow is represented by a solid line. For example, for basic computing unit PE11, the input vector includes the row vector a0 of the first matrix input in the row direction and the column vector b0 of the second matrix input in the column direction (i.e., b0[0:15]). For basic computing unit PE12, the input vector includes the row vector a0 of the first matrix input in the row direction and the column vector b1 of the second matrix input in the column direction (i.e., b1[0:15]). For basic computing unit PE21, due to the operation of the vector shift unit in the column direction, the column vector b0 of the second matrix (i.e., b0[0:15]) is shifted by one data unit to form the column vector b0[1:16]. Therefore, the input vector includes the row vector a1 and the column vector b0[1:16] of the first matrix input in the row direction. For the basic computing unit PE22, due to the operation of the vector shift unit shift in the column direction, the column vector b1 (i.e., b1[0:15]) of the second matrix is shifted by one data unit to form a column vector b1[1:16]. Therefore, the input vector includes the row vector a1 and column vector b1[1:16] of the first matrix input in the row direction.
[0093] As shown in Figures 9A and 9B , using a systolic array-based matrix multiplication device, a0 is dot-multiplied with b0[0:15], b1[0:15], b2[0:15], and b3[0:15]. A1 is dot-multiplied with b0[1:16], b1[1:16], b2[1:16], and b3[1:16]. The same applies to a2 and a3. The final result of the multiplication of the first matrix A and the second matrix B is a 4×4 dense matrix C.
[0094] Note that in this example, the diagonal window size in the matrix calculation result is 16. In actual applications, if the window size is larger than the maximum vector dimension that a basic computing unit (PE) can accept (in this example, the maximum vector dimension that a PE can accept is 16), multiple iterations of the calculation are required, and the intermediate results are accumulated to obtain the final result.
[0095] It should be noted that, taking the 4×4 PE array shown in Figure 6 as an example, in general mode, the PE array can complete at most one 4×n and n×4 matrix operation per calculation cycle; in the first sparse mode, the PE array can complete at most one 4×n and n×7 sparse matrix operation per calculation cycle; and in the second sparse mode, the PE array can complete at most one 4×(n+3) and (n+3)×7 sparse matrix operation per calculation cycle. Here, n refers to the maximum dimension of the input vector that a PE can receive. In the example shown in Figure 6, n=16.
[0096] Optionally, the matrix multiplication device based on the systolic array provided in this embodiment may further include a control signal generation module, where the control signal generation module is connected to the plurality of first control signal channels and the plurality of second control signal channels.
[0097] The control signal generating module is used to generate a first control signal and a second control signal according to a matrix multiplication mode.
[0098] The matrix multiplication mode may be a parameter transmitted from other devices, and the control signal generation module may generate the first control signal and the second control signal according to the matrix multiplication mode.
[0099] Another embodiment of the present application provides a matrix multiplication method based on a systolic array, which is applied to the matrix multiplication device based on a systolic array provided in the present application. FIG10 is a flow chart of the matrix multiplication method based on a systolic array provided in an embodiment of the present application. As shown in FIG10 , the matrix multiplication method based on a systolic array may include:
[0100] S1001: Acquire multiple row vectors of a first matrix, multiple column vectors of a second matrix, a first control signal, and a second control signal. The first control signal is used to control the operating state of a vector shift unit, and the second control signal is used to control whether the column vectors of the second matrix input into a basic computing unit are input in a column-wise direction or a diagonal direction.
[0101] Among them, the first matrix, the second matrix, the basic calculation unit, the vector displacement unit, the first control signal and the second control signal are described in the embodiment shown in Figure 6 and are not repeated here.
[0102] S1002. Input multiple row vectors of the first matrix into multiple rows of basic computing units in a one-to-one correspondence along the row direction of the systolic array; and, under the control of the first control signal and the second control signal, input multiple column vectors of the second matrix into multiple columns of basic computing units in a one-to-one correspondence along the column direction of the systolic array, or input multiple diagonal directions of basic computing units in a one-to-one correspondence along the diagonal direction of the systolic array.
[0103] Specifically, depending on whether the first matrix and the second matrix are sparse matrices or dense matrices, and based on the specific requirements of obtaining the sparse matrix or dense matrix through matrix multiplication, multiple row vectors of the first matrix are respectively input into multiple rows of the systolic array along the row direction of the systolic array, with each row of basic computing units corresponding to a row of the first matrix; multiple column vectors of the second matrix are respectively input into multiple columns or multiple diagonal lines of the systolic array along the column direction or diagonal direction of the systolic array, with each column of basic computing units corresponding to a column of the second matrix, or each basic computing unit in the diagonal direction corresponding to a column of the second matrix.
[0104] S1003 . Calculate the input vector using a basic computing unit based on a systolic array to obtain a matrix multiplication result of the first matrix and the second matrix.
[0105] The matrix multiplication method based on a systolic array provided in this embodiment is applied to the matrix multiplication device based on a systolic array provided in this application. By designing a systolic array architecture of reconfigurable basic computing units, sparse matrix multiplication operations can be implemented through hardware circuit design, reducing the computational complexity of matrix multiplication and improving computational efficiency. Moreover, sparse matrix multiplication operations and dense matrix multiplication operations can be implemented at the same time, supporting a variety of matrix multiplication operation scenarios. In scenarios such as Transformer models and sparse transformer models that require sparse matrix multiplication operations and / or dense matrix multiplication operations, computational complexity is reduced and computational efficiency is improved.
[0106] Optionally, the matrix multiplication method based on the systolic array may further include:
[0107] Get the matrix multiplication mode.
[0108] Among them, the matrix multiplication mode includes:
[0109] A general mode is used to perform a matrix multiplication operation on the first matrix and the second matrix to obtain a first result matrix.
[0110] The first sparse mode is configured to perform a matrix multiplication operation on the first matrix and the second matrix to obtain a second result matrix. In the general mode and the first sparse mode, the second matrix is input into the basic computing unit in different ways, and the way the second matrix is input into the basic computing unit is controlled by a second control signal.
[0111] The second sparse mode is configured to perform a matrix multiplication operation on the first matrix and the second matrix to obtain a third result matrix, wherein the row vectors in the first matrix are arranged in a diagonal form, and the second matrix is a dense matrix. The operating state of the vector shift unit is different in the general mode and the second sparse mode, and the operating state of the vector shift unit is controlled by the first control signal.
[0112] Optionally, in general mode, the first control signal is used to control the vector displacement unit to not work, and the second control signal is used to control the second matrix input into the basic computing unit through the second transmission channel to be input in the column direction, and each column of the basic computing unit corresponds to a column in the second matrix.
[0113] Optionally, in the first sparse mode, the first control signal is used to control the vector displacement unit to not work, and the second control signal is used to control the second matrix input into the basic computing unit through the fourth transmission channel to be input in a diagonal direction, and each diagonal direction of the basic computing unit corresponds to a column in the second matrix.
[0114] Optionally, in the second sparse mode, the first control signal is used to control the operation of the vector displacement unit, and the second control signal is used to control the column vector input of the second matrix into the basic computing unit through the third transmission channel, and each column of the basic computing unit corresponds to a column in the second matrix.
[0115] Optionally, the matrix multiplication method based on the systolic array may further include:
[0116] The matrix multiplication result of the first matrix and the second matrix is stored in dense matrix form.
[0117] For example, Figure 11 is a schematic diagram of matrix storage provided by an embodiment of the present application. As shown on the left side of Figure 11, the matrix multiplication result of the first matrix and the second matrix is a sparse matrix. Based on the characteristics of sparse matrices and dense matrices, in order to save storage space and improve data transmission efficiency, the sparse matrix can be stored in the form of a dense matrix, as shown on the right side of Figure 11.
[0118] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0119] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the embodiment of the method of this application, and its specific functions and the technical effects brought about are the same.
[0120] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0121] Figure 12 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. As shown in Figure 12, the electronic device 12 includes: at least one processor 120, a memory 121, and a computer program 122 stored in the memory 121 and executable on the at least one processor 120. When the processor 120 executes the computer program 122, the steps of any of the above-mentioned method embodiments are implemented.
[0122] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0123] The so-called memory may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).
[0124] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.
Claims
1. A matrix multiplication device based on a systolic array, characterized in that Comprising: A systolic array formed by arranging a plurality of basic computing units. Adjacent two of the basic computing units in each column of the systolic array are connected by a vector displacement unit. The basic computing unit is used for performing dot product operations of vectors, and the vector displacement unit is used for displacement operations of vectors; in the systolic array, a first transmission channel is established between adjacent two of the basic computing units in the row direction, a second transmission channel and / or a third transmission channel are established between adjacent two of the basic computing units in the column direction, and a fourth transmission channel is established between adjacent two of the basic computing units in the skew diagonal direction. The first transmission channel is used for inputting the row vectors of the first matrix, and at least one of the second transmission channel, the third transmission channel, and the fourth transmission channel is used for inputting the column vectors of the second matrix; A plurality of first control signal channels, communicating with the second transmission channel and / or the third transmission channel of the basic computing units in a plurality of column directions of the systolic array; the first control signal channels are used for transmitting first control signals, and the first control signals are used for controlling the working state of the vector displacement unit; A plurality of second control signal channels, communicating with the fourth transmission channel of the basic computing units in a plurality of skew diagonal directions of the systolic array; the second control signal channels are used for transmitting second control signals, and the second control signals are used for controlling that the column vectors of the second matrix input into the basic computing units are input in the column direction or the skew diagonal direction.
2. The device according to claim 1, characterized in that, The first control signal and the second control signal are related to the matrix multiplication mode; Wherein, the matrix multiplication mode includes: A general mode, used for performing matrix multiplication operations on the first matrix and the second matrix to obtain a first result matrix; A first sparse mode, used for performing matrix multiplication operations on the first matrix and the second matrix to obtain a second result matrix; wherein, in the general mode and the first sparse mode, the input manner of the second matrix into the basic computing units is different, and the manner of the second matrix input into the basic computing units is controlled by the second control signal; A second sparse mode, used for performing matrix multiplication operations on the first matrix and the second matrix to obtain a third result matrix, the row vectors in the first matrix are arranged in a skew diagonal form, and the second matrix is a dense matrix; wherein, in the general mode and the second sparse mode, the working state of the vector displacement unit is different, and the working state of the vector displacement unit is controlled by the first control signal.
3. The device according to claim 2, characterized in that, In the general mode, the first control signal is used for controlling the vector displacement unit not to work, the second control signal is used for controlling that the second matrix input into the basic computing units through the second transmission channel is input in the column direction, and each column of the basic computing units corresponds to one column in the second matrix.
4. The device according to claim 2, wherein In the first sparse mode, the first control signal is used to control the vector displacement unit not to work, and the second control signal is used to control the second matrix input to the basic computing unit through the fourth transmission channel to be input in the skew diagonal direction, and each column of the basic computing units in each skew diagonal direction corresponds to a column in the second matrix.
5. The device according to claim 2, characterized in that, In the second sparse mode, the first control signal is used to control the vector displacement unit to work, and the second control signal is used to control the column vector input of the second matrix input to the basic computing unit through the third transmission channel, and each column of the basic computing units corresponds to a column in the second matrix.
6. The device according to any one of claims 2 to 5, characterized in that, It further includes a control signal generation module, and the control signal generation module is connected to the multiple first control signal channels and the multiple second control signal channels; The control signal generation module is configured to generate the first control signal and the second control signal according to the matrix multiplication mode.
7. A matrix multiplication method based on a systolic array, characterized in that, Applied to the systolic array-based matrix multiplication device according to any one of claims 1 to 6, the method includes: Obtaining a plurality of row vectors of the first matrix, a plurality of column vectors of the second matrix, the first control signal, and the second control signal; the first control signal is used to control the working state of the vector displacement unit, and the second control signal is used to control the column vectors of the second matrix input to the basic computing unit to be input in the column direction or the skew diagonal direction; Inputting the plurality of row vectors of the first matrix into multiple rows of basic computing units one by one along the row direction of the systolic array; and, under the control of the first control signal and the second control signal, inputting the plurality of column vectors of the second matrix into multiple columns of basic computing units one by one along the column direction of the systolic array, or inputting them into multiple basic computing units in the skew diagonal direction one by one along the skew diagonal direction of the systolic array; Calculating the input vectors through the basic computing unit based on the systolic array to obtain the matrix multiplication result of the first matrix and the second matrix.
8. The method according to claim 7, wherein In the general mode, the first control signal is used to control the vector displacement unit not to work, and the second control signal is used to control the second matrix input to the basic computing unit through the second transmission channel to be input in the column direction, and each of the basic computing units in each skew diagonal direction corresponds to a column in the second matrix; the general mode is used to perform matrix multiplication on the first matrix and the second matrix to obtain a first result matrix; Or, In the first sparse mode, the first control signal is used to control the vector displacement unit not to work, and the second control signal is used to control the second matrix input to the basic computing unit through the fourth transmission channel to be input in the skew diagonal direction, and each of the basic computing units in each skew diagonal direction corresponds to a column in the second matrix; the first sparse mode is used to perform matrix multiplication on the first matrix and the second matrix to obtain a second result matrix; or, In the second sparse mode, the first control signal is used to control the operation of the vector displacement unit, and the second control signal is used to control the input of the column vectors of the second matrix input to the basic computing unit through the third transmission channel. Each basic computing unit corresponds to a column in the second matrix; the second sparse mode is used to perform matrix multiplication on the first matrix and the second matrix to obtain a third result matrix. The row vectors in the first matrix are arranged in a skew diagonal form, and the second matrix is a dense matrix.
9. The method according to claim 7 or 8, characterized in that, The method further includes: Storing the matrix multiplication result of the first matrix and the second matrix in the form of a dense matrix.
10. An electronic device, characterized in that, The electronic device includes the systolic array-based matrix multiplication device according to any one of claims 1 to 6; or, The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the systolic array-based matrix multiplication method according to any one of claims 7 to 9 is implemented.
Citation Information
Patent Citations
Hybrid QR decomposition-based least square FPGA solving device
CN101827044A
Matrix multiplier and processor
CN112434256A
Matrix multiplication device and method based on systolic array and electronic equipment
CN117908832A
Application specific integrated circuit accelerators
US10790828B1
Hardware accelerator for systolic matrix multiplication
US10915297B1