A general sparse matrix multiplication implementation method and device based on 2D systolic array
Patent Information
- Application Number
- CN202210847492.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-19
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2042-07-19
AI Technical Summary
除此之外,脉动数组中的数据流不能被选择操作中断
[0034] 1. This invention can accelerate sparse matrix multiplication on systolic arrays. By packing the sparsity of the matrix and working in conjunction with the reconfigured 2D systolic array, sparse matrix multiplication can be processed effectively.
Smart Images

Figure CN115328440B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to sparse matrix computation acceleration technology in processors, specifically to a general sparse matrix multiplication implementation method and apparatus based on a 2D systolic array. Background Technology
[0002] Sparse matrix computation is widely used in various applications, such as computational science, neural network training, and language processing models. However, directly calculating a sparse matrix as a dense matrix is a waste of computational power and storage space, as zero computation is meaningless. At the same time, sparsity provides the possibility of reducing computational overhead and accelerating computation. In the fields of machine learning and neural networks, sparsity is prevalent, mainly from two aspects: (1) activation sparsity comes from the relationship of activation functions, with a sparsity of about 50%. (2) weight sparsity, i.e., pruning algorithms and sparsity training algorithms, with a sparsity of up to 90%. The main computational overhead of neural network training is convolution operation (70% of training time). Converting the convolution operator into a matrix multiplication operator through the im2col method is a direction for hardware acceleration. A systolic array with a regular structure, uniform wiring, and high frequency allows data to flow in the computing unit and reduces the number of memory accesses. In addition, it has high data reusability and huge acceleration potential for matrix multiplication computation. The Google team designed a TPU with a systolic array as the core computing architecture. Furthermore, numerous works, such as those shown, demonstrate the great feasibility of pulsating arrays in accelerating matrix multiplication.
[0003] A systolic array is a hardware architecture in which many simple processing units are arranged according to a set pattern. Each processing unit is a multiply-accumulate unit, performing multiplication and addition in a single step. Data movement within a systolic array can only occur between adjacent processing units. This simple architecture offers good scalability, with system performance increasing proportionally to the array size. Therefore, systolic arrays are used to accelerate general matrix multiplication. In general, general matrix multiplication can perform almost any type of matrix multiplication. Even if the matrix is sparse, it can be computed as a dense matrix containing zeros. Sparse matrix multiplication is a matrix multiplication between sparse and dense matrices. General sparse matrix multiplication is a matrix multiplication between sparse and sparse matrices. Although the matrices in sparse matrix multiplication and general sparse matrix multiplication computations are sparse, they are stored as dense matrices. Currently, the main idea behind most sparse acceleration algorithms is to compress the sparse matrix into an indexed dense matrix and then use the index information to select non-zero values in the sparse matrix for computation. Unfortunately, the regularity of systolic arrays means that supporting such flexible selection operations is virtually impossible. In addition, the data flow in a systolic array cannot be interrupted by a selection operation. Therefore, supporting sparse computation in a systolic array is a challenge. Furthermore, balancing execution efficiency and on-chip resources is another challenge. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a general sparse matrix multiplication method and apparatus based on 2D systolic arrays to address the above-mentioned problems in the prior art. The present invention can accelerate sparse matrix multiplication on systolic arrays, effectively process sparse matrix multiplication, and ensure effectiveness, efficiency and performance under different resource constraints.
[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0006] A general sparse matrix multiplication method based on a 2D systolic array, wherein the processing steps of each processing unit in the 2D systolic array include:
[0007] S1. First load the compression matrix C*, then receive the input compression matrix A* and compression matrix B*. The compression matrix A*, compression matrix B* and compression matrix C* are all recorded in the format of index and index corresponding value. The index is the index of the corresponding value in the original matrix before zero value compression. The original block matrix is obtained by dividing the input matrix into blocks according to the pulsation scale of the 2D pulsation array.
[0008] S2. Determine whether the index of compression matrix A* or compression matrix B* is zero. If it is true, pass compression matrix A* and compression matrix B* directly to the next processing unit and exit; otherwise, jump to step S3.
[0009] S3. Concatenate the indices of compressed matrix A* and compressed matrix B* to obtain the concatenation index. Determine whether the concatenation index and the index of compressed matrix C* are the same. If they are the same, calculate A*×B*+C* and output the calculation result, and pass matrix A* and matrix B* to the next processing unit. Otherwise, calculate A*×B* and output the calculation result, and pass matrix A* and matrix B* to the next processing unit, and exit.
[0010] Optionally, in step S1, among the compression matrices A*, B*, and C*, the values of any compression matrix are all non-zero values or contain some zero values. When it contains some zero values, the index corresponding to the zero value is 0, and the values with index 0 are skipped during calculation.
[0011] Optionally, the compression matrix A input in step S1 * To obtain the compressed matrix A by zero-value compression of the zero values in the non-common dimension directions of the original matrix A and the original matrix B. * The index is the index of the value in the direction of the non-common dimension with the original matrix B, and the compressed matrix B *To obtain the compressed matrix B by zero-value compression of the zero values in the non-common dimension directions of the original matrix A. * The index is the index of the value in the direction of the non-common dimension with the original matrix A, and the compressed matrix C * The index is determined by the compression matrix A * index, compression matrix B * The indexes are concatenated to form the index.
[0012] Optionally, the zero-value compression of the original matrix A with respect to the zero values in the non-common dimension direction of the original matrix B refers to the zero-value compression of the original matrix A in the row direction, and the zero-value compression of the original matrix B with respect to the zero values in the non-common dimension direction of the original matrix A refers to the zero-value compression of the original matrix B in the column direction.
[0013] Optionally, the zero-value compression includes:
[0014] Step 1: Mark the zero values of the compressed matrix with marker 0 and the non-zero values of the compressed matrix with marker 1, and generate a mask using markers 0 and 1; generate a set of vertices for the non-common dimensions of the compressed matrix, where each vertex represents a non-common dimension of the matrix; generate an edge for the relationship between any two non-common dimensions of the compressed matrix, and define the edge corresponding to the two non-common dimensions as a valid edge when the two non-common dimensions have a non-zero value in the same common dimension, and generate a two-dimensional adjacency matrix to represent the relationship between the vertices.
[0015] Step 2: Group the vertex set according to the adjacency matrix to obtain the group number of each vertex in the vertex set, and generate a grouped vertex set; the grouping includes: (a) initializing the variable Ci for the group number; (b) assigning the group number Ci to the vertex with the highest degree among the ungrouped vertices in the vertex set, and combining the sets of vertices that are not adjacent to the vertex with the highest degree and are not grouped into a set of vertex that can be grouped together; (c) traversing the set of vertex that can be grouped together, adding the traversed vertices to group Ci, generating a set of vertex that can be grouped together for each vertex added, until the set of vertex that can be grouped together is empty; (d) updating the ungrouped vertices in the vertex set. If the ungrouped vertices in the vertex set are not empty, increment the variable Ci by 1 and jump to step (b); otherwise, determine that the grouping is complete.
[0016] Step 3: Merge rows / columns with the same group number into a new row / column to obtain the matrix after zero-value compression of the compressed matrix.
[0017] Optionally, the zero-value compression includes: inputting the matrix to be compressed row by row / column into a window queue along a non-common dimension direction. The window queue is a priority queue, and the parameters of the window queue include width, depth, and threshold. The width is the same as the number of data contained in the compressed row / column, the depth is the number of compressed rows / columns that the queue can hold, and the threshold is the maximum number of rows / columns to be merged during the compression process. Whenever a row / column is input into the window queue along a non-common dimension direction, it is determined whether the newly entered row / column can be merged with the rows / columns already stored in the window queue. If the data at each position does not conflict (meaning they are not both non-zero values at the same time), the two rows / columns are merged into one row / column. Otherwise, the newly entered row / column is inserted into another empty slot in the window queue. When the number of row / column mergings in the window queue reaches the threshold, the row / column is popped from the queue. When a new row / column enters the queue and there are no empty slots in the queue, the row / column with the most merges is popped from the queue. When no data is flowing in, all remaining rows / columns in the queue are popped.
[0018] Furthermore, the present invention also provides a general sparse matrix multiplication implementation apparatus for applying the aforementioned general sparse matrix multiplication implementation method based on 2D pulsating arrays, comprising:
[0019] The preprocessing module is used to divide the original input matrix into block matrices according to the scale of a 2D systolic array, and then use the block matrices as compressed matrices to perform zero-value compression and input the compressed data stream into the on-chip buffer.
[0020] An on-chip buffer is used to cache the compressed data stream of the input 2D systolic array and the compressed data obtained by the 2D systolic array. The compressed data stream includes a compressed matrix A* and a compressed matrix B*. The compressed data obtained by the 2D systolic array includes a compressed matrix C* and a final output compressed matrix Y*. The compressed matrix Y* is recorded in the format of an index and the corresponding index value.
[0021] A 2D pulsating array comprising N rows and N columns of processing units, the processing units being used to perform multiply-accumulate operations based on compression matrices A* and B* flowing in from the vertical or horizontal direction and a pre-loaded compression matrix C*;
[0022] The post-processing module is used to write the final output compressed matrix Y* from the compressed data in the on-chip buffer back to the off-chip memory, and restore the compressed matrix Y* to the uncompressed original matrix Y according to the index information.
[0023] The processing unit in the 2D pulse array includes:
[0024] The register set is used to cache the compressed matrix A output from the previous processing unit. * Compression matrix B* And multiple matrices C*0{0…n}, and output the compressed matrix A to the next processing unit. * Matrix B * Where matrix C*0{0…n} is the compression matrix A * A row and compression matrix B * The partial sum obtained by multiplying a series of terms;
[0025] A multiplier is used to compress matrix A. * Compression matrix B * Perform multiplication operations;
[0026] The first selector is used to select a matrix C* from {0…n}. 0i where i∈{0…n};
[0027] An adder is used to combine the multiplication result of the multiplier with the matrix C* output by the selector. 0i Perform an addition operation;
[0028] The second selector is used to select one of the addition operation result output from the adder and the selection result of the first selector as the output to the next processing unit of multiple matrices C*1{0…n}.
[0029] Control logic, used to determine the compression matrix A * Compression matrix B * Compression matrix C * The index is generated to generate control signals for selecting matrix C*0{0…n} to control the first selector, and to determine whether the index of compressed matrix A* or compressed matrix B* is zero. If it is, compressed matrix A* and matrix B* are directly passed to the next processing unit; if it is not, the indices of compressed matrix A* and compressed matrix B* are concatenated to obtain a concatenated index. It is then determined whether the concatenated index and the index of compressed matrix C* are the same. If they are the same, the multiplier and adder are called to calculate A*×B*+C*, and compressed matrix A* and compressed matrix B* are passed to the next processing unit. If they are not the same, the multiplier is called to calculate A*×B*, and compressed matrix A* and compressed matrix B* are passed to the next processing unit.
[0030] Optionally, the on-chip buffer, 2D systolic array, and post-processing module are integrated on the same chip, and the pre-processing module is embedded on the chip or located off the chip.
[0031] In addition, the present invention also provides an accelerator chip, including a chip body and a 2D systolic array embedded in the chip body, wherein the 2D systolic array is the general sparse matrix multiplication implementation device based on the 2D systolic array.
[0032] Furthermore, the present invention also provides a computer device including a microprocessor and a memory interconnected, characterized in that the microprocessor is programmed or configured to perform the steps of the general sparse matrix multiplication implementation method based on a 2D systolic array.
[0033] Compared with the prior art, the present invention has the following main advantages:
[0034] 1. This invention can accelerate sparse matrix multiplication on systolic arrays. By packing the sparsity of the matrix and working in conjunction with the reconfigured 2D systolic array, sparse matrix multiplication can be processed effectively.
[0035] 2. This invention uses a pulsating array as the computing core to reorder and compress sparse matrices, thereby reducing on-chip storage overhead, bandwidth, and communication costs.
[0036] 3. The method of the present invention can adapt to different scales of computing cores, bandwidth and resource constraints, and can bring out the acceleration potential of sparse matrix multiplication under limited resources, ensuring effectiveness, efficiency and performance under different resource constraints. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the basic process of the method in Embodiment 1 of the present invention.
[0038] Figure 2 This is a sparse data representation method in Embodiment 1 of the present invention.
[0039] Figure 3 This is an example of the pulse sparse motion calculation process in Embodiment 1 of the present invention.
[0040] Figure 4 This is a flowchart of the zero-value compression method in Embodiment 1 of the present invention.
[0041] Figure 5 This is a schematic diagram illustrating the principle of sparse matrix compression and multiplication implementation in Embodiment 1 of the present invention.
[0042] Figure 6 This is a schematic diagram illustrating the implementation principle of pulsation-based sparse matrix multiplication in Embodiment 1 of the present invention.
[0043] Figure 7 This is a schematic diagram of the device structure in Embodiment 1 of the present invention.
[0044] Figure 8 This is a schematic diagram of the processing unit in Embodiment 1 of the present invention.
[0045] Figure 9 This is an example of the microstructure and data flow of a single processing unit in Embodiment 1 of the present invention.
[0046] Figure 10This refers to the compression ratio of different block sizes under different sparsity conditions in Embodiment 1 of the present invention when there is no threshold.
[0047] Figure 11 This is an example of the composition of the compression matrix in Embodiment 1 of the present invention.
[0048] Figure 12 This is a visualization example of the compression results when different threshold settings are used in Embodiment 1 of the present invention.
[0049] Figure 13 The normalized computation time of the 8×8 systolic array is compared with that of the original systolic array.
[0050] Figure 14 This compares the normalized computation time of the 16×16 systolic array with that of the original systolic array.
[0051] Figure 15 Comparison of the speedup ratio of general sparse matrix multiplication for a 4096×4096 matrix with that of GPU.
[0052] Figure 16 Comparison of the speedup of general sparse matrix multiplication with GPU for a matrix with a sparsity of 0.99.
[0053] Figure 17 This is an example of the window queue processing procedure in Embodiment 2 of the present invention. Detailed Implementation
[0054] Example 1:
[0055] like Figure 1 As shown, this embodiment provides a general sparse matrix multiplication implementation method based on a 2D systolic array. The processing steps of each processing unit in the 2D systolic array include:
[0056] S1. First load the compression matrix C*, then receive the input compression matrix A* and compression matrix B*.
[0057] S2. Determine whether the index of compression matrix A* or compression matrix B* is zero. If it is true, pass compression matrix A* and compression matrix B* directly to the next processing unit and exit; otherwise, jump to step S3.
[0058] S3. Concatenate the indices of compressed matrix A* and compressed matrix B* to obtain the concatenation index. Determine whether the concatenation index and the index of compressed matrix C* are the same. If they are the same, calculate A*×B*+C* and output the calculation result, and pass matrix A* and matrix B* to the next processing unit. Otherwise, calculate A*×B* and output the calculation result, and pass matrix A* and matrix B* to the next processing unit, and exit.
[0059] like Figure 2As shown, compressed matrices A*, B*, and C* are all recorded using an index and its corresponding numerical value format. The index is the index of the corresponding numerical value in the original matrix before zero-value compression. The original block matrix is obtained by dividing the input matrix into blocks according to the pulsation scale of a 2D pulsation array. In this embodiment, the "index + numerical value" method generates a data stream that satisfies two-dimensional pulsation, ensuring the accuracy of matrix multiplication calculations. Besides storing non-zero elements, it is also necessary to store the index information of the corresponding non-zero elements. The index is used to select the elements to be multiplied, ensuring the validity of the multiplication result (i.e., the result of the multiplication is the value in the target matrix). Under this compression method, for the input matrices to be multiplied, A* and B*, only the index information of the non-common dimensions needs to be recorded; the complete index information (rows and columns) does not need to be recorded. Compared to the coordinate storage method, this reduces the index overhead by half. For the input cumulative matrix C* and the output result matrix Y*, two-dimensional coordinates need to be stored, and the index overhead is consistent with the coordinate storage method. For an input raw matrix, the processing steps include: first, generating a block matrix according to the size of the systolic array; during preprocessing of the block matrix, compression is performed as needed along the row or column direction; after processing, a compressed matrix (compressed matrix A*, compressed matrix B*) is output. The obtained compressed matrix will then be used... Figure 2 The data is stored in the specified format.
[0060] In this embodiment, among the compression matrices A*, B*, and C* mentioned in step S1, the values of any compression matrix are all non-zero values or contain some zero values. When it contains some zero values, the index corresponding to the zero value is 0, and the value with index 0 is skipped during calculation.
[0061] In this embodiment, the compression matrix A input in step S1 * To obtain the compressed matrix A by zero-value compression of the zero values in the non-common dimension directions of the original matrix A and the original matrix B. * The index is the index of the value in the direction of the non-common dimension with the original matrix B, and the compressed matrix B * To obtain the compressed matrix B by zero-value compression of the zero values in the non-common dimension directions of the original matrix A. * The index is the index of the value in the direction of the non-common dimension with the original matrix A, and the compressed matrix C * The index is determined by the compression matrix A * index, compression matrix B * The indexes are concatenated to form the index.
[0062] In this embodiment, zero-value compression of the zero values of the original matrix A and the original matrix B in the non-common dimension direction means zero-value compression of the original matrix A in the row direction, and zero-value compression of the zero values of the original matrix B and the original matrix A in the non-common dimension direction means zero-value compression of the zero values of the original matrix B in the column direction.
[0063] Figure 3 This is an example of the pulse sparse motion calculation process in this embodiment. An example of the process of multiplying the compressed matrix A* and the compressed matrix B* is shown below. Figure 3 As shown. First, the input data is represented in the format of "index + data". Specifically, in the compression matrix A*, the underlined black numbers represent row numbers (starting from 1), and the black numbers represent numerical values. Note that zero-value data has a row number of 0, the purpose of which is to skip zero-value calculations in the reconstructed processing unit. In the compression matrix B*, the underlined gray numbers represent column numbers, and the black numbers represent numerical values. In the first step, the compression matrix B* is loaded into the systolic array, and the tilted compression matrix A* flows into the calculation. Steps two through five are the process of systolic calculation. In the third step, the first set of output data, "43 3, 11 2", flows out, indicating that the value in the 4th row and 3rd column of the result matrix is 3, and the value in the 1st row and 1st column is 2. This is restored to obtain the output data. In the fifth step, the calculation is completed, and the calculation result is obtained, along with... Figure 2 The middle C is consistent.
[0064] Based on the zero-value compression principle mentioned above, the required zero-value compression method can be implemented as needed. In this embodiment, considering the needs of offline environments (where data generation and consumption occur on-chip and are constrained by on-chip resources), a zero-value compression method for offline environments is provided. Specifically, we abstract a sparse matrix A as an undirected graph G. The non-common dimension is abstracted as a vertex set V, where each vertex represents a row or column of matrix A, and the relationship between two rows / columns is abstracted as an edge set E. An edge is defined as valid when two rows / columns have non-zero values in the same column / row. Then, we divide V into K groups, each forming an independent set with no adjacent vertices. Our goal is to obtain the minimum value of K. At this point, we transform the problem into an optimal graph coloring problem.
[0065] like Figure 4 As shown, the zero-value compression for offline environments in this embodiment includes:
[0066] Step 1: Mark the zero values of the compressed matrix with marker 0 and the non-zero values of the compressed matrix with marker 1, and generate a mask using markers 0 and 1; generate a set of vertices for the non-common dimensions of the compressed matrix, where each vertex represents a non-common dimension of the matrix; generate an edge for the relationship between any two non-common dimensions of the compressed matrix, and define the edge corresponding to the two non-common dimensions as a valid edge when the two non-common dimensions have a non-zero value in the same common dimension, and generate a two-dimensional adjacency matrix to represent the relationship between the vertices.
[0067] Step 2: Group the vertex set according to the adjacency matrix to obtain the group number of each vertex in the vertex set, and generate a grouped vertex set; the grouping includes: (a) initializing the variable Ci for the group number; (b) assigning the group number Ci to the vertex with the highest degree among the ungrouped vertices in the vertex set, and combining the sets of vertices that are not adjacent to the vertex with the highest degree and are not grouped into a set of vertex that can be grouped together; (c) traversing the set of vertex that can be grouped together, adding the traversed vertices to group Ci, generating a set of vertex that can be grouped together for each vertex added, until the set of vertex that can be grouped together is empty; (d) updating the ungrouped vertices in the vertex set. If the ungrouped vertices in the vertex set are not empty, increment the variable Ci by 1 and jump to step (b); otherwise, determine that the grouping is complete.
[0068] Step 3: Merge rows / columns with the same group number directly into a new row / column to obtain the compressed matrix after zero-value compression. Taking matrix A as an example, finally filter and compress the data in A according to the group number to obtain the compressed matrix A*. Specifically, since vertices with the same group number do not conflict, that is, two rows / columns will not both have non-zero values at the same position in the corresponding column / row direction, rows / columns with the same group number can be directly merged into a new row / column.
[0069] Because the pulsating data stream flows unidirectionally with only one data outlet, when the compressed row / column flows into the pulsating stream for computation, multiple partial sums will exist in the same direction. To ensure the correctness of the result, these partial sums cannot be merged (due to different indices). Furthermore, the more vertices in a group, the more resources are required to ensure correctness. For example, if a compressed row is formed by merging four rows before compression, in the pulsating data stream, the row before compression produces four partial sums in four different row directions, while the compressed row produces four partial sums in only one row direction. While computational efficiency is improved, additional resources are needed to ensure correctness. During compression and grouping, we add a threshold to the input parameters to control the number of vertices in a group. By adjusting the threshold, we can control the number of rows / columns merged during the compression algorithm. When resources are sufficient, the threshold can be increased to improve merging and computational efficiency. Conversely, when resources are limited, a lower threshold is set, sacrificing some merging and computational efficiency to ensure correctness.
[0070] The basic idea of the general sparse matrix multiplication implementation method based on 2D pulsating arrays in this embodiment is as follows: First, consider the calculation rules of matrix multiplication. For a matrix A of size M×K and a matrix B of size K×N, multiplying them results in a matrix Y of size M×N. For the data a(i,k) in the i-th row and k-th column of A, multiply it with the data b(k,j) in the corresponding k-th row and k-th column of B to obtain the result y(i,j), which is the data in the i-th row and j-th column of matrix Y. To ensure the validity of the result y, the column number of data a(i,k) and the row number of data b(k,j) must be the same (both are k), while the row number i of a(i,k) and the column number j of b(k,j) can be arbitrary. In other words, to ensure the validity of the result, the common dimensions (columns of a(i,k) and rows of b(k,j)) of a(i,k) and b(k,j) cannot change; only the non-common dimensions (rows of a(i,k) and columns of b(k,j)) can be adjusted in order. When recording the indices of the non-common dimensions, even if the order of multiplication is adjusted, the correct result can still be obtained. Therefore, the compression algorithm compresses the matrices based on their non-common dimensions; matrix A is compressed by rows, and matrix B is compressed by columns. For example... Figure 5 As shown, matrix A is a 4×2 matrix, and matrix B is a 2×4 matrix. Multiplying them yields a 4×4 matrix Y. The diagram uses patterns to represent row and column information; identical patterns indicate data in the same row / column. Four different patterns represent the four rows of matrix A and the four columns of matrix B. During the compression algorithm, matrix A is compressed row-wise, merging rows 1 and 4. Row 3, lacking non-zero values, is filtered out and not stored, resulting in the compressed matrix A*, reduced to 2×2. Matrix B is compressed column-wise, merging columns 1 and 3. Column 2, lacking non-zero values, is filtered out, resulting in the compressed matrix B*, also reduced to 2×2. When multiplying the compressed matrices A* and B*, element-wise multiplication yields different combinations; the results of the same combination are accumulated. Y* retains only the non-zero values, and matrix Y is reconstructed based on the row and column information. Figure 6 The image shows the... Figure 5The matrix multiplication process involves mapping a fixed input data stream to a 2×2 two-dimensional systolic array for computation. When calculating Y = A*B, matrix B is first divided into two matrices, B0 and B1. Matrix B0 is loaded into the systolic array. After transposing matrix A, the data is skewed and then flows into the systolic array to be calculated with B0, yielding a partial result Y0. Then, B1 is loaded into the systolic array and calculated with A, yielding a partial result Y1. Y0 and Y1 are concatenated to obtain the final result Y, thus completing the matrix calculation. When performing multiplication of A* and B*, matrix B* is first loaded into the systolic array. After transposing matrix A*, the data flows skewed into the systolic array. The reconstructed processing unit determines whether to accumulate the results based on the index, ultimately obtaining Y*. In both matrix multiplication calculations, A×B has a total computation time of 64 FLOPs, 4 memory accesses (B0, A, B1, A), and a computation time of 12 steps (6 steps for a single block). A*×B* has a total computation time of 16 FLOPs, 2 memory accesses (B*, A*), and a computation time of 4 steps. Compared to traditional systolic arrays, the use of compression algorithms and reconstruction processing units achieves a total reduction of 4 times in computation, 2 times in memory access, and 3 times in time.
[0071] like Figure 7 As shown, this embodiment also provides a general sparse matrix multiplication implementation device for applying the general sparse matrix multiplication implementation method based on 2D pulsating arrays described above, including:
[0072] The preprocessing module is used to divide the original input matrix into block matrices according to the scale of a 2D systolic array, and then use the block matrices as compressed matrices to perform zero-value compression and input the compressed data stream into the on-chip buffer.
[0073] An on-chip buffer is used to cache the compressed data stream of the input 2D systolic array and the compressed data obtained by the 2D systolic array. The compressed data stream includes a compressed matrix A* and a compressed matrix B*. The compressed data obtained by the 2D systolic array includes a compressed matrix C* and a final output compressed matrix Y*. The compressed matrix Y* is recorded in the format of an index and the corresponding index value.
[0074] A 2D pulsating array comprising N rows and N columns of processing units, the processing units being used to perform multiply-accumulate operations based on compression matrices A* and B* flowing in from the vertical or horizontal direction and a pre-loaded compression matrix C*;
[0075] The post-processing module is used to write the final output compressed matrix Y* from the compressed data in the on-chip buffer back to the off-chip memory, and restore the compressed matrix Y* to the uncompressed original matrix Y according to the index information.
[0076] like Figure 8As shown, the processing unit in the 2D pulsating array includes:
[0077] The register set is used to cache the compressed matrix A output from the previous processing unit. * Compression matrix B * And multiple matrices C*0{0…n}, and output the compressed matrix A to the next processing unit. * Matrix B * Where matrix C*0{0…n} is the compression matrix A * A row and compression matrix B * The partial sum obtained by multiplying a series of terms;
[0078] A multiplier is used to compress matrix A. * Compression matrix B * Perform multiplication operations;
[0079] The first selector is used to select a matrix C* from {0…n}. 0i where i∈{0…n};
[0080] An adder is used to combine the multiplication result of the multiplier with the matrix C* output by the selector. 0i Perform an addition operation;
[0081] The second selector is used to select one of the addition operation result output from the adder and the selection result of the first selector as the output to the next processing unit of multiple matrices C*1{0…n}.
[0082] Control logic, used to determine the compression matrix A * Compression matrix B * Compression matrix C * The index is generated to generate control signals for selecting matrix C*0{0…n} to control the first selector, and to determine whether the index of compressed matrix A* or compressed matrix B* is zero. If it is, compressed matrix A* and matrix B* are directly passed to the next processing unit; if it is not, the indices of compressed matrix A* and compressed matrix B* are concatenated to obtain a concatenated index. It is then determined whether the concatenated index and the index of compressed matrix C* are the same. If they are the same, the multiplier and adder are called to calculate A*×B*+C*, and compressed matrix A* and compressed matrix B* are passed to the next processing unit. If they are not the same, the multiplier is called to calculate A*×B*, and compressed matrix A* and compressed matrix B* are passed to the next processing unit.
[0083] for Figure 6The multiplication of compressed matrices A* and B* shown can be calculated using a 2×2 systolic array. However, when obtaining C*, multiple valid outputs may occur in a single processing unit. To prevent multiple valid data from being accumulated into a single output value, the processing unit needs to be reconstructed. The original processing unit only requires three registers, A, B, and C, to complete the output Y = A×B + C. The reconstructed processing unit increases the size of the register storing C data and adds control logic for processing the index. Figure 7 The diagram shows a 16×16 pulsating array. Taking processing unit (15,1) as an example, when processing C*1 = A*×B* + C*0, for multiple possible input data C*0{0…n}, the register size is increased for storage. Before accumulation, the control logic selects the C* to be accumulated with A*×B* based on the indices of the compression matrix A*, B*, and C*. 0i (i∈{0…n}), after performing the addition operation, the control logic writes the result to the corresponding C*1 position, and then outputs all C*1s to the next adjacent processing unit. The control logic also has another function: when the compressed matrices A* and B* flow into the processing unit, if the index of any data is 0, then the compressed matrices A* and B* do not flow into the multiplier, but flow out to the adjacent processing unit together with the input compressed matrix C*, saving the overhead of multiplication and addition operations. Figure 7 The reconstructed processing units are arranged in a 16×16 two-dimensional systolic array. Data exchange occurs through an on-chip data buffer, which in turn exchanges data with off-chip high-bandwidth memory. Compression algorithms can preprocess the raw data either off-chip or on-chip. During off-chip processing, the generated compressed data stream is written to off-chip high-bandwidth memory, then read from and written to the on-chip buffer, and finally flows into the systolic array for computation. During on-chip processing, a preprocessing module is added between the high-bandwidth memory and the on-chip data buffer to perform data compression, and then data is read from the on-chip data buffer and flows into the systolic array. The post-processing module restores the data by writing the compressed data from the on-chip data buffer back to the high-bandwidth memory.
[0084] Under the control of the control logic, the processing unit's processing of the input compression matrix includes calculation and transmission. Calculation is optional, while transmission is mandatory. It is possible to choose not to calculate and only transmit, or to calculate and transmit simultaneously. Calculation includes multiply-add (calculating A*×B*+C*) and multiply (calculating A*×B*). This embodiment supports two data flow patterns: fixed input and fixed output. Fixed input means that one of the multiplication matrices is pre-fixed in the systolic array, while the other input matrix flows into the systolic array horizontally / vertically. In this case, the output result flows out of the systolic array vertically / horizontally. Similarly, fixed output means that the two input matrices flow into the systolic array horizontally and vertically respectively, and the output result is fixed in the systolic array. For example... Figure 9 To compute C* in two data stream modes out =A*×B*+C* in The diagram illustrates this. In fact, sparse matrices are difficult to compress into completely dense matrices. Even after preprocessing and compression, some zero-value data will still flow into the computation unit. To address this, we adjusted the processing unit. For zero-value data, the processing unit will not perform computational logic, thereby reducing computational overhead, but will simply pass it to the next adjacent unit. Figure 9 a and Figure 9 In the diagram, 'b' represents the input fixed data stream within the processing unit. Data B* is pre-loaded into the processing unit, A* flows in horizontally, and C*... in and C* out Flowing in and out vertically. For example... Figure 9 As shown in 'a', when A* and B* are non-zero values, the subsequent calculation logic will continue to be executed; in Figure 9 In 'b', since A* or B* is zero, no calculation logic is executed (gray line), only the passing logic is executed. Figure 9 c and Figure 9 In the diagram, 'd' represents the fixed data stream output within the processing unit. Data C* is pre-loaded into the processing unit, with A* flowing in horizontally and B* flowing in vertically. Similarly, whether computational logic is executed depends on whether A* or B* is zero.
[0085] In this embodiment, the on-chip buffer, 2D systolic array, and post-processing module are integrated on the same chip. In addition, it should be noted that the pre-processing module can be embedded on the chip or located off the chip as needed.
[0086] In this embodiment, the performance of the zero-value compression method and the performance of the systolic array with the support of the algorithm and reconstruction processing unit are analyzed. First, we need to clarify that for matrix multiplication larger than the systolic array size, the input matrix needs to be partitioned before being fed into the systolic array for computation. The partitioning strategy affects the overall computation time and other performance aspects. In our work, we do not discuss the performance of the partitioning strategy; we use the same partitioning strategy on both the original and reconstructed systolic arrays to ensure fairness in the comparison. Therefore, the following analysis uses the partitioned matrix that can be fed into the systolic array for computation as the input matrix. Let matrix A be an M×K partitioned matrix with sparsity S, the number of non-zero elements NZ, the data width dataWidth, the storage overhead Mb, and the systolic array size K×K.
[0087] After compressing matrix A by its row dimensions, we obtain a compressed matrix A* of size R×K with sparsity S. c The index width is idxWidth, and the storage overhead is M. C Then the following relationship holds:
[0088] NNZ=M*K*(1-S)=R*K*(1-Sc)≤R*K,
[0089] Mb = M * K * dataWidth
[0090] Mc=R*K*(dataWidth+idxWidth),
[0091] idxWidth = log2M,
[0092] In the above formula, NNZ represents the number of non-zero elements, M×K is the size of the block matrix, S is the sparsity of the block matrix, R×K is the size of the compression matrix A*, and S c Let M be the sparsity of the compression matrix A*, Mb be the storage overhead, and dataWidth be the data bit width. C Where idxWidth represents storage overhead, and idxWidth represents the index bit width. The compression ratio comRatio of matrix A can be obtained as:
[0093]
[0094] Assuming that the block size is set to 256×K during execution, and single-precision floating-point numbers are used for storage and calculation, a storage saving of 1.6 times can be achieved when the sparsity is 0.5.
[0095] For systolic arrays, the computation time is proportional to the amount of input data, equivalent to the size of the input matrix. In sparse matrix multiplication with only one-sided sparsity, let A be an M×K matrix with sparsity S and B be a K×N dense matrix. The computation time T of the systolic array for matrix A is... base satisfy:
[0096] T base ∝M×K×N.
[0097] If A* is a compressed R×K matrix, then the computation time T of the systolic array for the compressed matrix A* is... comp satisfy:
[0098] T comp ∝R×K×N,
[0099] Since the relationship between matrix A and compressed matrix A* satisfies:
[0100]
[0101] The speedup ratio can be obtained as:
[0102]
[0103] Furthermore, to further validate the method in this embodiment through experiments, we implemented a systolic array supporting sparse multiplication in Verilog RTL, and then used the compiler to estimate the chip area and total power consumption under the ASAP 7nm open-source PDK. We compared our design with a GPU (GeForce RT×2080Ti) using the cuSPARSE library in CUDA Toolkit 11.0 and a systolic array with the same number of processing units. We evaluated our design using randomly generated sparse matrices of different sizes, including the compression efficiency and multiplication performance of the sparse matrices. Additionally, we compared the performance of different CNNs, including VGG16, AlexNet, and ResNet18, on the CIFAR-10 and ImageNet datasets.
[0104] We used randomly generated 4096×4096 single-precision floating-point matrices with sparsity between 0.7 and 0.99 to test the thresholdless compression performance at different block sizes. Figure 10 As shown, when the sparsity is [0.7, 0.8, 0.9, 0.95, 0.99], the maximum potential compression ratio is [2.4, 3.5, 6.6, 12.3, 47.9] times. Generally speaking, the higher the sparsity, the larger the non-common dimension, and the smaller the common dimension, the higher the compression ratio. The experimental results are consistent with our analysis. Figure 11The composition of the compression matrix is shown, where (a) has no threshold setting and (b) has a threshold set to 4. Figure 11 Figure (a) shows the compositional analysis of a compressed matrix with a sparsity of 0.9 and a block size of 8×256. As can be seen from the figure, the highest ratio (labeled "5-in-1") of five rows from the original matrix being merged into one row in the compressed matrix is 36%. Generally, nearly 80% of the new rows in the compressed matrix come from combinations of four to six original rows merged into one new row, with only 9% of the rows remaining uncompressed.
[0105] Under the test conditions of a 4096×4096 matrix with a block size of 256×8 and a sparsity of 0.9, we set the threshold distribution to 2, 3, 4, and 8. Figure 12 The right-hand side of the diagram shows that when the threshold is set to 4, 95% of the rows in the compressed matrix are derived from the four rows of the original matrix. Furthermore, Figure 12 This is the result of visualizing a portion of the original matrix. When the threshold is set to 2, 3, 4, and 8, we can see... Figure 12 The right side of the image shows that 256 columns have been compressed to 130, 110, 95, and 39 columns (taking the maximum number of columns), with compression ratios close to each other. Therefore, the higher the threshold, the higher the density of non-zero values in the compression.
[0106] In addition, we present the number of additional buffers required by the processing unit under different threshold settings in Table 1.
[0107] Table 1: The amount of additional buffer required by the processing unit under different threshold settings.
[0108]
[0109] As shown in Table 1, the required number of additional buffers varies depending on the computational scenario. Specifically, in an 8×8 systolic array, the additional processing unit buffer required for sparse matrix multiplication is 2, 3, 4, and 8 for thresholds of 2, 3, 4, and 8, respectively, while the additional processing unit buffer required for general sparse matrix multiplication is 4, 8, 8, and 8, respectively. Regardless of the threshold, general sparse matrix multiplication requires more additional buffers than sparse matrix multiplication.
[0110] We performed matrix multiplication on the original systolic array and the modified systolic array, in which each processing unit has four additional buffers. Both systolic arrays are configured in two sizes: 8×8 and 16×16. Figure 13 and Figure 14The normalized computation time for performing sparse matrix multiplication and general sparse matrix multiplication is shown for the modified 8×8 and 16×16 systolic arrays as the sparsity increases from 0.3 to 0.9. Note that the sparsity of the two matrices in the general sparse matrix multiplication is the same. Compared to the original systolic array, the computation time is reduced for both configurations within the sparsity range of 0.5–0.9. When the sparsity is greater than 0.5, the computation time is reduced further. Figure 13 The 8×8 scale shown Figure 4 As shown, at a scale of 16×16, the modified systolic array achieves speedups of 1.4–4.2 times and 1.2–3.2 times in sparse matrix multiplication, respectively, and 2.0–7.1 times and 1.3–4.4 times in general sparse matrix multiplication, respectively. These experimental results are consistent with our analysis. For matrices with low sparsity, smaller-scale systolic arrays perform better. More importantly, even at low sparsity, the computation time of the modified systolic array does not exceed that of the original systolic array.
[0111] We tested the performance of the GPU and the modified systolic array in computing sparse matrix multiplication. To ensure a fair comparison, the GPU performance was reflected by the speedup ratio of cuSPARSE and cuBLAS, libraries for GPU-accelerated matrix multiplication. Similarly, the performance of the modified systolic array was also reflected by the speedup ratio compared to the original systolic array.
[0112] We configured the modified pulsation array with 256 processing units (16×16) and a matrix block size of 256×16. Figure 15 The speedup of the GPU and the modified systolic array is shown for 4096×4096 matrices with sparsity increasing from 0.3 to 0.99. At a sparsity of 0.99, the modified systolic array outperforms the GPU by 9.7 times. Even when processing low-sparse matrices, the modified systolic array maintains the same processing time as the original systolic array. If the sparsity is less than 0.98, cuSPARSE requires significantly more time. Furthermore, Figure 16 This paper describes the speedup performance of GPU and modified systolic array for square matrices ranging from 512 to 8192 at a sparsity of 0.99. The modified systolic array maintains a stable speedup across different scales, but cuSPARSE performs poorly with small matrices. Typically, cuSPARSE performs poorly, requiring larger matrices and higher sparsity (close to 0.98) to achieve good speedup benefits. However, in most cases, the modified systolic array does not yield negative gains.
[0113] Furthermore, we evaluated the neural network training in this embodiment. The RigL sparse training method was used to make the model's weights sparse, controlling the sparsity of the weights to 0.8 during the training phase. Additionally, due to the inherent ReLU function in the model, the sparsity of activations during training was between 0.3 and 0.6. We converted the convolution operation in the CNN into a matrix multiplication operation using im2ol and mapped it to a modified systolic array for acceleration. The results are shown in Figure 2.
[0114] Table 2: Comparison of floating-point computational complexity.
[0115]
[0116] Table 2 describes the computational complexity compared to dense floating-point computations. (See also...) Figure 2 It can be seen that the modified systolic array reduces the floating-point computation of VGG16, AlexNet and ResNet18 by 3.4 times, 3.2 times and 3.6 times respectively on the CIFAR-10 and ImageNet datasets.
[0117] In addition, Table 3 lists the area and power consumption of the raw systolic and sparse systolic processing units, which process 8-bit integer data (INT8) and 32-bit floating-point data (FP32), respectively.
[0118] Table 3: Statistics on processing unit area and power consumption.
[0119]
[0120]
[0121] As shown in Table 3, when processing INT8 data, due to the small data bit width, the index overhead is relatively high, resulting in a significant increase in sparse systolic area and power consumption, becoming 2.78 times and 2.87 times the original systolic area, respectively. However, when processing FP32, the index overhead decreases, and the sparse systolic area and power consumption become 1.28 times and 1.21 times the original systolic area, respectively. Furthermore, when performing the zero-value skipping operation, the power consumption of the sparse systolic is correspondingly reduced because no computational logic is used.
[0122] Table 4 shows the area and power consumption of sparse systolic arrays with sizes of 8×8 and 16×16 when processing INT8 and FP32 data types. We also provide the proportion of preprocessing and postprocessing.
[0123] Table 4: Statistics on sparse pulsation area and power consumption.
[0124]
[0125] See Figure 4It can be seen that, overall, the overhead of preprocessing and postprocessing is not high, accounting for no more than 7.5% in INT8 and no more than 3.5% in FP32. Experimental results show that for matrices with sparsity greater than 0.95, the general sparse matrix multiplication implementation method based on 2D systolic arrays in this embodiment achieves a compression ratio of more than 10 times; when the sparsity is 0.5 to 0.9, it improves the speed by 1.3 to 4.4 times compared to the baseline of matrix multiplication, and when the sparsity is 0.99, it is 9.7 times faster than cuSPARSE. In neural network training, we further reduced the floating-point computation by more than 3.2 times.
[0126] Furthermore, this embodiment also provides an accelerator chip, including a chip body and a 2D systolic array embedded within the chip body, wherein the 2D systolic array is the general sparse matrix multiplication implementation device based on the 2D systolic array described above. Additionally, this embodiment also provides a computer device, including a microprocessor and a memory interconnected, wherein the microprocessor is programmed or configured to execute the steps of the general sparse matrix multiplication implementation method based on the 2D systolic array described above.
[0127] Example 2:
[0128] This embodiment is basically the same as Embodiment 1, the main difference being the zero-value compression method used. The zero-value compression method used in this embodiment is geared towards online environments, where data generation and consumption occur off-chip and on-chip, respectively. Off-chip processing has more resources and fewer restrictions.
[0129] We propose a "window queue" design to meet the real-time generation and consumption requirements of on-chip data. The window queue is a priority queue-like structure. During preprocessing, sparse matrices flow into the window queue row by row / column. After relevant operations are performed, the sparsity of the rows / columns flowing out of the queue decreases, and there are fewer zero-value data. Specifically, zero-value compression in this embodiment includes: inputting the matrix to be compressed row by row / column into a window queue along non-common dimensions. The window queue is a priority queue, and its parameters include width, depth, and threshold. The width is the same as the number of data items contained in the compressed row / column, the depth is the number of compressed rows / columns that the queue can hold, and the threshold is the maximum number of rows / columns merged during the compression process. Whenever a row / column is input into the window queue along a non-common dimension, it is determined whether the newly entered row / column can be merged with the rows / columns already stored in the window queue. If the data at each position does not conflict (meaning they are not both non-zero values), the two rows / columns are merged into one row / column. Otherwise, the newly entered row / column is inserted into another empty slot in the window queue. When the number of row / column merges in the window queue reaches the threshold, the row / column is popped from the queue. When a new row / column enters the queue and there are no empty slots, the row / column with the most merges is popped from the queue. When no data enters, all remaining rows / columns in the queue are popped. Therefore, the operations mentioned above include: Enqueue operation: Determine if the newly entered row / column can be merged with existing rows / columns in the window queue. If the data at each position is non-conflicting (not simultaneously non-zero), merge the two rows / columns into one; otherwise, insert the newly entered row / column into another empty slot in the window queue. Dequeue operation: When the number of merges in the window queue reaches a threshold, pop that row / column from the queue; when a new row / column enters the queue and there are no empty slots, pop the row / column with the highest number of merges; when no data is entering, pop all remaining rows / columns from the queue. Figure 17 The diagram illustrates the process of a window queue with a width of 3, a depth of 2, and a threshold of 2 processing incoming data. The process includes: Step 1: Queue initialization; Step 2: Entry of the first column of data; Step 3: If new data conflicts with existing data in the queue, it is inserted into a new position; Step 4: The new data is merged with the data in the first column until the merging threshold is reached, then it is ready to be removed from the queue; Step 5: The new data is inserted into an empty space; Step 6: The data flow stops, and the remaining data is popped out sequentially.
[0130] In addition, this embodiment also provides a general sparse matrix multiplication implementation device for applying the general sparse matrix multiplication implementation method based on 2D pulsating array described above. The difference from Embodiment 1 is that the zero-value compression method used in its preprocessing module is different.
[0131] Furthermore, this embodiment also provides an accelerator chip, including a chip body and a 2D systolic array embedded within the chip body, wherein the 2D systolic array is the general sparse matrix multiplication implementation device based on the 2D systolic array described above. Additionally, this embodiment also provides a computer device, including a microprocessor and a memory interconnected, wherein the microprocessor is programmed or configured to execute the steps of the general sparse matrix multiplication implementation method based on the 2D systolic array described above.
[0132] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0133] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A general sparse matrix multiplication implementation method based on a 2D pulsating array, characterized in that, The processing steps for each processing unit in a 2D pulsating array include: S1. First, load the compression matrix C*, then receive the input compression matrix A* and compression matrix B*. The compression matrix A*, compression matrix B* and compression matrix C* are all recorded in the format of index and index corresponding value. The index is the index of the corresponding value in the original matrix before zero-value compression. The original matrix is obtained by dividing the input matrix into blocks according to the pulsation scale of the 2D pulsation array. S2. Determine whether the existence of a zero index in compression matrix A* or compression matrix B* is true. If true, directly pass compression matrix A* and compression matrix B* to the next processing unit and exit; otherwise, jump to step S3. S3. Concatenate the indices of compression matrix A* and compression matrix B* to obtain the concatenation index. Determine whether the concatenation index and the index of compression matrix C* are the same. If they are the same, calculate A*×B*+C* and output the calculation result, and pass compression matrix A* and compression matrix B* to the next processing unit. Otherwise, calculate A*×B* and output the calculation result, and pass compression matrix A* and compression matrix B* to the next processing unit, and exit. The compression matrix A input in step S1 * To obtain the compressed matrix A by zero-value compression of the zero values in the non-common dimension directions of the original matrix A and the original matrix B. * The index is the index of the value in the direction of the non-common dimension with the original matrix B, and the compressed matrix B * To obtain the compressed matrix B by zero-value compression of the zero values in the non-common dimension directions of the original matrix A. * The index is the index of the value in the direction of the non-common dimension with the original matrix A, and the compressed matrix C * The index is determined by the compression matrix A * index, compression matrix B * The index is concatenated to form the zero value. The zero value compression of the original matrix A with the non-common dimension direction of the original matrix B refers to the zero value compression of the original matrix A in the row direction. The zero value compression of the original matrix B with the non-common dimension direction of the original matrix A refers to the zero value compression of the original matrix B in the column direction.
2. The general sparse matrix multiplication implementation method based on 2D pulsating array according to claim 1, characterized in that, In step S1, among the three compression matrices A*, B*, and C*, the values of any compression matrix are all non-zero values or contain some zero values. When it contains some zero values, the index corresponding to the zero value is 0, and the value with index 0 is skipped during calculation.
3. The general sparse matrix multiplication implementation method based on 2D pulsating array according to claim 2, characterized in that, The zero-value compression includes: Step 1: Mark the zero values of the compressed matrix with marker 0 and the non-zero values of the compressed matrix with marker 1, and generate a mask using markers 0 and 1; generate a set of vertices for the non-common dimensions of the compressed matrix, where each vertex represents a non-common dimension of the matrix; generate an edge for the relationship between any two non-common dimensions of the compressed matrix, and define the edge corresponding to the two non-common dimensions as a valid edge when the two non-common dimensions have a non-zero value in the same common dimension, and generate a two-dimensional adjacency matrix to represent the relationship between the vertices. Step 2: Group the vertex set according to the adjacency matrix to obtain the group number of each vertex in the vertex set and generate a grouped vertex set; the grouping includes: (a) initializing the variable J of the group number; (b) assigning the group number J to the vertex with the largest degree among the ungrouped vertices in the vertex set, and combining the set of vertices that are not adjacent to the vertex with the largest degree and are not grouped into a set of vertex that can be grouped together; (c) traversing the set of vertex that can be grouped together, adding the vertices obtained from the traversal to group J, generating a set of vertex that can be grouped together for each vertex added, until the set of vertex that can be grouped together is empty; (d) updating the ungrouped vertices in the vertex set. If the ungrouped vertices in the vertex set are not empty, increment the variable J by 1 and jump to step (b); otherwise, determine that the grouping is complete. Step 3: Merge rows / columns with the same group number into a new row / column to obtain the matrix after zero-value compression of the compressed matrix.
4. The general sparse matrix multiplication implementation method based on 2D pulsating array according to claim 1, characterized in that, The zero-value compression process includes: inputting the matrix to be compressed row by row / column into a window queue along non-common dimensions. The window queue is a priority queue, and its parameters include width, depth, and a threshold. The width is the same as the number of data items in the compressed row / column, the depth is the number of compressed rows / columns that the queue can hold, and the threshold is the maximum number of rows / columns that can be merged during compression. Each time a row / column is input into the window queue along a non-common dimension, it is determined whether the newly input row / column can be merged with the rows / columns already stored in the window queue. If the data at each position does not conflict (meaning they are not both non-zero values), the two rows / columns are merged into one row / column. Otherwise, the newly input row / column is inserted into another empty slot in the window queue. When the number of row / column mergings in the window queue reaches the threshold, the row / column is popped from the queue. When a new row / column enters the queue and there are no empty slots, the row / column with the most merges is popped from the queue. When no data is input, all remaining rows / columns are popped from the queue.
5. A general sparse matrix multiplication implementation apparatus for applying the general sparse matrix multiplication implementation method based on a 2D pulsating array as described in any one of claims 1 to 4, characterized in that, include: The preprocessing module is used to divide the original input matrix into block matrices according to the scale of a 2D systolic array, and then use the block matrices as compressed matrices to perform zero-value compression and input the compressed data stream into the on-chip buffer. An on-chip buffer is used to cache the compressed data stream of the input 2D systolic array and the compressed data obtained by the 2D systolic array. The compressed data stream includes a compressed matrix A* and a compressed matrix B*. The compressed data obtained by the 2D systolic array includes a compressed matrix C* and a final output compressed matrix Y*. The compressed matrix Y* is recorded in the format of an index and the corresponding index value. A 2D pulsating array comprising N rows and N columns of processing units, the processing units being used to perform multiply-accumulate operations based on compression matrices A* and B* flowing in from the vertical or horizontal direction and a pre-loaded compression matrix C*; The post-processing module is used to write the final output compressed matrix Y* from the compressed data in the on-chip buffer back to the off-chip memory, and restore the compressed matrix Y* to the uncompressed original matrix Y according to the index information. The processing unit in the 2D pulse array includes: The register set is used to cache the compressed matrix A output from the previous processing unit. * Compression matrix B * and multiple matrices C* 0{0…n} And output the compressed matrix A to the next processing unit. * Compression matrix B * , where matrix C* 0{0…n} For the compression matrix A * A row and compression matrix B * The partial sum obtained by multiplying a series of terms; A multiplier is used to compress matrix A. * Compression matrix B * Perform multiplication operations; The first selector is used to select from matrix C* 0{0…n} Choose a matrix C* 0i where i∈{0…n}; An adder is used to combine the multiplication result of the multiplier with the matrix C* output by the selector. 0i Perform an addition operation; The second selector is used to select one of the addition operation result output from the adder and the selection result from the first selector as the output of multiple matrices C* to the next processing unit. 1{0…n} ; Control logic, used to determine the compression matrix A * Compression matrix B * Compression matrix C * Index generation is used to select matrix C* 0{0…n} The control signal controls the first selector and determines whether the index of compression matrix A* or compression matrix B* is zero. If it is, compression matrix A* and matrix B* are directly passed to the next processing unit. If it is not, the indices of compression matrix A* and compression matrix B* are concatenated to obtain a concatenated index. The concatenated index is then compared with the index of compression matrix C*. If they are the same, the multiplier and adder are called to calculate A*×B*+C*, and compression matrix A* and compression matrix B* are passed to the next processing unit. If they are different, the multiplier is called to calculate A*×B*, and compression matrix A* and compression matrix B* are passed to the next processing unit.
6. The general sparse matrix multiplication implementation apparatus according to claim 5, characterized in that, The on-chip buffer, 2D systolic array, and post-processing module are integrated on the same chip, while the pre-processing module is embedded on the chip or located off the chip.
7. A computer device comprising a microprocessor and a memory interconnected, characterized in that, The microprocessor is programmed or configured to perform the steps of the general sparse matrix multiplication implementation method according to any one of claims 1 to 4.
Citation Information
Patent Citations
A systolic array architecture for sparse matrix operations
CN110851779A
Universal FPGA / ASIC matrix-vector multiplication architecture
US20140108481A1