Sparse matrix operation method, processor and computing equipment
By rearranging non-zero elements of sparse matrix and performing external product operations, the inefficiency problem caused by zero elements participating in sparse matrix multiplication is solved, and more efficient calculations are achieved.
Patent Information
- Application Number
- CN202311544688.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
There are a large number of zero elements involved in the multiplication of sparse matrix, resulting in low efficiency.
By obtaining the numerical matrix and index matrix corresponding to the sparse matrix, rearrange non-zero elements and perform external product operations to avoid zero elements participating in the operation.
The computational efficiency of sparse matrix multiplication is improved and the computational complexity is reduced.
Smart Images

Figure CN120020761A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method for operating sparse matrices, a processor, and a computing device. Background Art
[0002] Matrix multiplication is widely used in various scientific computing scenarios. Among them, the density of matrices involved in different computing scenarios varies. Matrix multiplication is mainly divided into dense matrix multiplication and sparse matrix multiplication. Generally, dense matrix multiplication is mainly applied to scenarios such as traditional machine learning and neural network model calculations, while sparse matrix multiplication is mainly applied to scenarios such as graph computing, graph neural networks, and multigrid solvers.
[0003] Since a sparse matrix includes a large number of zero elements, during the execution of sparse matrix multiplication, a large number of zero elements participate in the matrix multiplication operation, resulting in low efficiency of matrix operations. Summary of the Invention
[0004] Embodiments of this application provide a method for operating sparse matrices, a processor, and a computing device, which can improve the efficiency of sparse matrix multiplication. The corresponding technical solutions are as follows:
[0005] In a first aspect, a method for operating a sparse matrix is provided. This method is executed by a processor, which includes a sparse matrix processing unit and a sparse matrix operation unit. The method includes:
[0006] The sparse matrix processing unit obtains a first numerical matrix and a first index matrix corresponding to a first sparse matrix, and a second numerical matrix and a second index matrix corresponding to a second sparse matrix. Among them, the first numerical matrix is obtained by rearranging the non-zero elements in each column of the first sparse matrix, the first index matrix is used to indicate the row numbers of the non-zero elements in the first numerical matrix in the first sparse matrix, the second numerical matrix is obtained by rearranging the non-zero elements in each row of the second sparse matrix, and the second index matrix is used to indicate the column numbers of the non-zero elements in the second numerical matrix in the second sparse matrix. The sparse matrix processing unit inputs the non-zero elements in each column of the first numerical matrix and the row numbers indicated by each column in the first index matrix, the non-zero elements in each row of the second numerical matrix and the column numbers indicated by each row in the second index matrix, into the sparse matrix operation unit. The sparse matrix operation unit performs an outer product operation on the non-zero elements in each column of the first numerical matrix and the non-zero elements in each row of the second numerical matrix based on the row numbers indicated by each column in the first index matrix and the column numbers indicated by each row in the second index matrix, to obtain the result matrix of the matrix multiplication operation of the first sparse matrix and the second sparse matrix.
[0007] In the solution shown in this application, the outer product operation can be performed on the non-zero elements in the first sparse matrix according to the row numbers of the non-zero elements in the first numerical matrix indicated by the first index matrix and the column numbers of the non-zero elements in the second numerical matrix indicated by the second index matrix in the second sparse matrix, so as to obtain the result matrix of the matrix multiplication of the first sparse matrix and the second sparse matrix. In this way, a large number of zero elements in the first sparse matrix and the second sparse matrix can be avoided from participating in the matrix multiplication operation process, and thus the operation efficiency of performing the sparse matrix multiplication operation can be improved.
[0008] In an implementable manner, the first numerical matrix includes a plurality of first numerical partitions, and the plurality of first numerical partitions are obtained by dividing the first sparse matrix into a plurality of first sparse matrix regions row by row and then rearranging each column of non-zero elements in each first sparse matrix region. The first index matrix includes a first index partition corresponding to each first numerical partition. The second numerical matrix includes a plurality of second numerical partitions, and the plurality of second numerical partitions are obtained by dividing the second sparse matrix into a plurality of second sparse matrix regions column by column and then rearranging each row of non-zero elements in each second sparse matrix region. The second index matrix includes a second index partition corresponding to each second numerical partition. In this way, by performing partition calculations on the first numerical matrix, the first index matrix, the second numerical matrix, and the second index matrix, the complexity of matrix operations can be reduced, and thus the operation efficiency of matrix operations can be improved.
[0009] In an implementable manner, each first numerical partition includes a plurality of first numerical blocks with a size of n×m, each first index partition includes a plurality of first index blocks with a size of n×m, each second numerical partition includes a plurality of second numerical blocks with a size of m×n, and each second index partition includes a plurality of second index blocks with a size of m×n. In this way, by performing block calculations on the first numerical partition, the first index partition, the second numerical partition, and the second index partition, the complexity of matrix operations can be further reduced, and the operation efficiency of matrix operations can be improved.
[0010] In an implementable manner, the sparse matrix processing unit inputs the non-zero elements in each column of the first numerical matrix, the row numbers indicated by each column of the first index matrix, the non-zero elements in each row of the second numerical matrix, and the column numbers indicated by each row of the second index matrix to the sparse matrix operation unit, including: The sparse matrix processing unit inputs, in the calculation order, each column element in the non-zero numerical block of each column in each first numerical block in each first numerical partition, each row element in the non-zero numerical block of each row in each second numerical block in each second numerical partition, the row numbers indicated by each column in the non-zero index block of each column in each first index block in each first index partition, and the column numbers indicated by each column in the non-zero index block of each row in each second index block in each second index partition to the sparse matrix operation unit in sequence. In the solution shown in this application, by inputting the non-zero numerical block and the corresponding index block to the sparse matrix operation unit for accumulation processing, it is possible to avoid the sparse matrix operation unit from performing accumulation processing on all-zero numerical blocks and all-zero index blocks, reduce the computational amount of the accumulation processing, and thus improve the operation efficiency of matrix operations.
[0011] In an implementable manner, the sparse matrix operation unit includes a sparse matrix outer product subunit and a sparse matrix accumulation subunit. The sparse matrix operation unit performs an outer product operation on the non-zero elements in each column of the first numerical matrix and the non-zero elements in each row of the second numerical matrix based on the row numbers indicated by each column of the first index matrix and the column numbers indicated by each row of the second index matrix, to obtain the result matrix of performing matrix multiplication on the first sparse matrix and the second sparse matrix, including: The sparse matrix outer product subunit performs an outer product operation on each column element of the first numerical block and each row element of the second numerical block input each time, to obtain a first intermediate result matrix. The sparse matrix outer product subunit forms a first intermediate index matrix corresponding to the first intermediate result matrix by using the row numbers indicated by each column of the first index block input each time and the column numbers indicated by each column of the second index block. The first intermediate index matrix is used to indicate the position information of each element in the first intermediate result matrix in the result matrix. The sparse matrix accumulation subunit accumulates multiple first intermediate result matrices based on multiple first intermediate index matrices to obtain the result matrix of performing matrix multiplication on the first sparse matrix and the second sparse matrix.
[0012] In an implementable manner, the processor includes multiple matrix registers. The method further includes: For the first first intermediate result matrix output by the sparse matrix outer product subunit and the corresponding first intermediate index matrix, the sparse matrix processing unit stores each row intermediate result in the first intermediate result matrix and the position information of each corresponding row in the first intermediate index matrix to the multiple matrix registers respectively.
[0013] The sparse matrix accumulation subunit accumulates multiple first intermediate result matrices based on multiple first intermediate index matrices to obtain the result matrix of the matrix multiplication operation of the first sparse matrix and the second sparse matrix, including: for the first intermediate result matrix and the corresponding first intermediate index matrix after the first one output by the sparse matrix accumulation subunit, the sparse matrix accumulation subunit accumulates the intermediate results of each row in the first intermediate result matrix and the intermediate results stored in multiple matrix registers respectively based on the row position information in the first intermediate index matrix and the position information stored in multiple matrix registers, and stores the accumulated intermediate results and the corresponding position information back to multiple matrix registers respectively. The sparse matrix accumulation subunit determines the result matrix based on the intermediate results and position information stored in multiple matrix registers. In the embodiments of the present application, multiple matrix registers cooperate with the sparse matrix accumulation subunit to implement highly parallel accumulation processing, thereby improving the efficiency of matrix operations.
[0014] In an implementable manner, the above method further includes: when the storage space of the matrix register reaches the storage space threshold, the sparse matrix processing unit stores the intermediate results and the corresponding position information stored in the matrix register in the form of a matrix into the memory to obtain a second intermediate result matrix and a second intermediate index matrix. The sparse matrix processing unit groups the second intermediate result matrix according to the number of matrix registers to obtain multiple groups of second intermediate result matrices.
[0015] The sparse matrix accumulation subunit determines the result matrix based on the intermediate results and position information stored in multiple matrix registers, including: the sparse matrix accumulation subunit accumulates the second intermediate result matrix based on the second intermediate index matrix corresponding to each group of second intermediate result matrices to obtain the result matrix. In the embodiments of the present application, by grouping and calculating the second intermediate result matrix, the complexity of accumulating the second intermediate result matrix can be reduced, thereby improving the efficiency of matrix operations.
[0016] In a second aspect, a processor is provided, the processor includes a sparse matrix processing unit and a sparse matrix operation unit, and the method includes:
[0017] The sparse matrix processing unit is configured to obtain a first numerical matrix and a first index matrix corresponding to a first sparse matrix, and a second numerical matrix and a second index matrix corresponding to a second sparse matrix, where the first numerical matrix is obtained by rearranging the non-zero elements in each column of the first sparse matrix, the first index matrix is used to indicate the row numbers of the non-zero elements in the first numerical matrix in the first sparse matrix, the second numerical matrix is obtained by rearranging the non-zero elements in each row of the second sparse matrix, and the second index matrix is used to indicate the column numbers of the non-zero elements in the second numerical matrix in the second sparse matrix.
[0018] A sparse matrix processing unit for inputting the non-zero elements in each column of the first numerical matrix, the row numbers indicated by each column of the first index matrix, the non-zero elements in each row of the second numerical matrix, and the column numbers indicated by each row of the second index matrix into a sparse matrix operation unit.
[0019] A sparse matrix operation unit for performing an outer product operation on the non-zero elements in each column of the first numerical matrix and the non-zero elements in each row of the second numerical matrix based on the row numbers indicated by each column of the first index matrix and the column numbers indicated by each row of the second index matrix, to obtain a result matrix of performing matrix multiplication on the first sparse matrix and the second sparse matrix.
[0020] In an implementable manner, the first numerical matrix includes a plurality of first numerical partitions, which are obtained by dividing the first sparse matrix into a plurality of first sparse matrix regions by rows, and then rearranging the non-zero elements in each column of each first sparse matrix region. The first index matrix includes first index partitions corresponding to each first numerical partition.
[0021] The second numerical matrix includes a plurality of second numerical partitions, which are obtained by dividing the second sparse matrix into a plurality of second sparse matrix regions by columns, and then rearranging the non-zero elements in each row of each second sparse matrix region. The second index matrix includes second index partitions corresponding to each second numerical partition.
[0022] In an implementable manner, each first numerical partition includes a plurality of first numerical blocks of size n×m, each first index partition includes a plurality of first index blocks of size n×m, each second numerical partition includes a plurality of second numerical blocks of size m×n, and each second index partition includes a plurality of second index blocks of size m×n.
[0023] In an implementable manner, the sparse matrix processing unit is used to sequentially input, according to the calculation order, the elements in each column of the non-zero numerical blocks in each column of the first numerical blocks in each first numerical partition, the elements in each row of the non-zero numerical blocks in each row of the second numerical blocks in each second numerical partition, the row numbers indicated by each column of the non-zero index blocks in each column of the first index blocks in each first index partition, and the column numbers indicated by each column of the non-zero index blocks in each row of the second index blocks in each second index partition into the sparse matrix operation unit.
[0024] In an implementable manner, the sparse matrix operation unit includes a sparse matrix outer product subunit and a sparse matrix accumulation subunit.
[0025] The sparse matrix outer product subunit is used to perform an outer product operation on the elements in each column of the first numerical block and the elements in each row of the second numerical block input each time, to obtain a first intermediate result matrix.
[0026] A sparse matrix outer product subunit, configured to form a first intermediate index matrix corresponding to a first intermediate result matrix by using the row numbers indicated by each column of the first index block input each time and the column numbers indicated by each column of the second index block, where the first intermediate index matrix is used to indicate the position information of each element in the first intermediate result matrix in the result matrix.
[0027] A sparse matrix accumulation subunit, configured to accumulate multiple first intermediate result matrices based on multiple first intermediate index matrices to obtain a result matrix of the matrix multiplication operation performed by the first sparse matrix and the second sparse matrix.
[0028] In an implementable manner, the processor includes multiple matrix registers. For the first first intermediate result matrix output by the sparse matrix outer product subunit and the corresponding first intermediate index matrix, a sparse matrix processing unit is configured to store the intermediate results of each row in the first intermediate result matrix and the position information of each row corresponding to the first intermediate index matrix into multiple matrix registers respectively.
[0029] For the first intermediate result matrix and the corresponding first intermediate index matrix after the first one output by the sparse matrix accumulation subunit, the sparse matrix accumulation subunit is configured to accumulate the intermediate results of each row in the first intermediate result matrix and the intermediate results stored in multiple matrix registers respectively based on the position information of each row in the first intermediate index matrix and the position information stored in multiple matrix registers respectively, and store the accumulated intermediate results and the corresponding position information back into multiple matrix registers respectively.
[0030] A sparse matrix accumulation subunit, configured to determine a result matrix based on the intermediate results and position information stored in multiple matrix registers.
[0031] In an implementable manner, when the storage space of the matrix register reaches the storage space threshold, a sparse matrix processing unit is configured to store the intermediate results and the corresponding position information stored in the matrix register into the memory in the form of a matrix to obtain a second intermediate result matrix and a second intermediate index matrix.
[0032] A sparse matrix processing unit, configured to group the second intermediate result matrix according to the number of matrix registers to obtain multiple groups of second intermediate result matrices.
[0033] A sparse matrix accumulation subunit, configured to accumulate the second intermediate result matrix based on the second intermediate index matrix corresponding to each group of second intermediate result matrices to obtain a result matrix.
[0034] In a third aspect, a computing device is provided, which includes a memory and a processor as described in the second aspect above. At least one instruction is stored in the memory, and when the processor executes the at least one instruction, the method described in the first aspect above can be implemented.
[0035] In a fourth aspect, a computer-readable storage medium is provided, which stores computer program code. When the computer program code is executed by a computing device, the computing device is caused to execute the method described in the first aspect above.
[0036] In a fifth aspect, a computer program product containing instructions is provided. When the computer program product runs on a computing device, the computing device is caused to execute the method described in the first aspect above. Description of the Drawings
[0037] Figure 1 is a schematic diagram of performing matrix multiplication operation in the related art;
[0038] Figure 2 is a schematic structural diagram of a computing device provided in an embodiment of the present application;
[0039] Figure 3 is a schematic diagram of a sparse compression matrix provided in an embodiment of the present application;
[0040] Figure 4 is a flowchart of a method for operating a sparse matrix provided in an embodiment of the present application;
[0041] Figure 5 is a schematic structural diagram of a processor provided in an embodiment of the present application;
[0042] Figure 6 is a schematic diagram of pre-partitioning a sparse matrix provided in an embodiment of the present application;
[0043] Figure 7 is a schematic diagram of a sparse matrix outer product sub-unit performing an outer product operation provided in an embodiment of the present application;
[0044] Figure 8 is a schematic diagram of a method for a sparse matrix accumulation sub-unit to perform an accumulation operation provided in an embodiment of the present application;
[0045] Figure 9 is a schematic flowchart of a sparse matrix accumulation sub-unit performing an accumulation process provided in an embodiment of the present application;
[0046] Figure 10 is a schematic diagram of a method for grouping a second intermediate result provided in an embodiment of the present application. Detailed Embodiments
[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0048] Matrix multiplication is widely used in various scientific computing scenarios. Among them, the density of matrices involved in different computing scenarios varies. Matrix multiplication is mainly divided into dense matrix multiplication and sparse matrix multiplication. Generally, dense matrix multiplication is mainly applied in scenarios such as traditional machine learning and neural network model calculations, while sparse matrix multiplication is mainly applied in scenarios such as graph computing, graph neural networks, and multigrid solvers.
[0049] Dense matrix multiplication means that the matrices for matrix multiplication operations are dense matrices, and sparse matrix multiplication means that the matrices for matrix multiplication operations are sparse matrices. Among them, a sparse matrix refers to a matrix in which the percentage of non-zero elements in all elements is very small (for example, less than 5%), and a dense matrix refers to a matrix in which the percentage of non-zero elements in all elements is relatively large.
[0050] Currently, the optimization methods for matrix multiplication operations are only effective for dense matrix multiplication. For example, the matrices for matrix multiplication operations can be tiled, and then the result matrices of the matrix multiplication operations can be calculated through the tiled matrices after tiling. Figure 1 is a schematic diagram of performing matrix multiplication operations in the related art. As Figure 1 shown, for matrices A and B for performing matrix multiplication operations, matrix A can be divided into tiled matrix A 0,0 , tiled matrix A 0,1 , tiled matrix A 1,0 , tiled matrix A 1,1 , and matrix B can be divided into tiled matrix B 0,0 , tiled matrix B 0,1 , tiled matrix B 1,0 , tiled matrix B 1,1 . Among them, for the result matrix C of performing matrix multiplication operations on matrices A and B, matrix C can be divided into tiled matrix C 0,0 , tiled matrix C 0,1 , tiled matrix C 1,0 , tiled matrix C 1,1 . Taking the calculation of tiled matrix C 0,0 as an example, tiled matrix C 0,0 is equal to the product of tiled matrix A 0,0 and tiled matrix B 0,0 plus the product of tiled matrix A 0,1 and tiled matrix B0,1 The sum of products. In this way, by reducing the size of the matrix, the complexity of matrix multiplication operations can be reduced, thereby improving the efficiency of matrix multiplication operations.
[0051] However, for sparse matrices, due to the large number of zero elements in sparse matrices, even after tiling the sparse matrix, there are still a large number of zero elements in the resulting tiled matrix, and there may even be a tiled matrix that is a zero matrix. In this way, during the process of matrix multiplication operations, a large number of zero elements are always involved in the operations, occupying a large amount of computing resources and storage resources, resulting in low execution efficiency of sparse matrix multiplication.
[0052] The embodiments of the present application provide a method for operating sparse matrices, which can optimize sparse matrix multiplication in combination with the characteristics of sparse matrices, and can further improve the execution efficiency of sparse matrix multiplication. Figure 2 is a computing device for implementing the method for operating sparse matrices provided by the embodiments of the present application. As Figure 2 shown, the computing device 200 may include: a bus 202, a processor 204, a memory 206. Optionally, the computing device 200 may further include a communication interface 208. The processor 204, the memory 206, and the communication interface 208 communicate with each other through the bus 202. The computing device 200 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 200. The computing device 200 may be a device for running a model, which may be a terminal or a server. When the computing device 200 is a terminal, the computing device 200 includes, but is not limited to, a desktop computer, a mobile phone, a notebook, a tablet computer, etc. When the computing device 200 is a server, the computing device 200 may be a single server or a server cluster composed of multiple servers, and may be a physical machine or a virtual machine or a container virtualized through virtual technology.
[0053] The bus 202 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 2 only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The bus 202 may include a path for transmitting information between various components of the computing device 200 (for example, the memory 206, the processor 204, the communication interface 208).
[0054] The processor 204 may include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP). Alternatively, the processor may be a system-on-chip (SOC) including one or more of the above-mentioned CPU, GPU, MP, etc. Among them, the processor 204 may further include a sparse matrix processing unit and a sparse matrix operation unit, etc. Among them, the sparse matrix processing unit may be a processor core, and the sparse matrix operation unit may be used to implement the multiplication operation of the sparse matrix provided in the embodiments of the present application.
[0055] The memory 206 may include volatile memory, such as random access memory (RAM). The memory 206 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD). Executable program code is stored in the memory 206, and the processor 204 executes the executable program code to implement the sparse matrix operation method provided in the embodiments of the present application.
[0056] The communication interface 208 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 200 and other devices or communication networks.
[0057] To facilitate understanding of the sparse matrix operation method provided in the embodiments of the present application, the sparse compression matrix adopted in the embodiments of the present application will be introduced first:
[0058] The sparse compression matrix includes a numerical (Data) matrix and an index (Index) matrix. The numerical matrix is obtained by continuously arranging the zero elements and non-zero elements in the sparse matrix row by row or column by column.
[0059] In one example, after the non-zero elements and zero elements included in each row (column) of the sparse matrix are continuously rearranged respectively, a numerical matrix corresponding to the sparse matrix can be obtained. The index matrix and the numerical matrix have the same size, and the distribution of zero elements and non-zero elements in the index matrix is consistent with the distribution of zero elements and non-zero elements in the numerical matrix. When the zero elements and non-zero elements in the sparse matrix are continuously rearranged by row to obtain the corresponding numerical matrix, each non-zero element in the index matrix is used to indicate the column number of the non-zero element at the same position in the numerical matrix in the sparse matrix. When the zero elements and non-zero elements in the sparse matrix are continuously rearranged by row and column to obtain the corresponding numerical matrix, each non-zero element in the index matrix is used to indicate the row number of the non-zero element at the same position in the numerical matrix in the sparse matrix. Among them, continuously rearranging the non-zero elements and zero elements respectively can be arranging the non-zero elements before the zero elements or arranging the non-zero elements after the zero elements.
[0060] As Figure 3 described, for the sparse matrix S, if the non-zero elements and zero elements are continuously rearranged by row, the corresponding numerical matrix A 1 , and the corresponding index matrix A 2 can be obtained. If the non-zero elements and zero elements are continuously rearranged by column, the corresponding numerical matrix B 1 , and the corresponding index matrix B 2 can be obtained. For example, the value at the 3rd row and 1st column of the index matrix A 2 is used to indicate that the element at the 3rd row and 1st column in the numerical matrix A 1 is in the 2nd column of the sparse matrix S. Another example, the value at the 3rd row and 3rd column of the index matrix B 2 is used to indicate that the element at the 3rd row and 3rd column in the numerical matrix B 1 is in the 4th row of the sparse matrix S. It should be noted that in the embodiments of the present application, the row numbers and column numbers of the matrix start from 1.
[0061] The embodiments of the present application provide a method for operating a sparse matrix. This operation method can, based on the above sparse compression matrix, perform the multiplication operation of the sparse matrix through the outer product of the matrix, which can improve the efficiency of the multiplication operation of the sparse matrix. Figure 4 is a schematic flowchart of a method for operating a sparse matrix provided by the embodiments of the present application. This method can be executed by the computing device 200 shown in the above Figure 2 , and specifically can be executed by the processor 204. As Figure 5As shown, the processor 204 may at least include a sparse matrix processing unit and a sparse matrix operation unit. Among them, the sparse matrix processing unit may refer to the hardware part in the processor 204 for processing the sparse matrix for performing sparse multiplication operations, such as a processor core. Or the sparse matrix processing unit may also refer to the software program running on the processor for processing the sparse matrix. The sparse matrix operation unit is the hardware unit in the processor 204 for performing operations on the sparse matrix in this application.
[0062] See Figure 4 , the operation method of the sparse matrix provided by the embodiment of this application includes:
[0063] Step 401, the sparse matrix processing unit obtains the first numerical matrix and the first index matrix corresponding to the first sparse matrix, and the second numerical matrix and the second index matrix corresponding to the second sparse matrix.
[0064] Among them, the first numerical matrix is obtained by rearranging the non-zero elements in each column of the first sparse matrix, and the first index matrix is used to indicate the row numbers of the non-zero elements in the first numerical matrix in the first sparse matrix. The second numerical matrix is obtained by rearranging the non-zero elements in each row of the second sparse matrix, and the second index matrix is used to indicate the column numbers of the non-zero elements in the second numerical matrix in the second sparse matrix.
[0065] The first sparse matrix and the second sparse matrix are two sparse matrices to be subjected to matrix multiplication operations. In one example, the first sparse matrix is the left matrix to be subjected to matrix multiplication operations, and the second sparse matrix is the right matrix to be subjected to matrix multiplication operations. The first numerical matrix corresponding to the first sparse matrix is a matrix obtained by continuously rearranging the zero elements and non-zero elements in each column of the first sparse matrix. For example, the non-zero elements can be arranged in the order of row numbers, and the remaining zero elements can be arranged behind the non-zero elements. The second numerical matrix corresponding to the second sparse matrix is a matrix obtained by continuously rearranging the zero elements and non-zero elements in each row of the second sparse matrix. For example, the non-zero elements can be arranged in the order of column numbers, and the remaining zero elements can be arranged behind the non-zero elements.
[0066] In implementation, an application program involving sparse matrix multiplication can run on a computing device. For example, the application program can be a high-performance computing (HPC) application, an artificial intelligence (AI) application, a graph computing application, etc. During the running of the application program, when sparse matrix multiplication is required, a matrix multiplication operation request can be sent to the processor. In response to the matrix multiplication operation request, the sparse matrix processing unit in the processor can read the first numerical matrix and the first index matrix corresponding to the first sparse matrix to be subjected to matrix multiplication operation, and the second numerical matrix and the second index matrix corresponding to the second sparse matrix from the memory or storage of the computing device (such as a hard disk).
[0067] If the first numerical matrix and the first index matrix corresponding to the first sparse matrix, and the second numerical matrix and the second index matrix corresponding to the second sparse matrix are not stored in the memory or storage, the sparse matrix processing unit can read the first sparse matrix and the second sparse matrix. Among them, the first sparse matrix and the second sparse matrix can be, but are not limited to, formats such as CSC (Compressed Sparse Column Format), CSR (Compressed Sparse Row Format), COO (Coordinate Format), etc. After the sparse matrix processing unit reads the first sparse matrix and the second sparse matrix, it can generate the first numerical matrix and the first index matrix corresponding to the first sparse matrix, and the second numerical matrix and the second index matrix corresponding to the second sparse matrix in the manner shown above Figure 3 and shown.
[0068] Step 402: The sparse matrix processing unit inputs the non-zero elements in each column of the first numerical matrix and the row numbers indicated by each column in the first index matrix, the non-zero elements in each row of the second numerical matrix, and the column numbers indicated by each row in the second index matrix into the sparse matrix operation unit.
[0069] The sparse matrix processing unit can input the row numbers indicated by each column of the first index matrix corresponding to the non-zero elements in each column of the first numerical matrix and the column numbers indicated by each row of the second index matrix corresponding to the non-zero elements in each row of the second numerical matrix into the sparse matrix operation unit in the order of performing the matrix outer product on the first numerical matrix and the second numerical matrix. That is, the non-zero elements of the first numerical matrix and the first index matrix in the first column are input into the sparse matrix operation unit in turn with the non-zero elements of the second numerical matrix and the second index matrix in the first row, the non-zero elements of the first numerical matrix and the first index matrix in the second column are input into the sparse matrix operation unit with the non-zero elements of the second numerical matrix and the second index matrix in the second row, …, and the non-zero elements of the first numerical matrix and the first index matrix in the nth column are input into the sparse matrix operation unit with the non-zero elements of the second numerical matrix and the second index matrix in the nth row.
[0070] In one example, if there are all zero elements in a certain column of the first numerical matrix or all zero elements in a certain row of the second numerical matrix, then it is possible to skip inputting the column with all zero elements in the first numerical matrix (and the first index matrix) and the corresponding row in the corresponding second numerical matrix (and the second index matrix) into the sparse matrix operation unit. Or, it is possible to skip inputting the row with all zero elements in the second numerical matrix (and the second index matrix) and the corresponding column in the corresponding first numerical matrix (and the first index matrix) into the sparse matrix operation unit. For example, if the i-th column in the first numerical matrix is all zero elements, then the i-th column of the first numerical matrix (and the first index matrix) and the i-th row of the second numerical matrix (and the second index matrix) can be not input into the sparse matrix operation unit. In implementation, the column numbers of the columns with all zero elements in the first numerical matrix or the row numbers of the rows with zero elements in the second numerical matrix can be marked by specified symbols. By skipping the columns with all zero elements and the rows with all zero elements input into the sparse matrix operation unit, the calculation amount of the sparse matrix operation unit can be reduced, and thus the calculation efficiency of the sparse matrix operation unit can be improved.
[0071] Step 403: The sparse matrix operation unit performs an outer product operation on the non-zero elements in each column of the first numerical matrix and the non-zero elements in each row of the second numerical matrix based on the row numbers indicated by each column of the first index matrix and the column numbers indicated by each row of the second index matrix, to obtain the result matrix of the matrix multiplication operation of the first sparse matrix and the second sparse matrix.
[0072] For each input of the non-zero elements in a column of the first numerical matrix and the non-zero elements in a row of the second numerical matrix, the sparse matrix operation unit can perform an outer product operation on the non-zero elements in this column and the non-zero elements in this row to obtain an intermediate result matrix. That is, the sparse matrix operation unit can multiply the i-th non-zero element in this column of non-zero elements by the j-th non-zero element in this row of non-zero elements to obtain the element value corresponding to the element in the i-th row and the j-th column in the intermediate result matrix during the outer product operation (subsequently referred to as the intermediate result).
[0073] For each column of non-zero elements in the first index matrix and each row of non-zero elements in the second index matrix in the input, the sparse matrix operation unit can, in the order of the outer product operation, form the intermediate index matrix corresponding to the intermediate result matrix with the non-zero elements of the column and the non-zero elements of the row. Each element in the intermediate index matrix is the position information of the intermediate result at the same position in the intermediate result matrix in the result matrix. For example, if the i-th non-zero element of a column of non-zero elements in the first index matrix is "2", and the j-th non-zero element of a row of non-zero elements in the second index matrix is "3", the formed position information is "2, 3", that is, the element value at the i-th row and j-th column in the corresponding intermediate index matrix is "2, 3". This element value is used to indicate the position of the intermediate result at the i-th row and j-th column in the intermediate result matrix in the result matrix, that is, the position of the intermediate result at the i-th row and j-th column in the result matrix is the 2nd row and 3rd column.
[0074] For each obtained intermediate result matrix, the sparse matrix operation unit can accumulate the elements in the intermediate result matrix according to the intermediate index matrix corresponding to each intermediate result matrix, that is, accumulate the elements with the same position information in the intermediate result matrix. After the accumulation, the accumulation result and the remaining intermediate results in each intermediate result matrix that have not been accumulated are the non-zero elements in the result matrix, and the position information corresponding to the accumulation result and the position information corresponding to the remaining intermediate results in the intermediate result matrix that have not been accumulated are the position information of the non-zero elements in the result matrix in the result matrix. In this way, the result matrix of the matrix multiplication of the first sparse matrix and the second sparse matrix can be obtained according to the accumulation result, the unaccumulated intermediate results, and the corresponding position information. After obtaining the result matrix corresponding to the first sparse matrix and the second sparse matrix, the sparse matrix processing unit can return the result to the application program that requests the sparse matrix multiplication.
[0075] In the embodiment of the present application, the first numerical matrix and the first index matrix corresponding to the first sparse matrix to be matrix-operated are obtained, and the second numerical matrix and the second index matrix corresponding to the second sparse matrix are obtained. According to the row numbers of the non-zero elements in the first numerical matrix indicated by the first index matrix in the first sparse matrix and the column numbers of the non-zero elements in the second numerical matrix indicated by the second index matrix in the second sparse matrix, the outer product operation is performed on the non-zero elements in the first sparse matrix and the second sparse matrix, and then the result matrix of the matrix multiplication of the first sparse matrix and the second sparse matrix is obtained. In this way, a large number of zero elements in the first sparse matrix and the second sparse matrix can be avoided from participating in the matrix multiplication operation process, and thus the operation efficiency of performing the sparse matrix multiplication operation can be improved.
[0076] The embodiments of the present application also provide a method for partitioning the above-mentioned first numerical matrix, first index matrix, second numerical matrix, and second numerical matrix, as follows:
[0077] The above-mentioned first numerical matrix includes a plurality of first numerical partitions. The plurality of first numerical partitions are obtained by dividing the first sparse matrix into a plurality of first sparse matrix regions row by row, and then rearranging each non-zero element in each column of each first sparse matrix region. The first index matrix includes a first index partition corresponding to each first numerical partition. The above-mentioned second numerical matrix includes a plurality of second numerical partitions. The plurality of second numerical partitions are obtained by dividing the second sparse matrix into a plurality of second sparse matrix regions column by column, and then rearranging each non-zero element in each row of each second sparse matrix region. The second index matrix includes a second index partition corresponding to each second numerical partition.
[0078] In implementation, pre-partition parameters for pre-partitioning the first sparse matrix can be set. Among them, the pre-partition parameters can be the number of rows for pre-partitioning the first sparse matrix or the number of columns for pre-partitioning the first sparse matrix (the number of rows is equal to the number of columns), or the number of partitions for partitioning the first sparse matrix and the second sparse matrix.
[0079] When the pre-partition parameters are the number of rows for pre-partitioning the first sparse matrix and the number of columns for pre-partitioning the first sparse matrix, the first sparse matrix can be partitioned row by row starting from the first row of the first sparse matrix according to the number of rows, to obtain a plurality of first sparse matrix regions corresponding to the first sparse matrix, that is, the number of rows included in each first sparse matrix region is equal to the number of rows. Among them, the number of rows in the last first sparse matrix region in the first sparse matrix can be equal to or less than the number of rows. Similarly, the second sparse matrix can be partitioned column by column starting from the first column of the second sparse matrix according to the number of columns, to obtain a plurality of second sparse matrix regions corresponding to the second sparse matrix, that is, the number of columns included in each second sparse matrix region is equal to the number of columns. Among them, the number of columns in the last second sparse matrix region in the second sparse matrix can be equal to or less than the number of columns.
[0080] When the pre-partition parameter is the number of partitions for partitioning the first sparse matrix and the second sparse matrix, the first sparse matrix can be evenly divided into a plurality of first sparse matrix regions in the column direction according to the number, and the second sparse matrix can be evenly divided into a plurality of second sparse matrix regions in the row direction according to the number.
[0081] Figure 6 is a schematic diagram of pre-partitioning a sparse matrix provided by the embodiments of the present application. In Figure 6Among them, the number of rows of the left matrix and the number of columns of the right matrix are both x, and the number of rows or columns corresponding to the pre-partitioning parameter is y, where x = 3y. Correspondingly, the left matrix can be evenly divided into 3 first sparse matrix regions in the column direction, with each first sparse matrix region including y rows, and the right matrix can be evenly divided into 3 second sparse matrix regions in the row direction, with each second sparse matrix region including y columns.
[0082] After dividing the first sparse matrix into multiple first sparse matrix regions, the zero elements in each column of each first sparse matrix region can be continuously rearranged to obtain each first numerical partition in the first numerical matrix. After dividing the second sparse matrix into multiple second sparse matrix regions, the zero elements in each row of each second sparse matrix region can be continuously rearranged to obtain each second numerical partition in the second numerical matrix. Correspondingly, a first index partition corresponding to each first numerical partition in the first numerical matrix and a second index partition corresponding to each second numerical partition in the second numerical matrix can be generated. Among them, the non-zero elements in the first index partition are used to indicate the row numbers of the non-zero elements at the same position in the corresponding first numerical partition in the first sparse matrix, and the non-zero elements in the second index partition are used to indicate the column numbers of the non-zero elements at the same position in the corresponding second numerical partition in the second sparse matrix.
[0083] In the embodiments of the present application, the first sparse matrix can be divided into multiple first sparse matrix regions, and the second sparse matrix can be divided into multiple second sparse matrix regions, and then the outer product operation is performed on each first sparse matrix region and multiple second sparse matrix regions respectively to implement the matrix multiplication operation of the first sparse matrix and the second sparse matrix. In this way, the matrix size in the outer product operation process can be reduced, and the efficiency of multiplying sparse matrices can be further improved.
[0084] In the embodiments of the present application, each first numerical partition can be further divided into multiple first numerical sub-blocks with a size of n×m, each first index partition can be further divided into multiple first index sub-blocks with a size of n×m, each second numerical partition can be further divided into multiple second numerical sub-blocks with a size of m×n, and each second index partition can be further divided into multiple second index sub-blocks with a size of m×n.
[0085] Among them, n×m and m×n are pre-set Tiling parameters for Tiling the first numerical partition and the second numerical partition. The first numerical block, the first index block, the second numerical block, and the second index block obtained after Tiling can be collectively referred to as Tiling blocks. In this application, the first numerical partition and the first index partition corresponding to the first sparse matrix, the Tiling blocks included in the first numerical partition and the first index partition, and the second numerical partition and the second index partition corresponding to the second sparse matrix, and the Tiling blocks included in the second numerical partition and the second index partition can be pre-partitioned, that is, the sparse matrix processing unit can directly obtain them from the memory or storage. Alternatively, it can also be partitioned by the sparse matrix processing unit after obtaining the first sparse matrix and the second sparse matrix (or the first numerical partition, the first index partition, the second numerical partition, and the second index partition).
[0086] In the embodiments of this application, by further dividing the first numerical partition, the first index partition, the second numerical partition, and the second index partition into Tiling blocks, and then performing outer product operations on the Tiling blocks respectively, the matrix size in the outer product operation process can be further reduced, and the efficiency of multiplying sparse matrices can be further improved.
[0087] The above pre-partitioning and Tiling partitioning of the first sparse matrix and the second sparse matrix can be performed before step 402. The following introduces the process of performing outer product operations on the Tiling blocks in this application to implement the multiplication operation of the first sparse matrix and the second sparse matrix:
[0088] In an implementable manner, the processing of the above step 402 may include: the sparse matrix processing unit inputs, according to the calculation order, each column element in the non-zero numerical blocks in each column of the first numerical blocks in each first numerical partition, each row element in the non-zero numerical blocks in each row of the second numerical blocks in each second numerical partition, the row numbers indicated by each column of the non-zero index blocks in each column of the first index blocks in each first index partition, and the column numbers indicated by each column of the non-zero index blocks in each row of the second index blocks in each second index partition, into the sparse matrix operation unit.
[0089] In practice, each first numerical partition in the first numerical matrix can be respectively operated with the second numerical partition in each second numerical matrix. For example Figure 6As shown, the first numerical partition a can first perform matrix operations with the second numerical partition a, the second numerical partition b, and the second numerical partition c respectively. Then, the first numerical partition b can perform matrix operations with the second numerical partition a, the second numerical partition b, and the second numerical partition c respectively. Finally, the first numerical partition c can perform matrix operations with the second numerical partition a, the second numerical partition b, and the second numerical partition c respectively. Among them, when the first numerical partition and the second numerical partition perform operations, it is only the operation between the non-zero matrix blocks included in the first numerical partition and the second numerical partition. Since there will be a large number of zero matrix blocks, that is, Tiling blocks all of which are zero, after pre-partitioning and Tiling of the first sparse matrix and the second sparse matrix, only performing operations on the non-zero matrix blocks in the first numerical partition and the second numerical partition can greatly reduce the Tiling blocks participating in the operation. Therefore, the efficiency of matrix operations by the subsequent sparse matrix operation unit can be improved.
[0090] For any first numerical partition and second numerical partition performing operations, each first numerical sub-block belonging to a non-zero matrix block in each Tiling column of the first numerical partition can perform a matrix outer product operation with each second numerical sub-block belonging to a non-zero block in the corresponding Tiling row of the second numerical partition. For example, each first numerical sub-block belonging to a non-zero block in the i-th Tiling column of the first numerical partition can perform a matrix outer product operation with each second numerical sub-block belonging to a non-zero block in the i-th Tiling row of the second numerical partition. Among them, the Tiling column refers to the column composed of Tiling blocks in the first numerical partition, the Tiling row refers to the row composed of Tiling blocks in the second numerical partition, and the Tiling row corresponding to the Tiling column means that the column number of the Tiling column is the same as the row number of the Tiling row.
[0091] As Figure 6 shown, the three non-zero Tiling blocks in the first Tiling column of the first numerical partition a can perform outer product operations with the two non-zero Tiling blocks in the first Tiling row of the second numerical partition a respectively. The two non-zero Tiling blocks in the second Tiling column of the first numerical partition a can perform outer product operations with the two non-zero Tiling blocks in the second Tiling row of the second numerical partition a respectively.
[0092] A technician can preset the calculation order of the first numerical partition and the second numerical partition, as well as the calculation order of each column of the first numerical sub-block and the corresponding row of the second numerical sub-block in the first numerical partition and the second numerical partition. In implementation, the sparse matrix processing unit can input each column element in the non-zero numerical sub-block of each column of the first numerical sub-block in each first numerical partition, each row element in the non-zero numerical sub-block of each row of the second numerical sub-block in each second numerical partition, each column element in the non-zero index sub-block of each column of the first index partition, and each row element in the non-zero index sub-block of each row of the second index partition into the sparse matrix operation unit in sequence according to the set calculation order.
[0093] In an example, the sparse matrix operation unit provided by the embodiment of the present application includes a sparse matrix outer product sub-unit and a sparse matrix accumulation sub-unit. The processing in step 403 may include:
[0094] The sparse matrix outer product sub-unit performs an outer product operation on each column element of the first numerical sub-block and each row element of the second numerical sub-block input each time to obtain a first intermediate result matrix. The row numbers indicated by each column of the first index sub-block input each time and the column numbers indicated by each column of the second index sub-block are combined to form a first intermediate index matrix corresponding to the first intermediate result matrix. The first intermediate index matrix is used to indicate the position information of each element in the first intermediate result matrix in the result matrix.
[0095] Figure 7 It is a schematic diagram of a method for a sparse matrix outer product sub-unit provided by the embodiment of the present application to perform an outer product operation. As Figure 7 shown, the inputs of the sparse matrix outer product sub-unit are the first numerical sub-block (Value1) and the corresponding first index sub-block (Index1), the second numerical sub-block (Value2), and the corresponding second index sub-block (Index2). The outputs of the sparse matrix outer product sub-unit are multiple groups of first intermediate result matrices and corresponding first intermediate index matrices. Among them, the number of columns of the first intermediate result matrix (the first intermediate index matrix) output by the sparse matrix outer product sub-unit is equal to the number of columns of the input first numerical sub-block (or the number of rows of the second numerical sub-block).
[0096] In implementation, a Tiling block row-column vector matching device can be run in the sparse matrix processing unit. The Tiling block row-column vector matching device can input a column element in the first numerical sub-block and a corresponding row element in the second numerical sub-block, a column element in the first index sub-block and a corresponding row element in the second index sub-block into the sparse matrix outer product sub-unit, so that the sparse matrix outer product sub-unit outputs a first intermediate index matrix and the corresponding first intermediate index matrix. For example, Value1 is a 4×2 identity matrix, and Value2 is a 2×4 identity matrix. Two first intermediate index matrices can be generated by the sparse matrix outer product subunit, namely and and two 2×4 identity matrices can be generated.
[0097] In one example, a technician can preset a sparse outer product (SpSMEFMOPA) instruction. During the sparse matrix multiplication operation, the SpSMEFMOPA instruction can be called to implement the outer product operation of the sparse matrix outer product subunit on the input first numerical block and second numerical block. In addition, the technician can set the specific circuit structure of the sparse matrix outer product subunit according to the method for performing the outer product operation provided by the embodiments of the present application. The specific circuit structure of the sparse matrix outer product subunit will not be described in detail in the embodiments of the present application.
[0098] The sparse matrix accumulation subunit accumulates multiple first intermediate result matrices based on multiple first intermediate index matrices to obtain the result matrix of the matrix multiplication operation of the first sparse matrix and the second sparse matrix.
[0099] Figure 8 is a schematic diagram of a method for performing an accumulation operation by a sparse matrix accumulation subunit provided by the embodiments of the present application. As Figure 8 shown, the inputs of the sparse matrix accumulation subunit are the first intermediate result matrix (Value3) and the corresponding first intermediate index matrix (Index2), the first intermediate result matrix (Value4) and the corresponding first intermediate index matrix (Index4). In the case where there is no identical position information in the input second intermediate index matrix, the sparse matrix accumulation subunit does not need to accumulate the second intermediate result matrix. The output second numerical index matrix and the corresponding second intermediate index matrix of the sparse matrix accumulation subunit are the input first numerical index matrix and first intermediate index matrix. In the case where there is identical position information in the input second intermediate index matrix, the sparse matrix accumulation subunit can accumulate the intermediate results corresponding to the two first intermediate result matrices at the identical position information respectively to obtain an accumulation result, and output the accumulation result and the corresponding position information, as well as the intermediate results that have not been accumulated and the corresponding position information. The accumulation result and the intermediate results that have not been accumulated are the elements included in the second intermediate result matrix, and the accumulation result and the intermediate results that have not been accumulated and the corresponding position information are the elements included in the second intermediate index matrix.
[0100] In one example, a technician can preset a Sparse Bitwise Summed Matrix Multiplication and Element-Wise Addition (SpSMEFADD) instruction. During the sparse matrix multiplication operation, by invoking this SpSMEFADD instruction, the sparse matrix accumulation subunit can perform the accumulation process on the first intermediate result matrix. Additionally, the technician can set the specific circuit structure of the sparse matrix accumulation subunit according to the method for performing the accumulation process provided by the embodiments of the present application. The specific circuit structure of the sparse matrix accumulation subunit will not be described in detail in the embodiments of the present application. Further, the number of the sparse matrix outer product subunit and the sparse matrix accumulation subunit included in the sparse matrix operation unit is not limited, and can be one respectively, or multiple respectively.
[0101] In one example, the processor further includes an on-chip multi-channel storage device, which may include multiple storage channels, and each storage channel may correspond to a matrix register. The matrix register can store the accumulation result, the intermediate result that has not been accumulated, and the corresponding position information stored by the sparse matrix accumulation subunit in the form of a matrix. Among them, a multi-channel data mapping device may also run in the sparse matrix processing unit. The multi-channel data mapping device can store the intermediate results included in the second intermediate result matrix output by the sparse matrix accumulation subunit and the position information corresponding to the second intermediate index matrix to the matrix registers on multiple storage channels according to the set storage strategy. For example, the multi-channel data mapping device can store the rows in the second intermediate value matrix and the second index value matrix to different matrix registers according to the number of matrix registers and the size of the second intermediate value matrix. Among them, the matrix register can store the intermediate results and the corresponding position information in the form of a matrix respectively.
[0102] The following further gives an exemplary introduction to the accumulation process of performing matrix multiplication on the first sparse matrix and the second sparse matrix by using the accumulation operation unit based on the sparse matrix accumulation subunit:
[0103] Step S1: The sparse matrix processing unit stores the intermediate results of each row in the first intermediate result matrix output by the sparse matrix outer product subunit and the position information of each row corresponding to the first intermediate index matrix to multiple matrix registers respectively.
[0104] Since there is no intermediate result corresponding to the same position information in the first intermediate result matrix, the first intermediate result matrix output by the sparse matrix outer product subunit does not need to be accumulated first. The intermediate results of each row in the first intermediate result matrix and the corresponding position information of each row in the first intermediate index matrix can be stored in multiple matrix registers respectively. In one example, the number of rows of the intermediate index matrix and the corresponding intermediate value matrix stored in each matrix register can be determined according to the number of matrix registers. For example, if the intermediate index matrix includes 4 rows and there are 4 matrix registers, one row of intermediate results in the first intermediate result matrix and one row of position information in the first intermediate index matrix can be stored in each matrix register.
[0105] Step S2: For the first subsequent first intermediate result matrix and the corresponding first intermediate index matrix output by the sparse matrix accumulation subunit, the sparse matrix accumulation subunit accumulates the intermediate results of each row in the first intermediate result matrix and the intermediate results stored in multiple matrix registers respectively based on the position information of each row in the first intermediate index matrix and the position information stored in multiple matrix registers respectively, and stores the accumulated intermediate results and the corresponding position information back to multiple matrix registers respectively.
[0106] After the first first intermediate result matrix and the corresponding first intermediate index matrix output by the sparse matrix accumulation subunit, the sparse matrix processing unit can input the intermediate results and the corresponding position information in the second and subsequent first intermediate result matrices output by the sparse matrix accumulation subunit into the sparse matrix accumulation subunit, and input the intermediate results and the corresponding position information stored in multiple matrix registers into the sparse matrix accumulation subunit.
[0107] The sparse matrix accumulation subunit can accumulate the intermediate results according to the input position information, that is, accumulate the intermediate results with the same position information to obtain the accumulated result. The position information corresponding to the accumulated result remains unchanged and is still the position information before accumulation. The sparse matrix accumulation subunit can output the accumulated result and the corresponding position information, and output the intermediate results that have not been accumulated and the corresponding position information. Among them, the accumulated result can also be called the intermediate result. The accumulated result and the intermediate results that have not been accumulated are the intermediate results included in the second intermediate result matrix, and the position information corresponding to the accumulated result and the intermediate results that have not been accumulated is the information corresponding to the second intermediate index matrix.
[0108] Among them, the size of the second intermediate result matrix can be the same as the size of the first intermediate result matrix. In one example, the sparse matrix processing unit can form the second intermediate result matrix by the accumulated result and the intermediate results that have not been accumulated output by the sparse matrix accumulation subunit in a row-first or column-first manner, and form the second intermediate index matrix by the corresponding position information.
[0109] Figure 9 It is a schematic flowchart of the accumulation process performed by a sparse matrix accumulation subunit provided in an embodiment of the present application. As Figure 9 shown, a channel data loading device, a channel data write-back device, and a channel data write-to-memory device can be run in the sparse matrix processing unit. Among them:
[0110] The channel data loading device can sequentially load the intermediate results and position information stored in the matrix registers on each storage channel according to the set loading order. Among them, the loaded intermediate results and position information can be the intermediate results in the first first intermediate result matrix output by the sparse matrix outer product subunit and the position information in the first intermediate index matrix, or can also be the intermediate results in the second intermediate result matrix output by the sparse matrix accumulation subunit and the position information in the second intermediate index matrix.
[0111] After the channel data loading device loads the intermediate results and the corresponding position information stored in a matrix register each time, it can input the loaded intermediate results and position information into the sparse matrix accumulation subunit. At the same time, the sparse matrix outer product subunit can input the intermediate results in the first intermediate result matrix output and the position information corresponding to the first intermediate index matrix into the sparse matrix accumulation subunit ( Figure 9 not shown in the figure). In an example, a row of intermediate results in the first intermediate result matrix output by the sparse matrix outer product subunit and the intermediate results and the corresponding position information stored in a matrix register loaded by the channel data loading device can be input into the sparse matrix accumulation subunit for accumulation processing. By calling SpSMEAdd to execute, the sparse matrix accumulation subunit accumulates the input intermediate result information according to the position information input each time, outputs the accumulated result and the corresponding position information, and outputs the intermediate results and the corresponding position information that have not been accumulated.
[0112] The channel data write-back device can write back the intermediate results in the second intermediate result matrix output by the sparse matrix accumulation subunit each time and the corresponding position information in the second intermediate index matrix to the matrix register, as Figure 9 shown in the matrix registers corresponding to storage channels 1 to 4 respectively. In an example, in one accumulation process of the sparse matrix accumulation subunit, the matrix register loaded by the channel data loading device and the matrix register written by the channel data write-back device are the same matrix register. When the channel data write-back device writes the intermediate results and the corresponding position information to the matrix register, it can perform overwrite writing on the matrix register.
[0113] Step S3: When the storage space of the matrix register reaches the storage space threshold, the sparse matrix processing unit stores the intermediate result and the corresponding position information stored in the matrix register in the memory in the form of a matrix, obtaining a second intermediate result matrix and a second intermediate index matrix.
[0114] During the execution of Step S2, when the storage space of any matrix register reaches the storage space threshold, the matrix register, the stored intermediate result, and the corresponding position information can be stored in the memory in the form of a matrix. Among them, the matrix storing the intermediate result is the second intermediate result matrix, and the matrix storing the position information is the second intermediate index matrix. The storage space of the matrix register reaching the storage space threshold can mean that the storage space of the matrix register is full.
[0115] Step S4: The sparse matrix processing unit groups the second intermediate result matrix according to the number of matrix registers, obtaining multiple groups of second intermediate result matrices.
[0116] After Steps S2 and S3, the memory will store multiple second intermediate result matrices and the corresponding second intermediate index matrix for each second intermediate result matrix. Since there may still be intermediate results corresponding to the same position information in the second intermediate result matrix. Therefore, the sparse matrix accumulation subunit needs to further perform accumulation processing on the multiple second intermediate result matrices in the memory. To improve the efficiency of the accumulation processing on the second intermediate result matrix and reduce the number of accumulations on the second intermediate result matrix, a branch merge scheduling device can also be run in the sparse matrix processing unit. As Figure 10 shown, Figure 10 is a schematic diagram of a method for grouping the second intermediate result provided by an embodiment of the present application. The branch merge scheduling device can group the multiple second intermediate results stored in the memory, and then the sparse matrix accumulation subunit can perform accumulation processing on the second intermediate result matrices included in each group, thereby reducing the number of accumulations on the second intermediate result matrix. Among them, each group of second intermediate result matrices also includes the second intermediate index matrix corresponding to the second intermediate result matrix.
[0117] In one example, the second intermediate result matrix can be grouped according to the number of matrix registers. For example, the number of groups corresponding to the second intermediate result matrix is equal to the number of matrix registers. The accumulation results corresponding to the multiple groups of second intermediate result matrices can be stored in different matrix registers respectively, thereby improving the utilization rate of the matrix registers. As Figure 10 shown, the memory can include 8 intermediate result matrices, and the processor includes 4 matrix registers. Then the 8 intermediate result matrices can be divided into 4 groups, and each group includes 2 intermediate result matrices.
[0118] Step S5: The sparse matrix accumulator unit accumulates the second intermediate result matrices based on the corresponding second intermediate index matrices in each group to obtain a result matrix.
[0119] After grouping the second intermediate result matrices, the second intermediate result matrices included in each group and the corresponding second intermediate index matrices can be sequentially input to the sparse matrix accumulator unit row by row or column by column, and the sparse matrix accumulator unit further accumulates the second intermediate result matrices in each group. For the intermediate result after accumulation and the corresponding position information, they can be written back to the matrix register again through the above channel data write-back device.
[0120] In a possible case, in the above step S3, the matrix register also stores some intermediate results and the corresponding position information because there are matrix registers that do not reach the storage space threshold in step S3. In this case, the sparse matrix accumulator unit can accumulate the remaining intermediate results in the matrix register and the intermediate results included in the second intermediate result matrices in the corresponding groups.
[0121] During the process of the channel data write-back device storing the intermediate results and the corresponding position information back to the matrix register again, if the matrix register reaches the storage space threshold, the intermediate results and the corresponding position information in the matrix register can be stored in the memory in the form of a matrix again through the channel data write-to-memory device. Among them, the matrix corresponding to the intermediate result stored in step S5 can still be called the second intermediate result matrix, and similarly, the matrix corresponding to the stored position information can still be called the second intermediate index matrix.
[0122] After completing the accumulation of the above multiple groups of second intermediate result matrices, multiple second intermediate result matrices and the corresponding intermediate index matrices can be obtained in the memory again. For the multiple second intermediate result matrices and the corresponding intermediate index matrices obtained again, the processing of steps S4 to S5 can be performed again until there are no intermediate results with the same corresponding position information in the intermediate results after being accumulated by the sparse matrix accumulator unit.
[0123] In an example, after completing the accumulation of the second intermediate result matrices after each grouping, the intermediate results and the corresponding position information stored in the matrix register can be stored in the memory in the form of a matrix, and the number of elements included in the second intermediate result matrices stored in the memory again can be counted. If the number of elements counted twice in a row is the same, it can be considered that the accumulation of the intermediate results is completed. Then, based on the second intermediate result matrix and the corresponding second intermediate index matrix stored in the current memory, the result matrix of the matrix multiplication operation of the first sparse matrix and the second sparse matrix can be restored.
[0124] All of the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, and the embodiments of the present application will not elaborate on them one by one.
[0125] An embodiment of the present application also provides a processor, as Figure 5 shown. The processor includes a sparse matrix processing unit and a sparse matrix operation unit. The operations of the sparse matrix by the processor in the above embodiments may include:
[0126] The sparse matrix processing unit is configured to obtain a first numerical matrix and a first index matrix corresponding to a first sparse matrix, and a second numerical matrix and a second index matrix corresponding to a second sparse matrix. Wherein, the first numerical matrix is obtained by rearranging the non-zero elements in each column of the first sparse matrix, the first index matrix is used to indicate the row numbers of the non-zero elements in the first numerical matrix in the first sparse matrix, the second numerical matrix is obtained by rearranging the non-zero elements in each row of the second sparse matrix, and the second index matrix is used to indicate the column numbers of the non-zero elements in the second numerical matrix in the second sparse matrix.
[0127] The sparse matrix processing unit is configured to input the non-zero elements in each column of the first numerical matrix and the row numbers indicated by each column in the first index matrix, the non-zero elements in each row of the second numerical matrix and the column numbers indicated by each row in the second index matrix into the sparse matrix operation unit.
[0128] The sparse matrix operation unit is configured to perform an outer product operation on the non-zero elements in each column of the first numerical matrix and the non-zero elements in each row of the second numerical matrix based on the row numbers indicated by each column in the first index matrix and the column numbers indicated by each row in the second index matrix, to obtain a result matrix of the matrix multiplication operation of the first sparse matrix and the second sparse matrix.
[0129] In an implementable manner, the first numerical matrix includes a plurality of first numerical partitions. The plurality of first numerical partitions are obtained by dividing the first sparse matrix into a plurality of first sparse matrix regions by rows, and then rearranging the non-zero elements in each column of each first sparse matrix region. The first index matrix includes a first index partition corresponding to each first numerical partition.
[0130] The second numerical matrix includes a plurality of second numerical partitions. The plurality of second numerical partitions are obtained by dividing the second sparse matrix into a plurality of second sparse matrix regions by columns, and then rearranging the non-zero elements in each row of each second sparse matrix region. The second index matrix includes a second index partition corresponding to each second numerical partition.
[0131] In an implementable manner, each first numerical partition includes a plurality of first numerical blocks with a size of n×m, each first index partition includes a plurality of first index blocks with a size of n×m, each second numerical partition includes a plurality of second numerical blocks with a size of m×n, and each second index partition includes a plurality of second index blocks with a size of m×n.
[0132] In an implementable manner, a sparse matrix processing unit is configured to input, in a calculation order, each column element in the non-zero numerical blocks of each column of the first numerical blocks in each first numerical partition, each row element in the non-zero numerical blocks of each row of the second numerical blocks in each second numerical partition, the row numbers indicated by each column of the non-zero index blocks in each column of the first index blocks in each first index partition, and the column numbers indicated by each column of the non-zero index blocks in each row of the second index blocks in each second index partition, into the sparse matrix operation unit in sequence.
[0133] In an implementable manner, the sparse matrix operation unit includes a sparse matrix outer product subunit and a sparse matrix accumulation subunit.
[0134] The sparse matrix outer product subunit is configured to perform an outer product operation on each column element of the first numerical block and each row element of the second numerical block input each time, to obtain a first intermediate result matrix.
[0135] The sparse matrix outer product subunit is configured to form a first intermediate index matrix corresponding to the first intermediate result matrix from the row numbers indicated by each column of the first index block and the column numbers indicated by each column of the second index block input each time. The first intermediate index matrix is used to indicate the position information of each element in the first intermediate result matrix in the result matrix.
[0136] The sparse matrix accumulation subunit is configured to accumulate a plurality of first intermediate result matrices based on a plurality of first intermediate index matrices, to obtain a result matrix of the matrix multiplication operation performed by the first sparse matrix and the second sparse matrix.
[0137] In an implementable manner, the processor includes a plurality of matrix registers.
[0138] For the first first intermediate result matrix output by the sparse matrix outer product subunit and the corresponding first intermediate index matrix, the sparse matrix processing unit is configured to store each row intermediate result in the first intermediate result matrix and the position information of each corresponding row in the first intermediate index matrix into a plurality of matrix registers respectively.
[0139] For the first intermediate result matrix after the first output of the sparse matrix accumulation subunit and the corresponding first intermediate index matrix, the sparse matrix accumulation subunit is configured to accumulate the intermediate results of each row in the first intermediate result matrix and the intermediate results stored in multiple matrix registers respectively based on the row position information in the first intermediate index matrix and the position information stored in multiple matrix registers, and store the accumulated intermediate results and the corresponding position information back into multiple matrix registers respectively.
[0140] The sparse matrix accumulation subunit is configured to determine a result matrix based on the intermediate results and position information stored in multiple matrix registers.
[0141] In an implementable manner, when the storage space of the matrix register reaches the storage space threshold, the sparse matrix processing unit is configured to store the intermediate results and the corresponding position information stored in the matrix register in the form of a matrix into the memory to obtain a second intermediate result matrix and a second intermediate index matrix.
[0142] The sparse matrix processing unit is configured to group the second intermediate result matrix according to the number of matrix registers to obtain multiple groups of second intermediate result matrices.
[0143] The sparse matrix accumulation subunit is configured to accumulate the second intermediate result matrix based on the second intermediate index matrix corresponding to each group of second intermediate result matrices to obtain a result matrix.
[0144] The processor provided in the embodiment of the present application can be used to execute the operation method of the sparse matrix provided in the embodiment of the present application. For the specific execution process, see the introduction content of the operation method of the sparse matrix in the above embodiment, which will not be elaborated here. By using the processor provided in the embodiment of the present application to execute the sparse multiplication operation, the processor can obtain the first numerical matrix and the first index matrix corresponding to the first sparse matrix to be subjected to matrix operation, the second numerical matrix and the second index matrix corresponding to the second sparse matrix, and perform an outer product operation on the non-zero elements in the first sparse matrix and the second sparse matrix according to the row numbers of the non-zero elements in the first numerical matrix indicated by the first index matrix in the first sparse matrix and the column numbers of the non-zero elements in the second numerical matrix indicated by the second index matrix in the second sparse matrix, so as to obtain the result matrix of the matrix multiplication of the first sparse matrix and the second sparse matrix. In this way, a large number of zero elements in the first sparse matrix and the second sparse matrix can be avoided from participating in the matrix multiplication operation process, and thus the operation efficiency of the sparse matrix multiplication operation can be improved.
[0145] The embodiments of the present application also provide a computer program product containing instructions. The computer program product may be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute the operation method of the sparse matrix provided by the embodiments of the present application.
[0146] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the operation method of the sparse matrix provided by the embodiments of the present application.
[0147] In the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions. It should be understood that there is no logical or chronological dependence between "first" and "second", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first numerical matrix may be referred to as the second numerical matrix, and similarly, the second numerical matrix may be referred to as the first numerical matrix. The first numerical matrix and the second numerical matrix can both be collectively referred to as the numerical matrix, and in some cases, they may be separate and different numerical matrices.
[0148] In the present application, the meaning of the term "at least one" refers to one or more, and the meaning of the term "multiple" refers to two or more.
[0149] The above description is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A sparse matrix operation method, characterized in that: The method is executed by a processor, the processor includes a sparse matrix processing unit and a sparse matrix operation unit, and the method includes: The sparse matrix processing unit obtains a first numerical matrix and a first index matrix corresponding to the first sparse matrix, and a second numerical matrix and a second index matrix corresponding to the second sparse matrix, wherein the first numerical matrix is obtained by rearranging the non-zero elements in each column of the first sparse matrix, and the first index matrix is used to indicate the row numbers of the non-zero elements in the first numerical matrix in the first sparse matrix, and the second numerical matrix is obtained by rearranging the non-zero elements in each row of the second sparse matrix, and the second index matrix is used to indicate the column numbers of the non-zero elements in the second numerical matrix in the second sparse matrix; The sparse matrix processing unit inputs the non-zero elements of each column in the first numerical matrix and the row number indicated by each column in the first index matrix, and the non-zero elements of each row in the second numerical matrix and the column number indicated by each row in the second index matrix to the sparse matrix operation unit; The sparse matrix operation unit performs outer product operations on the non-zero elements of each column in the first numerical matrix and the non-zero elements of each row in the second numerical matrix based on the row number indicated by each column in the first index matrix and the column number indicated by each row in the second index matrix to obtain a result matrix of matrix multiplication operations performed on the first sparse matrix and the second sparse matrix.
2. The method according to claim 1, characterized in that The first numerical matrix includes a plurality of first numerical partitions, wherein the plurality of first numerical partitions are obtained by dividing the first sparse matrix into a plurality of first sparse matrix regions by rows, and then rearranging non-zero elements in each column of each first sparse matrix region, and the first index matrix includes a first index partition corresponding to each first numerical partition; The second numerical matrix includes multiple second numerical partitions, which are obtained by dividing the second sparse matrix into multiple second sparse matrix areas by columns, and then rearranging the non-zero elements in each row of each second sparse matrix area. The second index matrix includes second index partitions corresponding to each second numerical partition.
3. The method according to claim 2, characterized in that Each of the first numerical partitions includes a plurality of first numerical blocks of size n×m, each of the first index partitions includes a plurality of first index blocks of size n×m, each of the second numerical partitions includes a plurality of second numerical blocks of size m×n, and each of the second index partitions includes a plurality of second index blocks of size m×n.
4. The method according to claim 3, characterized in that The sparse matrix processing unit inputs the non-zero elements of each column in the first numerical matrix and the row number indicated by each column in the first index matrix, and the non-zero elements of each row in the second numerical matrix and the column number indicated by each row in the second index matrix to the sparse matrix operation unit, including: The sparse matrix processing unit inputs, in order of calculation, each column element in the non-zero numerical block in each column of the first numerical block in each first numerical partition, each row element in the non-zero numerical block in each row of the second numerical block in each second numerical partition, the row number indicated by each column of the non-zero index block in each column of the first index block in each first index partition, and the column number indicated by each column of the non-zero index block in each row of the second index block in each second index partition into the sparse matrix operation unit.
5. The method according to claim 4, characterized in that: The sparse matrix operation unit includes a sparse matrix outer product subunit and a sparse matrix accumulation subunit. The sparse matrix operation unit performs an outer product operation on non-zero elements in each column of the first numerical matrix and non-zero elements in each row of the second numerical matrix based on the row number indicated by each column in the first index matrix and the column number indicated by each row in the second index matrix to obtain a result matrix of the matrix multiplication operation performed on the first sparse matrix and the second sparse matrix, including: The sparse matrix outer product subunit performs an outer product operation on each column element of the first numerical block and each row element of the second numerical block input each time to obtain a first intermediate result matrix; The sparse matrix outer product subunit combines the row number indicated by each column of the first index block input each time and the column number indicated by each column of the second index block to form a first intermediate index matrix corresponding to the first intermediate result matrix, wherein the first intermediate index matrix is used to indicate the position information of each element in the first intermediate result matrix in the result matrix; The sparse matrix accumulation subunit accumulates multiple first intermediate result matrices based on multiple first intermediate index matrices to obtain a result matrix of performing matrix multiplication operation of the first sparse matrix and the second sparse matrix.
6. The method according to claim 5, characterized in that The processor includes a plurality of matrix registers, and the method further includes: For a first first intermediate result matrix output by the sparse matrix outer product subunit and a corresponding first intermediate index matrix, the sparse matrix processing unit stores each row of intermediate results in the first intermediate result matrix and each row position information corresponding to the first intermediate index matrix in the plurality of matrix registers respectively; The sparse matrix accumulation subunit accumulates a plurality of first intermediate result matrices based on a plurality of first intermediate index matrices to obtain a result matrix of matrix multiplication operation performed by the first sparse matrix and the second sparse matrix, including: For the first intermediate result matrix output by the sparse matrix accumulation subunit and the corresponding first intermediate index matrix, the sparse matrix accumulation subunit accumulates the intermediate results of each row in the first intermediate result matrix and the intermediate results respectively stored in the multiple matrix registers based on the position information of each row in the first intermediate index matrix and the position information respectively stored in the multiple matrix registers, and stores the accumulated intermediate results and the corresponding position information back to the multiple matrix registers respectively; The sparse matrix accumulation subunit determines the result matrix based on the intermediate results and position information stored in the plurality of matrix registers.
7. The method according to claim 6, characterized in that The method further comprises: When the storage space of the matrix register reaches the storage space threshold, the sparse matrix processing unit stores the intermediate results and corresponding position information stored in the matrix register into the memory in the form of a matrix to obtain a second intermediate result matrix and a second intermediate index matrix; The sparse matrix processing unit groups the second intermediate result matrix according to the number of the matrix registers to obtain multiple groups of second intermediate result matrices; The sparse matrix accumulation subunit determines the result matrix based on the intermediate results and position information stored in the plurality of matrix registers, including: The sparse matrix accumulation subunit accumulates the second intermediate result matrices based on the second intermediate index matrix corresponding to each group of second intermediate result matrices to obtain the result matrix.
8. A processor, characterized in that: The processor includes a sparse matrix processing unit and a sparse matrix operation unit, and the method includes: The sparse matrix processing unit is used to obtain a first numerical matrix and a first index matrix corresponding to the first sparse matrix, and a second numerical matrix and a second index matrix corresponding to the second sparse matrix, wherein the first numerical matrix is obtained by rearranging the non-zero elements in each column of the first sparse matrix, and the first index matrix is used to indicate the row numbers of the non-zero elements in the first numerical matrix in the first sparse matrix, and the second numerical matrix is obtained by rearranging the non-zero elements in each row of the second sparse matrix, and the second index matrix is used to indicate the column numbers of the non-zero elements in the second numerical matrix in the second sparse matrix; The sparse matrix processing unit is used to input the non-zero elements of each column in the first numerical matrix and the row number indicated by each column in the first index matrix, and the non-zero elements of each row in the second numerical matrix and the column number indicated by each row in the second index matrix to the sparse matrix operation unit; The sparse matrix operation unit is used to perform outer product operations on non-zero elements in each column of the first numerical matrix and non-zero elements in each row of the second numerical matrix based on the row number indicated by each column in the first index matrix and the column number indicated by each row in the second index matrix, so as to obtain a result matrix of matrix multiplication operations performed on the first sparse matrix and the second sparse matrix.
9. The processor according to claim 1, wherein: The first numerical matrix includes a plurality of first numerical partitions, wherein the plurality of first numerical partitions are obtained by dividing the first sparse matrix into a plurality of first sparse matrix regions by rows, and then rearranging non-zero elements in each column of each first sparse matrix region, and the first index matrix includes a first index partition corresponding to each first numerical partition; The second numerical matrix includes multiple second numerical partitions, which are obtained by dividing the second sparse matrix into multiple second sparse matrix areas by columns, and then rearranging the non-zero elements in each row of each second sparse matrix area. The second index matrix includes second index partitions corresponding to each second numerical partition.
10. The processor according to claim 9, characterized in that Each of the first numerical partitions includes a plurality of first numerical blocks of size n×m, each of the first index partitions includes a plurality of first index blocks of size n×m, each of the second numerical partitions includes a plurality of second numerical blocks of size m×n, and each of the second index partitions includes a plurality of second index blocks of size m×n.
11. The processor according to claim 10, characterized in that The sparse matrix processing unit is used to input, in order of calculation, each column element in the non-zero numerical block in each column of the first numerical block in each first numerical partition, each row element in the non-zero numerical block in each row of the second numerical block in each second numerical partition, the row number indicated by each column of the non-zero index block in each column of the first index block in each first index partition, and the column number indicated by each column of the non-zero index block in each row of the second index block in each second index partition into the sparse matrix operation unit.
12. The processor according to claim 11, characterized in that: The sparse matrix operation unit includes a sparse matrix outer product subunit and a sparse matrix accumulation subunit; The sparse matrix outer product subunit is used to perform an outer product operation on each column element of the first numerical block and each row element of the second numerical block input each time to obtain a first intermediate result matrix; The sparse matrix outer product subunit is used to combine the row number indicated by each column of the first index block input each time and the column number indicated by each column of the second index block to form a first intermediate index matrix corresponding to the first intermediate result matrix, wherein the first intermediate index matrix is used to indicate the position information of each element in the first intermediate result matrix in the result matrix; The sparse matrix accumulation subunit is used to accumulate multiple first intermediate result matrices based on multiple first intermediate index matrices to obtain a result matrix of performing matrix multiplication operation on the first sparse matrix and the second sparse matrix.
13. The processor according to claim 12, characterized in that The processor comprises a plurality of matrix registers, For the first first intermediate result matrix output by the sparse matrix outer product subunit and the corresponding first intermediate index matrix, the sparse matrix processing unit is used to store each row of intermediate results in the first intermediate result matrix and each row position information corresponding to the first intermediate index matrix in the plurality of matrix registers respectively; For the first intermediate result matrix output by the sparse matrix accumulation subunit and the corresponding first intermediate index matrix, the sparse matrix accumulation subunit is used to accumulate the intermediate results of each row in the first intermediate result matrix and the intermediate results respectively stored in the multiple matrix registers based on the position information of each row in the first intermediate index matrix and the position information respectively stored in the multiple matrix registers, and store the accumulated intermediate results and the corresponding position information back to the multiple matrix registers respectively; The sparse matrix accumulation subunit is used to determine the result matrix based on the intermediate results and position information stored in the multiple matrix registers.
14. The processor according to claim 13, characterized in that When the storage space of the matrix register reaches the storage space threshold, the sparse matrix processing unit is used to store the intermediate results and corresponding position information stored in the matrix register into the memory in the form of a matrix to obtain a second intermediate result matrix and a second intermediate index matrix; The sparse matrix processing unit is used to group the second intermediate result matrices according to the number of the matrix registers to obtain multiple groups of second intermediate result matrices; The sparse matrix accumulation subunit is used to accumulate the second intermediate result matrices based on the second intermediate index matrix corresponding to each group of second intermediate result matrices to obtain the result matrix.
15. A computing device, characterized in that: The computing device comprises a memory and a processor as described in any one of claims 8 to 14 above, wherein the memory stores at least one instruction, and the processor executes the at least one instruction to execute the method as described in any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program code, and when the computer program code is executed by a computing device, the computing device is caused to perform the method according to any one of claims 1 to 7.
17. A computer program product comprising instructions, characterized in that When the computer program product is executed on a computing device, the computing device is caused to execute the method according to any one of claims 1 to 7.