Operation method for sparse matrix, processor and computing device
By obtaining and rearranging the numerical and index matrices of sparse matrices, using outer product operations to avoid zero elements, solving the problem of low multiplication efficiency of sparse matrices and achieving more efficient operations.
Patent Information
- Application Number
- PCT/CN2024/100797
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-06-21
- Publication Date
- 2025-05-22
AI Technical Summary
The prior art is inefficient when performing sparse matrix multiplication because a large number of zero elements participate in the operation.
By obtaining the numerical matrix and index matrix corresponding to the sparse matrix, rearranging non-zero elements and using external product operations, avoiding zero elements participating in the operation, thereby improving operation efficiency.
It effectively avoids the participation of a large number of zero elements in the sparse matrix and improves the computational efficiency of sparse matrix multiplication.
Smart Images

Figure CN2024100797_22052025_PF_FP_ABST
Abstract
Description
Sparse matrix operation method, processor and computing device
[0001] This application claims priority to Chinese patent application No. 202311544688.5 filed on November 17, 2023, entitled “Sparse Matrix Operation Method, Processor and Computing Device,” the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a sparse matrix operation method, processor, and computing device. Background Art
[0003] Matrix multiplication is widely used in various scientific computing scenarios. Due to the varying density of the matrices involved, matrix multiplication is primarily categorized into dense and sparse matrix multiplication. Dense matrix multiplication is typically used in traditional machine learning and neural network model computations, while sparse matrix multiplication is primarily used in graph computing, graph neural networks, and multigrid solvers.
[0004] Since a sparse matrix includes a large number of zero elements, a large number of zero elements will participate in the matrix multiplication operation during the execution of sparse matrix multiplication, which leads to low efficiency of matrix operation.
[0005] Summary of the Invention
[0006] The embodiments of the present application provide a sparse matrix operation method, processor, and computing device, which can improve the efficiency of sparse matrix multiplication operations. The corresponding technical solutions are as follows:
[0007] In a first aspect, a sparse matrix operation method is provided. The method is executed by a processor, the processor including a sparse matrix processing unit and a sparse matrix operation unit. The method includes:
[0008] The sparse matrix processing unit obtains a first numerical matrix and a first index matrix corresponding to the first sparse matrix, and a second numerical matrix and a second index matrix corresponding to the second sparse matrix, wherein the first numerical matrix is obtained by rearranging the non-zero elements in each column of the first sparse matrix, the first index matrix is used to indicate the row numbers of the non-zero elements in the first numerical matrix in the first sparse matrix, and the second numerical matrix is obtained by rearranging the non-zero elements in each row of the second sparse matrix, and the second index matrix is used to indicate the column numbers of the non-zero elements in the second numerical matrix in the second sparse matrix. The sparse matrix processing unit inputs the non-zero elements in each column of the first numerical matrix and the row numbers indicated by each column in the first index matrix, the non-zero elements in each row of the second numerical matrix, and the column numbers indicated by each row in the second index matrix to the sparse matrix operation unit. The sparse matrix operation unit performs an outer product operation on the non-zero elements in each column of the first numerical matrix and the non-zero elements in each row of the second numerical matrix based on the row numbers indicated by each column in the first index matrix and the column numbers indicated by each row in the second index matrix, to obtain a result matrix of the matrix multiplication operation performed on the first sparse matrix and the second sparse matrix.
[0009] In the solution shown in the present application, the non-zero elements in the first numerical matrix indicated by the first index matrix can be used to perform an outer product operation on the non-zero elements in the first sparse matrix and the second sparse matrix according to the row numbers of the non-zero elements in the first numerical matrix indicated by the first index matrix in the first sparse matrix and the column numbers of the non-zero elements in the second numerical matrix indicated by the second index matrix in the second sparse matrix, thereby obtaining a result matrix of the matrix multiplication of the first sparse matrix and the second sparse matrix. In this way, a large number of zero elements in the first sparse matrix and the second sparse matrix can be avoided from participating in the matrix multiplication operation process, thereby improving the efficiency of performing sparse matrix multiplication operations.
[0010] In one implementable manner, the first numerical matrix includes a plurality of first numerical partitions, which are obtained by dividing the first sparse matrix into a plurality of first sparse matrix regions by rows, and then rearranging the non-zero elements in each column of each first sparse matrix region, and the first index matrix includes a first index partition corresponding to each first numerical partition. The second numerical matrix includes a plurality of second numerical partitions, which are obtained by dividing the second sparse matrix into a plurality of second sparse matrix regions by columns, and then rearranging the non-zero elements in each row of each second sparse matrix region, and the second index matrix includes a second index partition corresponding to each second numerical partition. In this way, by performing partition calculations on the first numerical matrix, the first index matrix, the second numerical matrix, and the second index matrix, the complexity of matrix operations can be reduced, thereby improving the operational efficiency of matrix operations.
[0011] In one achievable embodiment, each first value partition includes multiple first value blocks of size n×m, each first index partition includes multiple first index blocks of size n×m, each second value partition includes multiple second value blocks of size m×n, and each second index partition includes multiple second index blocks of size m×n. Thus, by performing block calculations on the first value partitions, first index partitions, second value partitions, and second index partitions, the complexity of matrix operations can be further reduced and the efficiency of matrix operations can be improved.
[0012] In an achievable manner, the sparse matrix processing unit inputs the non-zero elements of each column in the first numerical matrix and the row number indicated by each column in the first index matrix, the non-zero elements of each row in the second numerical matrix and the column number indicated by each row in the second index matrix into the sparse matrix operation unit, including: the sparse matrix processing unit inputs the elements of each column in the non-zero numerical block in each column of the first numerical block in each first numerical partition, the elements of each row in the non-zero numerical block in each row of the second numerical block in each second numerical partition, the row number indicated by each column of the non-zero index block in each column of the first index block in each first index partition, and the column number indicated by each column of the non-zero index block in each row of the second index block in each second index partition into the sparse matrix operation unit in sequence. In the scheme shown in the present application, by inputting the non-zero numerical blocks and the corresponding index blocks into the sparse matrix operation unit for cumulative processing, the sparse matrix operation unit can be prevented from performing cumulative processing on all-zero numerical blocks and all-zero index blocks, which can reduce the amount of calculation for cumulative processing and thereby improve the efficiency of matrix operations.
[0013] In one implementable manner, the sparse matrix operation unit includes a sparse matrix outer product subunit and a sparse matrix accumulation subunit. The sparse matrix operation unit performs an outer product operation on the non-zero elements of each column of the first numerical matrix and the non-zero elements of each row of the second numerical matrix based on the row number indicated by each column in the first index matrix and the column number indicated by each row in the second index matrix, to obtain a result matrix of the matrix multiplication operation performed on the first sparse matrix and the second sparse matrix, including: the sparse matrix outer product subunit performs an outer product operation on each column element of the first numerical block input each time and each row element of the second numerical block each time to obtain a first intermediate result matrix. The sparse matrix outer product subunit combines the row number indicated by each column of the first index block input each time and the column number indicated by each column of the second index block each time to form a first intermediate index matrix corresponding to the first intermediate result matrix, and the first intermediate index matrix is used to indicate the position information of each element in the first intermediate result matrix in the result matrix. The sparse matrix accumulation subunit accumulates multiple first intermediate result matrices based on multiple first intermediate index matrices to obtain a result matrix of the matrix multiplication operation performed on the first sparse matrix and the second sparse matrix.
[0014] In one implementable manner, the processor includes multiple matrix registers, and the method also includes: for the first first intermediate result matrix output by the sparse matrix outer product subunit and the corresponding first intermediate index matrix, the sparse matrix processing unit stores the intermediate results of each row in the first intermediate result matrix and the position information of each row corresponding to the first intermediate index matrix in multiple matrix registers respectively.
[0015] The sparse matrix accumulation subunit accumulates multiple first intermediate result matrices based on multiple first intermediate index matrices to obtain a result matrix of matrix multiplication operation performed by the first sparse matrix and the second sparse matrix, including: for the first intermediate result matrix after the first intermediate index matrix output by the sparse matrix accumulation subunit and the corresponding first intermediate index matrix, the sparse matrix accumulation subunit accumulates the intermediate results of each row in the first intermediate result matrix and the intermediate results stored in the multiple matrix registers based on the position information of each row in the first intermediate index matrix and the position information stored in the multiple matrix registers, and stores the accumulated intermediate results and the corresponding position information back to the multiple matrix registers. The sparse matrix accumulation subunit determines the result matrix based on the intermediate results and position information stored in the multiple matrix registers. In an embodiment of the present application, the multiple matrix registers cooperate with the sparse matrix accumulation subunit to realize highly parallel accumulation processing, thereby improving the efficiency of matrix operations.
[0016] In one achievable embodiment, the method further includes: when the storage space of the matrix register reaches a storage space threshold, the sparse matrix processing unit stores the intermediate results and corresponding position information stored in the matrix register in the form of a matrix to the memory, thereby obtaining a second intermediate result matrix and a second intermediate index matrix. The sparse matrix processing unit groups the second intermediate result matrices according to the number of matrix registers, thereby obtaining multiple groups of second intermediate result matrices.
[0017] The sparse matrix accumulation subunit determines a result matrix based on the intermediate results and position information stored in multiple matrix registers, including: the sparse matrix accumulation subunit accumulates the second intermediate result matrices based on the second intermediate index matrix corresponding to each group of second intermediate result matrices to obtain a result matrix. In an embodiment of the present application, by grouping the second intermediate result matrices, the complexity of the accumulation processing of the second intermediate result matrices can be reduced, thereby improving the efficiency of the matrix budget.
[0018] In a second aspect, a processor is provided, the processor including a sparse matrix processing unit and a sparse matrix operation unit, and a method including:
[0019] A sparse matrix processing unit is used to obtain a first numerical matrix and a first index matrix corresponding to the first sparse matrix, and a second numerical matrix and a second index matrix corresponding to the second sparse matrix, wherein the first numerical matrix is obtained by rearranging the non-zero elements in each column of the first sparse matrix, and the first index matrix is used to indicate the row numbers of the non-zero elements in the first numerical matrix in the first sparse matrix; the second numerical matrix is obtained by rearranging the non-zero elements in each row of the second sparse matrix, and the second index matrix is used to indicate the column numbers of the non-zero elements in the second numerical matrix in the second sparse matrix.
[0020] The sparse matrix processing unit is used to input the non-zero elements of each column in the first numerical matrix and the row number indicated by each column in the first index matrix, the non-zero elements of each row in the second numerical matrix and the column number indicated by each row in the second index matrix into the sparse matrix operation unit.
[0021] The sparse matrix operation unit is used to perform an outer product operation on the non-zero elements in each column of the first numerical matrix and the non-zero elements in each row of the second numerical matrix based on the row number indicated by each column in the first index matrix and the column number indicated by each row in the second index matrix, so as to obtain a result matrix of the matrix multiplication operation performed on the first sparse matrix and the second sparse matrix.
[0022] In one implementable manner, the first numerical matrix includes multiple first numerical partitions, which are obtained by dividing the first sparse matrix into multiple first sparse matrix regions by rows, and then rearranging the non-zero elements in each column of each first sparse matrix region. The first index matrix includes a first index partition corresponding to each first numerical partition.
[0023] The second numerical matrix includes multiple second numerical partitions, which are obtained by dividing the second sparse matrix into multiple second sparse matrix areas by columns, and then rearranging the non-zero elements in each row of each second sparse matrix area. The second index matrix includes second index partitions corresponding to each second numerical partition.
[0024] In one implementable manner, each first numerical partition includes multiple first numerical blocks of size n×m, each first index partition includes multiple first index blocks of size n×m, each second numerical partition includes multiple second numerical blocks of size m×n, and each second index partition includes multiple second index blocks of size m×n.
[0025] In one implementable manner, the sparse matrix processing unit is used to input, in order of calculation, each column element in the non-zero numerical block in each column of the first numerical block in each first numerical partition, each row element in the non-zero numerical block in each row of the second numerical block in each second numerical partition, the row number indicated by each column of the non-zero index block in each column of the first index block in each first index partition, and the column number indicated by each column of the non-zero index block in each row of the second index block in each second index partition into the sparse matrix operation unit.
[0026] In one implementable manner, the sparse matrix operation unit includes a sparse matrix outer product subunit and a sparse matrix accumulation subunit.
[0027] The sparse matrix outer product subunit is used to perform outer product operations on each column element of the first numerical block and each row element of the second numerical block input each time to obtain a first intermediate result matrix.
[0028] The sparse matrix outer product subunit is used to combine the row number indicated by each column of the first index block input each time and the column number indicated by each column of the second index block to form a first intermediate index matrix corresponding to the first intermediate result matrix. The first intermediate index matrix is used to indicate the position information of each element in the first intermediate result matrix in the result matrix.
[0029] The sparse matrix accumulation subunit is used to accumulate multiple first intermediate result matrices based on multiple first intermediate index matrices to obtain a result matrix of performing matrix multiplication operation on the first sparse matrix and the second sparse matrix.
[0030] In one implementable manner, the processor includes multiple matrix registers. For the first intermediate result matrix output by the sparse matrix outer product subunit and the corresponding first intermediate index matrix, the sparse matrix processing unit is used to store the intermediate results of each row in the first intermediate result matrix and the position information of each row corresponding to the first intermediate index matrix in the multiple matrix registers respectively.
[0031] For the first intermediate result matrix and the corresponding first intermediate index matrix output by the sparse matrix accumulation subunit, the sparse matrix accumulation subunit is used to accumulate the intermediate results of each row in the first intermediate result matrix and the intermediate results stored in the multiple matrix registers based on the position information of each row in the first intermediate index matrix and the position information stored in the multiple matrix registers, and store the accumulated intermediate results and the corresponding position information back to the multiple matrix registers.
[0032] The sparse matrix accumulation subunit is used to determine a result matrix based on intermediate results and position information stored in the plurality of matrix registers.
[0033] In one feasible manner, when the storage space of the matrix register reaches the storage space threshold, the sparse matrix processing unit is used to store the intermediate results and corresponding position information stored in the matrix register in the form of a matrix to the memory to obtain a second intermediate result matrix and a second intermediate index matrix.
[0034] The sparse matrix processing unit is used to group the second intermediate result matrices according to the number of matrix registers to obtain multiple groups of second intermediate result matrices.
[0035] The sparse matrix accumulation subunit is used to accumulate the second intermediate result matrices based on the second intermediate index matrix corresponding to each group of second intermediate result matrices to obtain a result matrix.
[0036] In a third aspect, a computing device is provided, which includes a memory and a processor as described in the second aspect above, wherein the memory stores at least one instruction, and the processor executes the at least one instruction to implement the method described in the first aspect above.
[0037] In a fourth aspect, a computer-readable storage medium is provided, which stores computer program code. When the computer program code is executed by a computing device, the computing device executes the method described in the first aspect above.
[0038] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed on a computing device, causes the computing device to execute the method as described in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] FIG1 is a schematic diagram of a matrix multiplication operation in the related art;
[0040] FIG2 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0041] FIG3 is a schematic diagram of a sparse compression matrix provided in an embodiment of the present application;
[0042] FIG4 is a flow chart of a sparse matrix operation method provided in an embodiment of the present application;
[0043] FIG5 is a schematic diagram of the structure of a processor provided in an embodiment of the present application;
[0044] FIG6 is a schematic diagram of pre-partitioning a sparse matrix according to an embodiment of the present application;
[0045] FIG7 is a schematic diagram of a sparse matrix outer product subunit performing an outer product operation provided in an embodiment of the present application;
[0046] FIG8 is a schematic diagram of a method for performing an accumulation operation by a sparse matrix accumulation subunit provided in an embodiment of the present application;
[0047] FIG9 is a schematic diagram of a flow chart of accumulation processing performed by a sparse matrix accumulation subunit according to an embodiment of the present application;
[0048] FIG10 is a schematic diagram of a method for grouping second intermediate results provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0050] Matrix multiplication is widely used in various scientific computing scenarios. The density of the matrices involved varies across different computing scenarios, and matrix multiplication is primarily categorized into dense and sparse matrix multiplication. Dense matrix multiplication is typically used in traditional machine learning and neural network model calculations, while sparse matrix multiplication is primarily used in graph computing, graph neural networks, and multigrid solvers.
[0051] Dense matrix multiplication refers to multiplying a dense matrix, while sparse matrix multiplication refers to multiplying a sparse matrix. A sparse matrix is one in which the percentage of non-zero elements is very small (e.g., less than 5%), while a dense matrix is one in which the percentage of non-zero elements is large.
[0052] The current optimization method for matrix multiplication is only effective for dense matrix multiplication. For example, the matrix that can be used for matrix multiplication can be Tiled (cut into blocks), and then the result matrix of the matrix multiplication can be calculated through the Tiling matrix after Tiling. Figure 1 is a schematic diagram of matrix multiplication in related technology. As shown in Figure 1, for matrix A and matrix B that perform matrix multiplication, matrix A can be divided into Tiling matrix A and B. 0,0 , Tiling matrix A 0,1 , Tiling matrix A 1,0 , Tiling matrix A 1,1 , the matrix B can be divided into the Tiling matrix B 0,0 , Tiling matrix B 0,1 , Tiling matrix B 1,0 , Tiling matrix B 1,1 Among them, the result matrix C of matrix multiplication operation performed on matrix A and matrix B can be divided into tiling matrix C 0,0 , Tiling matrix C 0,1, Tiling matrix C 1,0 , Tiling matrix C 1,1 To calculate the Tiling matrix C 0,0 For example, the Tiling matrix C 0,0 Equal to the Tiling matrix A 0,0 and the tiling matrix B 0,0 The product of and the tiling matrix A 0,1 and the tiling matrix B 0,1 This can reduce the size of the matrix and thus reduce the complexity of the matrix multiplication operation, thereby improving the efficiency of the matrix multiplication operation.
[0053] However, sparse matrices contain a large number of zero elements. Even after tiling, the resulting tiling matrix still contains a large number of zero elements, and even some tiling matrices are zero matrices. This constant presence of zero elements during matrix multiplication consumes significant computing and storage resources, resulting in low sparse matrix multiplication efficiency.
[0054] An embodiment of the present application provides a sparse matrix operation method that can optimize sparse matrix multiplication by combining the characteristics of sparse matrices, thereby further improving the execution efficiency of sparse matrix multiplication. Figure 2 is a computing device for executing a sparse matrix operation method provided by an embodiment of the present application. As shown in Figure 2, the computing device 200 may include: a bus 202, a processor 204, and a memory 206. The optional computing device 200 may also include a communication interface 208. The processor 204, the memory 206, and the communication interface 208 communicate via the bus 202. The computing device 200 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 200. The computing device 200 may be a device for running a model, a terminal, or a server. When the computing device 200 is a terminal, the computing device 200 includes but is not limited to a desktop computer, a mobile phone, a notebook, a tablet computer, etc. When the computing device 200 is a server, the computing device 200 may be a separate server or a server cluster composed of multiple servers. It may be a physical machine or a virtual machine or container virtualized by virtualization technology.
[0055] Bus 202 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, value buses, control buses, and the like. For ease of illustration, FIG2 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 202 may include a path for transmitting information between various components of computing device 200 (e.g., memory 206, processor 204, and communication interface 208).
[0056] The processor 204 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP). Alternatively, the processor may be a system on chip (SOC) including one or more of the above-mentioned CPU, GPU, MP, etc. The processor 204 may further include a sparse matrix processing unit and a sparse matrix operation unit. The sparse matrix processing unit may be a processor core, and the sparse matrix operation unit may be used to implement the sparse matrix multiplication operation provided in the embodiment of the present application.
[0057] The memory 206 may include a volatile memory, such as a random access memory (RAM). The memory 206 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD). The memory 206 stores an executable program code, and the processor 204 executes the executable program code to implement the sparse matrix operation method provided in the embodiment of the present application.
[0058] The communication interface 208 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 200 and other devices or a communication network.
[0059] To facilitate understanding of the sparse matrix operation method provided in the embodiment of the present application, the sparse compressed matrix used in the embodiment of the present application is first introduced below:
[0060] The sparse compressed matrix includes a data matrix and an index matrix. The data matrix is obtained by sequentially arranging the zero and non-zero elements in the sparse matrix by row or column.
[0061] In one example, the non-zero elements and zero elements included in each row (column) of the sparse matrix can be rearranged continuously to obtain a numerical matrix corresponding to the sparse matrix. The index matrix and the numerical matrix have the same size, and the distribution of zero elements and non-zero elements in the index matrix is consistent with the distribution of zero elements and non-zero elements in the numerical matrix. When the zero elements and non-zero elements in the sparse matrix are rearranged continuously by row to obtain the corresponding numerical matrix, each non-zero element in the index matrix is used to indicate the column number of the non-zero elements at the same position in the numerical matrix in the sparse matrix. When the zero elements and non-zero elements in the sparse matrix are rearranged continuously by row and column to obtain the corresponding numerical matrix, each non-zero element in the index matrix is used to indicate the row number of the non-zero elements at the same position in the numerical matrix in the sparse matrix. Wherein, the non-zero elements and zero elements are rearranged continuously, and the non-zero elements can be arranged before the zero elements or after the zero elements.
[0062] As shown in Figure 3, for the sparse matrix S, if the non-zero elements and zero elements are continuously rearranged by row, the corresponding numerical matrix A1 and the corresponding index matrix A2 can be obtained. If the non-zero elements and zero elements are continuously rearranged by column, the corresponding numerical matrix B1 and the corresponding index matrix B2 can be obtained. For example, the value of the 3rd row and 1st column in the index matrix A2 is used to indicate that the element in the 3rd row and 1st column of the numerical matrix A1 is in the 2nd column of the sparse matrix S. For another example, the value of the 3rd row and 3rd column in the index matrix B2 is used to indicate that the element in the 3rd row and 3rd column of the numerical matrix B1 is in the 4th row of the sparse matrix S. It should be noted that in the embodiment of the present application, the row numbers and column numbers of the matrices are arranged starting from 1.
[0063] An embodiment of the present application provides a method for operating a sparse matrix. This method can realize the multiplication operation of a sparse matrix by performing the outer product of the matrix on the basis of the above-mentioned sparse compressed matrix, thereby improving the efficiency of multiplication operation on the sparse matrix. Figure 4 is a flow chart of a method for operating a sparse matrix provided by an embodiment of the present application. The method can be executed by the computing device 200 shown in Figure 2 above, and specifically can be executed by the processor 204. As shown in Figure 5, the processor 204 may include at least a sparse matrix processing unit and a sparse matrix operation unit. Among them, the sparse matrix processing unit may refer to a hardware part in the processor 204 for processing a sparse matrix that performs a sparse multiplication operation, such as a processor core. Or the sparse matrix processing unit may also refer to a software program that is run by the processor to process a sparse matrix. The sparse matrix operation unit is a hardware unit in the processor 204 for operating the sparse matrix in this application.
[0064] Referring to FIG4 , the sparse matrix operation method provided in the embodiment of the present application includes:
[0065] Step 401: The sparse matrix processing unit obtains a first numerical matrix and a first index matrix corresponding to a first sparse matrix, and a second numerical matrix and a second index matrix corresponding to a second sparse matrix.
[0066] The first numerical matrix is obtained by rearranging the non-zero elements in each column of the first sparse matrix, and the first index matrix is used to indicate the row numbers of the non-zero elements in the first numerical matrix in the first sparse matrix. The second numerical matrix is obtained by rearranging the non-zero elements in each row of the second sparse matrix, and the second index matrix is used to indicate the column numbers of the non-zero elements in the second numerical matrix in the second sparse matrix.
[0067] The first sparse matrix and the second sparse matrix are two sparse matrices to be subjected to matrix multiplication. In one example, the first sparse matrix is the left matrix to be subjected to matrix multiplication, and the second sparse matrix is the right matrix to be subjected to matrix multiplication. The first numerical matrix corresponding to the first sparse matrix is the matrix obtained by continuously rearranging the zero elements and non-zero elements in each column of the first sparse matrix. For example, the non-zero elements can be arranged in the order of row numbers, and the remaining zero elements can be arranged behind the non-zero elements. The second numerical matrix corresponding to the second sparse matrix is the matrix obtained by continuously rearranging the zero elements and non-zero elements in each row of the second sparse matrix. For example, the non-zero elements can be arranged in the order of column numbers, and the remaining zero elements can be arranged behind the non-zero elements.
[0068] In implementation, an application involving sparse matrix multiplication may be run in a computing device, such as a high performance computing (HPC) application, an artificial intelligence (AI) application, a graph computing application, etc. During operation, when the application needs to perform sparse matrix multiplication, it may send a matrix multiplication operation request to the processor. In response to the matrix multiplication operation request, the sparse matrix processing unit in the processor may read a first numerical matrix and a first index matrix corresponding to a first sparse matrix to be subjected to the matrix multiplication operation, and a second numerical matrix and a second index matrix corresponding to a second sparse matrix from the memory or storage (such as a hard disk) of the computing device.
[0069] If the first numerical matrix and the first index matrix corresponding to the first sparse matrix and the second numerical matrix and the second index matrix corresponding to the second sparse matrix are not stored in the memory or storage, the sparse matrix processing unit can read the first sparse matrix and the second sparse matrix. Wherein, the first sparse matrix and the second sparse matrix can be, but are not limited to, CSC (Compressed Sparse Column Format, compressed sparse column format), CSR (Compressed Sparse Row Format, compressed sparse row format), COO (Coordinate Format, coordinate format) and other formats. After reading the first sparse matrix and the second sparse matrix, the sparse matrix processing unit can generate the first numerical matrix and the first index matrix corresponding to the first sparse matrix, and the second numerical matrix and the second index matrix corresponding to the second sparse matrix in the manner shown in Figure 3 above.
[0070] Step 402: The sparse matrix processing unit inputs the non-zero elements of each column in the first numerical matrix and the row number indicated by each column in the first index matrix, the non-zero elements of each row in the second numerical matrix and the column number indicated by each row in the second index matrix to the sparse matrix operation unit.
[0071] The sparse matrix processing unit can input the row number indicated by each column in the first index matrix corresponding to the non-zero elements of each column in the first numerical matrix and the column number indicated by each row in the second index matrix corresponding to the non-zero elements of each row in the second numerical matrix into the sparse matrix operation unit in the order of performing the matrix outer product on the first numerical matrix and the second numerical matrix. That is, the non-zero elements in the first column of the first numerical matrix and the first index matrix and the non-zero elements in the first row of the second numerical matrix and the second index matrix, the non-zero elements in the second column of the first numerical matrix and the first index matrix and the non-zero elements in the second row of the second numerical matrix and the second index matrix, ..., the non-zero elements in the nth column of the first numerical matrix and the first index matrix and the non-zero elements in the nth row of the second numerical matrix and the second index matrix are input into the sparse matrix operation unit in sequence.
[0072] In one example, if there is a column of the first numerical matrix that is all zero elements or a row of the second numerical matrix that is all zero elements, then the column of the first numerical matrix (and the first index matrix) that is all zero elements and the corresponding row in the corresponding second numerical matrix (and the second index matrix) can be skipped and input into the sparse matrix operation unit, or the row of the second numerical matrix (and the second index matrix) that is all zero elements and the corresponding column in the corresponding first numerical matrix (and the first index matrix) can be skipped and input into the sparse matrix operation unit. For example, if the i-th column in the first numerical matrix is a zero element, then the i-th column of the first numerical matrix (and the first index matrix) and the i-th row in the second numerical matrix (and the second index matrix) can be not input into the sparse matrix operation unit. In implementation, the column number of the first numerical matrix that is all zero elements or the row number of the second numerical matrix that is zero elements can be marked by specifying a symbol. By skipping the columns of all zero elements and the rows of all zero elements and inputting them into the sparse matrix operation unit, the amount of calculation of the sparse matrix operation unit can be reduced, thereby improving the calculation efficiency of the sparse matrix operation unit.
[0073] Step 403: The sparse matrix operation unit performs an outer product operation on the non-zero elements in each column of the first numerical matrix and the non-zero elements in each row of the second numerical matrix based on the row number indicated by each column in the first index matrix and the column number indicated by each row in the second index matrix to obtain a result matrix of the matrix multiplication operation performed on the first sparse matrix and the second sparse matrix.
[0074] For each input of a column of non-zero elements in the first numerical matrix and a row of non-zero elements in the second numerical matrix, the sparse matrix operation unit may perform an outer product operation on the column of non-zero elements and the row of non-zero elements to obtain an intermediate result matrix. That is, the sparse matrix operation unit may multiply the i-th non-zero element in the column of non-zero elements with the j-th non-zero element in the row of non-zero elements to obtain the element value corresponding to the element in the i-th row and j-th column of the intermediate result matrix during the outer product operation (hereinafter referred to as the intermediate result).
[0075] For each input of a column of non-zero elements in the first index matrix and a row of non-zero elements in the second index matrix, the sparse matrix operation unit can form an intermediate index matrix corresponding to the intermediate result matrix with the column of non-zero elements and the row of non-zero elements in the order of the outer product operation. Wherein, each element in the intermediate index matrix is the position information of the intermediate result at the same position in the intermediate result matrix in the result matrix. For example, the i-th non-zero element of a column of non-zero elements in the first index matrix is "2", and the j-th non-zero element of a row of non-zero elements in the second index matrix is "3", then the position information formed is "2, 3", that is, the element value of the i-th row and j-th column in the corresponding intermediate index matrix is "2, 3". This element value is used to indicate the position of the intermediate result of the i-th row and j-th column of the intermediate result matrix in the result matrix, that is, the position of the intermediate result of the i-th row and j-th column in the result matrix is the 2nd row and 3rd column.
[0076] For each intermediate result matrix obtained, the sparse matrix operation unit can accumulate the elements in the intermediate result matrix according to the intermediate index matrix corresponding to each intermediate result matrix, that is, accumulate the elements corresponding to the same position information in the intermediate result matrix. After the accumulation, the accumulated results and the intermediate results remaining in each intermediate result matrix that have not been accumulated are the non-zero elements in the result matrix, and the position information corresponding to the accumulated results and the position information corresponding to the intermediate results remaining in the intermediate result matrix that have not been accumulated are the position information of the non-zero elements in the result matrix in the result matrix. In this way, the result matrix of the matrix multiplication operation performed by the first sparse matrix and the second sparse matrix can be obtained based on the accumulated results, the intermediate results that have not been accumulated, and the corresponding position information. After obtaining the result matrices corresponding to the first sparse matrix and the second sparse matrix, the sparse matrix processing unit can return the result to the application requesting sparse matrix multiplication.
[0077] In an embodiment of the present application, a first numerical matrix and a first index matrix corresponding to a first sparse matrix to be subjected to matrix operation, a second numerical matrix and a second index matrix corresponding to a second sparse matrix are obtained, and an outer product operation is performed on the non-zero elements in the first numerical matrix indicated by the first index matrix in the first sparse matrix and a column number in the second sparse matrix indicated by the second index matrix on the non-zero elements in the second numerical matrix, thereby obtaining a result matrix of matrix multiplication of the first sparse matrix and the second sparse matrix. In this way, a large number of zero elements in the first sparse matrix and the second sparse matrix can be avoided from participating in the matrix multiplication operation process, thereby improving the efficiency of performing sparse matrix multiplication operations.
[0078] The embodiment of the present application further provides a method for partitioning the first numerical matrix, the first index matrix, the second numerical matrix, and the second numerical matrix as follows:
[0079] The first numerical matrix includes a plurality of first numerical partitions, each of which is obtained by dividing the first sparse matrix into a plurality of first sparse matrix regions by rows, and then rearranging the non-zero elements in each column of each first sparse matrix region. The first index matrix includes a first index partition corresponding to each first numerical partition. The second numerical matrix includes a plurality of second numerical partitions, each of which is obtained by dividing the second sparse matrix into a plurality of second sparse matrix regions by columns, and then rearranging the non-zero elements in each row of each second sparse matrix region. The second index matrix includes a second index partition corresponding to each second numerical partition.
[0080] In implementation, a pre-partition parameter for pre-partitioning the first sparse matrix may be set, wherein the pre-partition parameter may be the number of rows for pre-partitioning the first sparse matrix or the number of columns for pre-partitioning the first sparse matrix (the number of rows is equal to the number of columns), or the number of partitions of the first sparse matrix and the second sparse matrix.
[0081] When the pre-partitioning parameters are the number of rows for pre-partitioning the first sparse matrix and the number of columns for pre-partitioning the first sparse matrix, the first sparse matrix can be partitioned in sequence starting from the first row of the first sparse matrix according to the number of rows, to obtain a plurality of first sparse matrix regions corresponding to the first sparse matrix, that is, the number of rows included in each first sparse matrix region is equal to the number of rows. Wherein, the number of rows in the last first sparse matrix region in the first sparse matrix can be equal to or less than the number of rows. Similarly, the second sparse matrix can be partitioned starting from the first column of the second sparse matrix according to the number of columns, to obtain a plurality of second sparse matrix regions corresponding to the second sparse matrix, that is, the number of columns included in each second sparse matrix region is equal to the number of columns. Wherein, the number of columns in the last second sparse matrix region in the second sparse matrix can be equal to or less than the number of columns.
[0082] When the pre-partition parameter is the number of partitions of the first sparse matrix and the second sparse matrix, the first sparse matrix can be evenly divided into multiple first sparse matrix regions in the column direction according to the number, and the second sparse matrix can be evenly divided into multiple second sparse matrix regions in the row direction according to the number.
[0083] Figure 6 is a schematic diagram of pre-partitioning a sparse matrix provided by an embodiment of the present application. In Figure 6, the number of rows of the left matrix and the number of columns of the right matrix are both x, and the number of rows or columns corresponding to the pre-partitioning parameter is y, where x=3y. Accordingly, the left matrix can be divided into three first sparse matrix regions in the column direction, each of which includes y rows, and the right matrix can be divided into three second sparse matrix regions in the row direction, each of which includes y columns.
[0084] After the first sparse matrix is divided into a plurality of first sparse matrix areas, the zero elements in each column of each first sparse matrix area can be continuously rearranged to obtain a first numerical matrix including each first numerical partition. After the second sparse matrix is divided into a plurality of second sparse matrix areas, the zero elements in each row of each second sparse matrix area can be continuously rearranged to obtain a second numerical matrix including each second numerical partition. Accordingly, a first index partition corresponding to each first numerical partition in the first numerical matrix and a second index partition corresponding to each second numerical partition in the second numerical matrix can be generated. The non-zero elements in the first index partition are used to indicate the row number of the non-zero elements at the same position in the corresponding first numerical partition in the first sparse matrix, and the non-zero elements in the second index partition are used to indicate the column number of the non-zero elements at the same position in the corresponding second numerical partition in the second sparse matrix.
[0085] In an embodiment of the present application, a first sparse matrix can be divided into a plurality of first sparse matrix regions, and a second sparse matrix can be divided into a plurality of second sparse matrix regions. Then, an outer product operation is performed on each first sparse matrix region and each of the plurality of second sparse matrix regions to implement a matrix multiplication operation of the first sparse matrix and the second sparse matrix. This can reduce the matrix size during the outer product operation and further improve the efficiency of the sparse matrix multiplication operation.
[0086] In an embodiment of the present application, each first numerical partition can be further divided into multiple first numerical blocks of size n×m, each first index partition can be further divided into multiple first index blocks of size n×m, each second numerical partition can be further divided into multiple second numerical blocks of size m×n, and each second index partition can be further divided into multiple second index blocks of size m×n.
[0087] Among them, n×m and m×n are pre-set tiling parameters for tiling the first numerical partition and the second numerical partition, and the first numerical block, the first index block, the second numerical block and the second index block obtained after tiling can be collectively referred to as Tiling blocks. In the present application, the first numerical partition and the first index partition corresponding to the first sparse matrix, the Tiling blocks included in the first numerical partition and the first index partition, and the second numerical partition and the second index partition corresponding to the second sparse matrix, the Tiling blocks included in the second numerical partition and the second index partition, can be pre-divided, that is, the sparse matrix processing unit can obtain them directly from the memory or storage. Alternatively, the sparse matrix processing unit can divide the first sparse matrix and the second sparse matrix (or the first numerical partition, the first index partition, the second numerical partition and the second index partition) after obtaining them.
[0088] In an embodiment of the present application, the first numerical partition, the first index partition, the second numerical partition and the second index partition can be further divided into Tiling blocks, and then outer product operations can be performed on the Tiling blocks respectively. This can further reduce the matrix size during the outer product operation and further improve the efficiency of multiplication operations on sparse matrices.
[0089] The above-mentioned pre-partitioning and tiling division of the first sparse matrix and the second sparse matrix can be performed before step 402. The following describes the process of performing outer product operation on the tiling block to realize the multiplication operation of the first sparse matrix and the second sparse matrix in this application:
[0090] In one implementable manner, the processing of the above-mentioned step 402 may include: the sparse matrix processing unit inputs, in accordance with the calculation order, each column element in the non-zero numerical block in each column of the first numerical block in each first numerical partition, each row element in the non-zero numerical block in each row of the second numerical block in each second numerical partition, the row number indicated by each column of the non-zero index block in each column of the first index block in each first index partition, and the column number indicated by each column of the non-zero index block in each row of the second index block in each second index partition into the sparse matrix operation unit.
[0091] In implementation, each first numerical partition in the first numerical matrix can be operated with each second numerical partition in the second numerical matrix. As shown in Figure 6, the first numerical partition a can first perform matrix operations with the second numerical partition a, the second numerical partition b and the second numerical partition c respectively, and then the first numerical partition b can perform matrix operations with the second numerical partition a, the second numerical partition b and the second numerical partition c respectively, and finally the first numerical partition c can perform matrix operations with the second numerical partition a, the second numerical partition b and the second numerical partition c respectively. Among them, when the first numerical partition and the second numerical partition are operated, it is only the operation between the non-zero matrix blocks included in the first numerical partition and the second numerical partition. Since there will be a large number of zero matrix blocks, that is, Tiling blocks that are all zero, after the first sparse matrix and the second sparse matrix are pre-partitioned and Tiling is performed, only the non-zero matrix blocks in the first numerical partition and the second numerical partition are operated, which can greatly reduce the Tiling blocks involved in the operation, so the efficiency of the matrix operation of the subsequent sparse matrix operation unit can be improved.
[0092] For any first numerical partition and second numerical partition that are operated, each first numerical block belonging to a non-zero matrix block in each Tiling column in the first numerical partition can be subjected to a matrix outer product operation with each second numerical block belonging to a non-zero block in the corresponding Tiling row in the second numerical partition. For example, each first numerical block belonging to a non-zero block in the i-th Tiling column in the first numerical partition can be subjected to a matrix outer product operation with each second numerical block belonging to a non-zero block in the i-th Tiling row in the second numerical partition. Wherein, the Tiling column refers to the column composed of the Tiling blocks in the first numerical partition, the Tiling row refers to the row composed of the Tiling blocks in the second numerical partition, and the Tiling row corresponding to the Tiling column means that the column number of the Tiling column is the same as the row number of the Tiling row.
[0093] As shown in Figure 6, the three non-zero tiling blocks in the first tiling column of the first numerical partition a can be outer-producted with the two non-zero tiling blocks in the first tiling row of the second numerical partition a. The two non-zero tiling blocks in the second tiling column of the first numerical partition a can be outer-producted with the two non-zero tiling blocks in the second tiling row of the second numerical partition a.
[0094] The technicians can pre-set the calculation order of the first numerical partition and the second numerical partition, as well as the calculation order of each column of the first numerical block and each row of the corresponding second numerical block in the first numerical partition and the second numerical partition. In implementation, the sparse matrix processing unit can input each column element of the non-zero numerical block in each column of the first numerical block in each first numerical partition, each row element of the non-zero numerical block in each row of the second numerical block in each second numerical partition, each column element of the non-zero index block in each column of the first index block in each first index partition, and each row element of the non-zero index block in each row of the second index block in each second index partition into the sparse matrix operation unit in sequence according to the set calculation order.
[0095] In one example, the sparse matrix operation unit provided in the embodiment of the present application includes a sparse matrix outer product subunit and a sparse matrix accumulation subunit. The processing of the above step 403 may include:
[0096] The sparse matrix outer product subunit performs an outer product operation on each column element of the first numerical block and each row element of the second numerical block for each input, to obtain a first intermediate result matrix. The row number indicated by each column of the first index block and the column number indicated by each column of the second index block for each input are combined to form a first intermediate index matrix corresponding to the first intermediate result matrix. The first intermediate index matrix is used to indicate the position information of each element in the first intermediate result matrix in the result matrix.
[0097] Figure 7 is a schematic diagram of a method for a sparse matrix outer product subunit to perform an outer product operation provided by an embodiment of the present application. As shown in Figure 7, the input of the sparse matrix outer product subunit is a first numerical block (Value1) and a corresponding first index block (Index1), a second numerical block (Value2) and a corresponding second index block (Index2). The output of the sparse matrix outer product subunit is a plurality of groups of first intermediate result matrices and corresponding first intermediate index matrices. Among them, the first intermediate result matrix (first intermediate index matrix) output by the sparse matrix outer product subunit is equal to the number of columns of the input first numerical block (or the number of rows of the second numerical block).
[0098] In implementation, a tiling block row-column vector matching device may be run in the sparse matrix processing unit, and the tiling block row-column vector matching device may input a column of elements in the first numerical block and a row of elements in the second numerical block, and a column of elements in the first index block and a row of elements in the second index block into the sparse matrix outer product subunit, so that the sparse matrix outer product subunit outputs a first intermediate index matrix and a corresponding first intermediate index matrix. For example, Value1 is a 4×2 identity matrix, Value2 is a 2×4 identity matrix, The sparse matrix outer product subunit can generate two first intermediate index matrices, which are and And can generate two 2×4 identity matrices.
[0099] In one example, a technician can pre-set a sparse outer product (SpSMEFMOPA) instruction. During the sparse matrix multiplication operation, the technician can call the SpSMEFMOPA instruction to implement the outer product operation of the sparse matrix outer product subunit on the first numerical block and the second numerical block of the input. In addition, the technician can set the specific circuit structure of the sparse matrix outer product subunit according to the method for performing the outer product operation of the sparse matrix outer product subunit provided in the embodiment of the present application. The specific circuit structure of the sparse matrix outer product subunit will not be described in detail in the embodiment of the present application.
[0100] The sparse matrix accumulation subunit accumulates multiple first intermediate result matrices based on multiple first intermediate index matrices to obtain a result matrix of performing a matrix multiplication operation on the first sparse matrix and the second sparse matrix.
[0101] Fig. 8 is a schematic diagram of a method for performing accumulation operation of a sparse matrix accumulation subunit provided by an embodiment of the present application. As shown in Figure 8, the input of the sparse matrix accumulation subunit is the first intermediate result matrix (Value3) and the corresponding first intermediate index matrix (Index2), the first intermediate result matrix (Value4) and the corresponding first intermediate index matrix (Index4). When the same position information does not exist in the second intermediate index matrix of the input, the sparse matrix accumulation subunit does not need to accumulate the second intermediate result matrix, and the second numerical index matrix and the corresponding second intermediate index matrix of the output of the sparse matrix accumulation subunit are the first numerical index matrix and the first intermediate index matrix of the input. When the same position information exists in the second intermediate index matrix of the input, the sparse matrix accumulation subunit can accumulate the intermediate results corresponding to the two first intermediate result matrices for the same position information to obtain the accumulation result, and output the accumulation result and the corresponding position information, as well as the intermediate result and the corresponding position information that have not been accumulated. The accumulated result and the intermediate result that has not been accumulated are elements included in the second intermediate result matrix, and the accumulated result and the intermediate result that has not been accumulated and the corresponding position information are elements included in the second intermediate index matrix.
[0102] In one example, a technician can pre-set a sparse bitwise addition (SpSMEFADD) instruction. During the sparse matrix multiplication operation, the sparse matrix accumulation subunit can be called to implement the accumulation processing of the first intermediate result matrix by the sparse matrix accumulation subunit. In addition, the technician can set the specific circuit structure of the sparse matrix accumulation subunit according to the method for performing accumulation processing by the sparse matrix outer product subunit provided in the embodiment of the present application. The specific circuit structure of the sparse matrix accumulation subunit will not be described in detail in the embodiment of the present application. In addition, the number of sparse matrix outer product subunits and sparse matrix accumulation subunits included in the sparse matrix operation unit is not limited, and can be one or more.
[0103] In one example, the processor further includes an on-chip multi-channel storage device, which may include multiple storage channels, and each storage channel may correspond to a matrix register. The matrix register may store the accumulated results stored by the sparse matrix accumulation subunit and the intermediate results that have not been accumulated and the corresponding position information in the form of a matrix. In particular, a multi-channel data mapping device may also be run in the sparse matrix processing unit. The multi-channel data mapping device may store the intermediate results and the position information corresponding to the second intermediate index matrix included in the second intermediate result matrix output by the sparse matrix accumulation subunit in the matrix registers on the multiple storage channels according to the set storage strategy. For example, the multi-channel data mapping device may store the rows in the second intermediate value matrix and the second index value matrix in different matrix registers according to the number of matrix registers and the size of the second intermediate value matrix. In particular, the intermediate results and the corresponding position information may be stored in the matrix registers in the form of a matrix.
[0104] The following further describes an exemplary accumulation process based on the sparse matrix accumulation subunit executing the accumulation operation unit to obtain the first sparse matrix and the second sparse matrix to perform the matrix multiplication operation:
[0105] Step S1: The sparse matrix processing unit stores the intermediate results of each row in the first intermediate result matrix output by the sparse matrix outer product sub-unit and the position information of each row corresponding to the first intermediate index matrix in a plurality of matrix registers respectively.
[0106] Since there is no intermediate result corresponding to the same position information in the first intermediate result matrix, the first intermediate result matrix output by the sparse matrix outer product subunit can be not accumulated first, and the intermediate results of each row of the first intermediate result matrix and the corresponding position information of each row in the first intermediate index matrix can be stored in multiple matrix registers respectively. In one example, the number of rows of the intermediate index matrix and the corresponding intermediate value matrix stored in each matrix register can be determined based on the number of multiple matrix registers. For example, if the intermediate index matrix includes 4 rows and the number of matrix registers is 4, then each matrix register can store a row of intermediate results in the first intermediate result matrix and store a row of position information in the first intermediate index matrix.
[0107] Step S2: For the first intermediate result matrix and the corresponding first intermediate index matrix output by the sparse matrix accumulation subunit, the sparse matrix accumulation subunit accumulates the intermediate results of each row in the first intermediate result matrix and the intermediate results stored in the multiple matrix registers based on the position information of each row in the first intermediate index matrix and the position information stored in the multiple matrix registers, and stores the accumulated intermediate results and the corresponding position information back to the multiple matrix registers.
[0108] After the sparse matrix accumulation subunit outputs the first first intermediate result matrix and the corresponding first intermediate index matrix, the sparse matrix processing unit can input the intermediate results and corresponding position information in the second and subsequent first intermediate result matrices output by the sparse matrix accumulation subunit into the sparse matrix accumulation subunit, and input the intermediate results and corresponding position information stored in multiple matrix registers into the sparse matrix accumulation subunit.
[0109] The sparse matrix accumulation subunit can accumulate the intermediate results according to the input position information, that is, accumulate the intermediate results with the same position information to obtain the accumulated result. The position information corresponding to the accumulated result remains unchanged and is still the position information before accumulation. The sparse matrix accumulation subunit can output the accumulated result and the corresponding position information, and output the intermediate result that has not been accumulated and the corresponding position information. Among them, the accumulated result can also be called an intermediate result, and the accumulated result and the intermediate result that have not been accumulated are the intermediate results included in the second intermediate result matrix, and the position information corresponding to the accumulated result and the intermediate result that have not been accumulated is the information corresponding to the second intermediate index matrix.
[0110] In which, the size of the second intermediate result matrix can be the same as the size of the first intermediate result matrix. In one example, the sparse matrix processing unit can combine the accumulated results output by the sparse matrix accumulation subunit and the intermediate results that have not been accumulated into a second intermediate result matrix in a row-first or column-first manner, and combine the corresponding position information into a second intermediate index matrix.
[0111] FIG9 is a flow chart of an accumulation process performed by a sparse matrix accumulation subunit according to an embodiment of the present application. As shown in FIG9 , a channel data loading device, a channel data writing device, and a channel data writing memory device may be run in the sparse matrix processing unit.
[0112] The channel data loading device can sequentially load the intermediate results and position information stored in the matrix register on each storage channel according to a set loading order. The loaded intermediate results and position information can be the intermediate results in the first intermediate result matrix output by the sparse matrix outer product subunit and the position information in the first intermediate index matrix, or can also be the intermediate results in the second intermediate result matrix output by the sparse matrix accumulation subunit and the position information in the second intermediate index matrix.
[0113] After the channel data loading device loads the intermediate results and corresponding position information stored in a matrix register each time, the loaded intermediate results and position information can be input into the sparse matrix accumulation subunit. At the same time, the sparse matrix outer product subunit can input the intermediate results in the output first intermediate result matrix and the position information corresponding to the first intermediate index matrix into the sparse matrix accumulation subunit (not shown in Figure 9). In one example, a row of intermediate results in the first intermediate result matrix output by the sparse matrix outer product subunit and the intermediate results and corresponding position information stored in a matrix register loaded by the channel data loading device can be input into the sparse matrix accumulation subunit for accumulation processing. By calling SpSMEAdd to execute, the sparse matrix accumulation subunit accumulates the input intermediate result information according to the position information input each time, outputs the accumulated result and the corresponding position information, and outputs the intermediate result and the corresponding position information that have not been accumulated.
[0114] The channel data write-back device can write the intermediate results in the second intermediate result matrix outputted by the sparse matrix accumulation subunit each time and the corresponding position information in the second intermediate index matrix back to the matrix register, such as the matrix registers corresponding to storage channels 1 to 4 in FIG9 . In one example, in one accumulation process of the sparse matrix accumulation subunit, the matrix register loaded by the channel data loading device and the matrix register written by the channel data write-back device are the same matrix register. When the channel data write-back device writes the intermediate results and the corresponding position information to the matrix register, the matrix register can be overwritten.
[0115] Step S3: When the storage space of the matrix register reaches the storage space threshold, the sparse matrix processing unit stores the intermediate results and corresponding position information stored in the matrix register into the memory in the form of a matrix to obtain a second intermediate result matrix and a second intermediate index matrix.
[0116] During step S2, when the storage space of any matrix register reaches the storage space threshold, the matrix register, the stored intermediate results, and the corresponding position information may be stored in the memory in the form of a matrix. The matrix storing the intermediate results is the second intermediate result matrix, and the matrix storing the position information is the second intermediate index matrix. The fact that the storage space of the matrix register reaches the storage space threshold may mean that the storage space of the matrix register is full.
[0117] Step S4: The sparse matrix processing unit groups the second intermediate result matrices according to the number of matrix registers to obtain multiple groups of second intermediate result matrices.
[0118] After step S2 and step S3, the memory will store multiple second intermediate result matrices and the second intermediate index matrix corresponding to each second intermediate result matrix. Since there may still be intermediate results corresponding to the same position information in the second intermediate result matrix. Therefore, a sparse matrix accumulation subunit is required to further perform accumulation processing on the multiple second intermediate result matrices in the memory. In order to improve the efficiency of the accumulation processing of the second intermediate result matrix and reduce the number of times the second intermediate result matrix is accumulated, a branch merge scheduling device can also be run in the sparse matrix processing unit. As shown in Figure 10, Figure 10 is a schematic diagram of a method for grouping second intermediate results provided by an embodiment of the present application. The branch merge scheduling device can group multiple second intermediate results stored in the memory, and then the sparse matrix accumulation subunit can perform accumulation processing on the second intermediate result matrix included in each group, thereby reducing the number of times the second intermediate result matrix is accumulated. Among them, each group of second intermediate result matrices also includes a second intermediate index matrix corresponding to the second intermediate result matrix.
[0119] In one example, the second intermediate result matrices can be grouped according to the number of matrix registers, for example, the number of groups corresponding to the second intermediate result matrices is equal to the number of matrix registers. The accumulated results corresponding to multiple groups of second intermediate result matrices can be stored in different matrix registers, thereby improving the utilization of the matrix registers. As shown in Figure 10, the memory may include 8 intermediate result matrices and the processor includes 4 matrix registers. Then, the 8 intermediate result matrices can be divided into 4 groups, each group including 2 intermediate result matrices.
[0120] Step S5: The sparse matrix accumulation subunit accumulates the second intermediate result matrices based on the second intermediate index matrix corresponding to each group of second intermediate result matrices to obtain a result matrix.
[0121] After the second intermediate result matrices are grouped, the second intermediate result matrices and corresponding second intermediate index matrices included in each group can be input into the sparse matrix accumulation subunit in sequence by row or column, and the sparse matrix accumulation subunit further accumulates the second intermediate result matrices of each group. The accumulated intermediate results and corresponding position information can be written back to the matrix register again through the channel data write-back device.
[0122] In one possible scenario, in step S3, some intermediate results and corresponding position information are still stored in the matrix register because some matrix registers do not reach the storage space threshold in step S3. In this case, the sparse matrix accumulation subunit can accumulate the intermediate results remaining in the matrix register and the intermediate results included in the second intermediate result matrix in the corresponding group.
[0123] During the process of the channel data write-back device storing the intermediate results and corresponding position information in the matrix register again, if the matrix register reaches the storage space threshold, the intermediate results and corresponding position information in the matrix register can be stored in the memory in the form of a matrix again through the channel data write-memory device. The matrix corresponding to the intermediate results stored in the memory in step S5 can still be referred to as the second intermediate result matrix, and similarly, the matrix corresponding to the position information stored in the memory can still be referred to as the second intermediate index matrix.
[0124] After completing the accumulation of the plurality of sets of second intermediate result matrices, a plurality of second intermediate result matrices and corresponding intermediate index matrices may be obtained in the memory again. Steps S4 to S5 may be performed again on the plurality of second intermediate result matrices and corresponding intermediate index matrices obtained again until no intermediate result corresponding to the same position information exists among the intermediate results accumulated by the sparse matrix accumulation subunit.
[0125] In one example, after the accumulation of the second intermediate result matrix after each grouping is completed, the intermediate results stored in the matrix register and the corresponding position information can be stored in the memory in the form of a matrix, and the number of elements included in the second intermediate result matrix stored again in the memory can be counted. If the number of elements counted twice in a row is the same, it can be considered that the accumulation of the intermediate results is completed. Then, based on the second intermediate result matrix and the corresponding second intermediate index matrix stored in the current memory, the result matrix of the matrix multiplication operation performed on the first sparse matrix and the second sparse matrix can be restored.
[0126] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and the embodiments of the present application will not be described in detail one by one.
[0127] The present application also provides a processor, as shown in FIG5 , which includes a sparse matrix processing unit and a sparse matrix operation unit. The processor may perform the sparse matrix operation method in the above embodiment, which may include:
[0128] A sparse matrix processing unit is used to obtain a first numerical matrix and a first index matrix corresponding to the first sparse matrix, and a second numerical matrix and a second index matrix corresponding to the second sparse matrix, wherein the first numerical matrix is obtained by rearranging the non-zero elements in each column of the first sparse matrix, and the first index matrix is used to indicate the row numbers of the non-zero elements in the first numerical matrix in the first sparse matrix; the second numerical matrix is obtained by rearranging the non-zero elements in each row of the second sparse matrix, and the second index matrix is used to indicate the column numbers of the non-zero elements in the second numerical matrix in the second sparse matrix.
[0129] The sparse matrix processing unit is used to input the non-zero elements of each column in the first numerical matrix and the row number indicated by each column in the first index matrix, the non-zero elements of each row in the second numerical matrix and the column number indicated by each row in the second index matrix into the sparse matrix operation unit.
[0130] The sparse matrix operation unit is used to perform an outer product operation on the non-zero elements in each column of the first numerical matrix and the non-zero elements in each row of the second numerical matrix based on the row number indicated by each column in the first index matrix and the column number indicated by each row in the second index matrix, so as to obtain a result matrix of the matrix multiplication operation performed on the first sparse matrix and the second sparse matrix.
[0131] In one implementable manner, the first numerical matrix includes multiple first numerical partitions, which are obtained by dividing the first sparse matrix into multiple first sparse matrix regions by rows, and then rearranging the non-zero elements in each column of each first sparse matrix region. The first index matrix includes a first index partition corresponding to each first numerical partition.
[0132] The second numerical matrix includes multiple second numerical partitions, which are obtained by dividing the second sparse matrix into multiple second sparse matrix areas by columns, and then rearranging the non-zero elements in each row of each second sparse matrix area. The second index matrix includes second index partitions corresponding to each second numerical partition.
[0133] In one implementable manner, each first numerical partition includes multiple first numerical blocks of size n×m, each first index partition includes multiple first index blocks of size n×m, each second numerical partition includes multiple second numerical blocks of size m×n, and each second index partition includes multiple second index blocks of size m×n.
[0134] In one implementable manner, the sparse matrix processing unit is used to input, in order of calculation, each column element in the non-zero numerical block in each column of the first numerical block in each first numerical partition, each row element in the non-zero numerical block in each row of the second numerical block in each second numerical partition, the row number indicated by each column of the non-zero index block in each column of the first index block in each first index partition, and the column number indicated by each column of the non-zero index block in each row of the second index block in each second index partition into the sparse matrix operation unit.
[0135] In one implementable manner, the sparse matrix operation unit includes a sparse matrix outer product subunit and a sparse matrix accumulation subunit.
[0136] The sparse matrix outer product subunit is used to perform outer product operations on each column element of the first numerical block and each row element of the second numerical block input each time to obtain a first intermediate result matrix.
[0137] The sparse matrix outer product subunit is used to combine the row number indicated by each column of the first index block input each time and the column number indicated by each column of the second index block to form a first intermediate index matrix corresponding to the first intermediate result matrix. The first intermediate index matrix is used to indicate the position information of each element in the first intermediate result matrix in the result matrix.
[0138] The sparse matrix accumulation subunit is used to accumulate multiple first intermediate result matrices based on multiple first intermediate index matrices to obtain a result matrix of performing matrix multiplication operation on the first sparse matrix and the second sparse matrix.
[0139] In one possible implementation, the processor includes a plurality of matrix registers.
[0140] For the first first intermediate result matrix and the corresponding first intermediate index matrix output by the sparse matrix outer product subunit, the sparse matrix processing unit is used to store the intermediate results of each row in the first intermediate result matrix and the position information of each row corresponding to the first intermediate index matrix in multiple matrix registers respectively.
[0141] For the first intermediate result matrix and the corresponding first intermediate index matrix output by the sparse matrix accumulation subunit, the sparse matrix accumulation subunit is used to accumulate the intermediate results of each row in the first intermediate result matrix and the intermediate results stored in the multiple matrix registers based on the position information of each row in the first intermediate index matrix and the position information stored in the multiple matrix registers, and store the accumulated intermediate results and the corresponding position information back to the multiple matrix registers.
[0142] The sparse matrix accumulation subunit is used to determine a result matrix based on intermediate results and position information stored in the plurality of matrix registers.
[0143] In one feasible manner, when the storage space of the matrix register reaches the storage space threshold, the sparse matrix processing unit is used to store the intermediate results and corresponding position information stored in the matrix register in the form of a matrix to the memory to obtain a second intermediate result matrix and a second intermediate index matrix.
[0144] The sparse matrix processing unit is used to group the second intermediate result matrices according to the number of matrix registers to obtain multiple groups of second intermediate result matrices.
[0145] The sparse matrix accumulation subunit is used to accumulate the second intermediate result matrix based on the second intermediate index matrix corresponding to each group of second intermediate result matrices to obtain a result matrix.
[0146] The processor provided in the embodiment of the present application can be used to perform the operation method of the sparse matrix provided in the embodiment of the present application. For the specific execution process, the introduction content of the operation method of the sparse matrix in the above embodiment can be seen, which will not be repeated here. The processor provided in the embodiment of the present application is used to perform sparse multiplication operation. The processor can obtain the first numerical matrix and the first index matrix corresponding to the first sparse matrix to be subjected to matrix operation, the second numerical matrix and the second index matrix corresponding to the second sparse matrix, and the row number of the non-zero elements in the first numerical matrix indicated by the first index matrix in the first sparse matrix and the column number of the non-zero elements in the second numerical matrix indicated by the second index matrix in the second sparse matrix. The outer product operation is performed on the non-zero elements in the first sparse matrix and the second sparse matrix, and then the result matrix of the matrix multiplication performed by the first sparse matrix and the second sparse matrix is obtained. In this way, a large number of zero elements in the first sparse matrix and the second sparse matrix can be avoided from participating in the operation process of matrix multiplication, thereby improving the operation efficiency of performing sparse matrix multiplication operation.
[0147] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the sparse matrix operation method provided in the present application.
[0148] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform the sparse matrix operation method provided in the embodiment of the present application.
[0149] In this application, the words such as term "first", "second" are used to distinguish the same items or similar items with basically the same effect and function. It should be understood that there is no logical or temporal dependency between "first" and "second", nor is quantity and execution order limited. It should also be understood that although the following description uses the terms first, second, etc. to describe various elements, these elements should not be restricted by the terms. These terms are only used to distinguish an element from another element. For example, without departing from the scope of various examples, the first numerical matrix can be referred to as the second numerical matrix, and similarly, the second numerical matrix can be referred to as the first numerical matrix. The first numerical matrix and the second numerical matrix can both be collectively referred to as numerical matrices, and in some cases, can be separate and different numerical matrices.
[0150] The term "at least one" in this application means one or more, and the term "plurality" in this application means two or more.
[0151] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A sparse matrix operation method, characterized in that: The method is executed by a processor, the processor includes a sparse matrix processing unit and a sparse matrix operation unit, and the method includes: The sparse matrix processing unit obtains a first numerical matrix and a first index matrix corresponding to the first sparse matrix, and a second numerical matrix and a second index matrix corresponding to the second sparse matrix, wherein the first numerical matrix is obtained by rearranging the non-zero elements in each column of the first sparse matrix, and the first index matrix is used to indicate the row numbers of the non-zero elements in the first numerical matrix in the first sparse matrix, and the second numerical matrix is obtained by rearranging the non-zero elements in each row of the second sparse matrix, and the second index matrix is used to indicate the column numbers of the non-zero elements in the second numerical matrix in the second sparse matrix; The sparse matrix processing unit inputs the non-zero elements of each column in the first numerical matrix and the row number indicated by each column in the first index matrix, and the non-zero elements of each row in the second numerical matrix and the column number indicated by each row in the second index matrix to the sparse matrix operation unit; The sparse matrix operation unit performs outer product operations on the non-zero elements of each column in the first numerical matrix and the non-zero elements of each row in the second numerical matrix based on the row number indicated by each column in the first index matrix and the column number indicated by each row in the second index matrix to obtain a result matrix of matrix multiplication operations performed on the first sparse matrix and the second sparse matrix.
2. The method according to claim 1, characterized in that The first numerical matrix includes a plurality of first numerical partitions, wherein the plurality of first numerical partitions are obtained by dividing the first sparse matrix into a plurality of first sparse matrix regions by rows, and then rearranging non-zero elements in each column of each first sparse matrix region, and the first index matrix includes a first index partition corresponding to each first numerical partition; The second numerical matrix includes multiple second numerical partitions, which are obtained by dividing the second sparse matrix into multiple second sparse matrix areas by columns, and then rearranging the non-zero elements in each row of each second sparse matrix area. The second index matrix includes second index partitions corresponding to each second numerical partition.
3. The method according to claim 2, characterized in that Each of the first numerical partitions includes a plurality of first numerical blocks of size n×m, each of the first index partitions includes a plurality of first index blocks of size n×m, each of the second numerical partitions includes a plurality of second numerical blocks of size m×n, and each of the second index partitions includes a plurality of second index blocks of size m×n.
4. The method according to claim 3, characterized in that The sparse matrix processing unit inputs the non-zero elements of each column in the first numerical matrix and the row number indicated by each column in the first index matrix, and the non-zero elements of each row in the second numerical matrix and the column number indicated by each row in the second index matrix to the sparse matrix operation unit, including: The sparse matrix processing unit inputs, in order of calculation, each column element in the non-zero numerical block in each column of the first numerical block in each first numerical partition, each row element in the non-zero numerical block in each row of the second numerical block in each second numerical partition, the row number indicated by each column of the non-zero index block in each column of the first index block in each first index partition, and the column number indicated by each column of the non-zero index block in each row of the second index block in each second index partition into the sparse matrix operation unit.
5. The method according to claim 4, characterized in that: The sparse matrix operation unit includes a sparse matrix outer product subunit and a sparse matrix accumulation subunit. The sparse matrix operation unit performs an outer product operation on non-zero elements in each column of the first numerical matrix and non-zero elements in each row of the second numerical matrix based on the row number indicated by each column in the first index matrix and the column number indicated by each row in the second index matrix to obtain a result matrix of the matrix multiplication operation performed on the first sparse matrix and the second sparse matrix, including: The sparse matrix outer product subunit performs an outer product operation on each column element of the first numerical block and each row element of the second numerical block input each time to obtain a first intermediate result matrix; The sparse matrix outer product subunit combines the row number indicated by each column of the first index block input each time and the column number indicated by each column of the second index block to form a first intermediate index matrix corresponding to the first intermediate result matrix, wherein the first intermediate index matrix is used to indicate the position information of each element in the first intermediate result matrix in the result matrix; The sparse matrix accumulation subunit accumulates multiple first intermediate result matrices based on multiple first intermediate index matrices to obtain a result matrix of performing matrix multiplication operation of the first sparse matrix and the second sparse matrix.
6. The method according to claim 5, characterized in that The processor includes a plurality of matrix registers, and the method further includes: For a first first intermediate result matrix output by the sparse matrix outer product subunit and a corresponding first intermediate index matrix, the sparse matrix processing unit stores each row of intermediate results in the first intermediate result matrix and each row position information corresponding to the first intermediate index matrix in the plurality of matrix registers respectively; The sparse matrix accumulation subunit accumulates a plurality of first intermediate result matrices based on a plurality of first intermediate index matrices to obtain a result matrix of matrix multiplication operation performed by the first sparse matrix and the second sparse matrix, including: For the first intermediate result matrix output by the sparse matrix accumulation subunit and the corresponding first intermediate index matrix, the sparse matrix accumulation subunit accumulates the intermediate results of each row in the first intermediate result matrix and the intermediate results respectively stored in the multiple matrix registers based on the position information of each row in the first intermediate index matrix and the position information respectively stored in the multiple matrix registers, and stores the accumulated intermediate results and the corresponding position information back to the multiple matrix registers respectively; The sparse matrix accumulation subunit determines the result matrix based on the intermediate results and position information stored in the plurality of matrix registers.
7. The method according to claim 6, characterized in that The method further comprises: When the storage space of the matrix register reaches the storage space threshold, the sparse matrix processing unit stores the intermediate results and corresponding position information stored in the matrix register into the memory in the form of a matrix to obtain a second intermediate result matrix and a second intermediate index matrix; The sparse matrix processing unit groups the second intermediate result matrix according to the number of the matrix registers to obtain multiple groups of second intermediate result matrices; The sparse matrix accumulation subunit determines the result matrix based on the intermediate results and position information stored in the plurality of matrix registers, including: The sparse matrix accumulation subunit accumulates the second intermediate result matrices based on the second intermediate index matrix corresponding to each group of second intermediate result matrices to obtain the result matrix.
8. A processor, characterized in that: The processor includes a sparse matrix processing unit and a sparse matrix operation unit, and the method includes: The sparse matrix processing unit is used to obtain a first numerical matrix and a first index matrix corresponding to the first sparse matrix, and a second numerical matrix and a second index matrix corresponding to the second sparse matrix, wherein the first numerical matrix is obtained by rearranging the non-zero elements in each column of the first sparse matrix, and the first index matrix is used to indicate the row numbers of the non-zero elements in the first numerical matrix in the first sparse matrix, and the second numerical matrix is obtained by rearranging the non-zero elements in each row of the second sparse matrix, and the second index matrix is used to indicate the column numbers of the non-zero elements in the second numerical matrix in the second sparse matrix; The sparse matrix processing unit is used to input the non-zero elements of each column in the first numerical matrix and the row number indicated by each column in the first index matrix, and the non-zero elements of each row in the second numerical matrix and the column number indicated by each row in the second index matrix to the sparse matrix operation unit; The sparse matrix operation unit is used to perform outer product operations on non-zero elements in each column of the first numerical matrix and non-zero elements in each row of the second numerical matrix based on the row number indicated by each column in the first index matrix and the column number indicated by each row in the second index matrix, so as to obtain a result matrix of matrix multiplication operations performed on the first sparse matrix and the second sparse matrix.
9. The processor according to claim 1, wherein: The first numerical matrix includes a plurality of first numerical partitions, wherein the plurality of first numerical partitions are obtained by dividing the first sparse matrix into a plurality of first sparse matrix regions by rows, and then rearranging non-zero elements in each column of each first sparse matrix region, and the first index matrix includes a first index partition corresponding to each first numerical partition; The second numerical matrix includes multiple second numerical partitions, which are obtained by dividing the second sparse matrix into multiple second sparse matrix areas by columns, and then rearranging the non-zero elements in each row of each second sparse matrix area. The second index matrix includes second index partitions corresponding to each second numerical partition.
10. The processor according to claim 9, characterized in that Each of the first numerical partitions includes a plurality of first numerical blocks of size n×m, each of the first index partitions includes a plurality of first index blocks of size n×m, each of the second numerical partitions includes a plurality of second numerical blocks of size m×n, and each of the second index partitions includes a plurality of second index blocks of size m×n.
11. The processor according to claim 10, characterized in that The sparse matrix processing unit is used to input, in order of calculation, each column element in the non-zero numerical block in each column of the first numerical block in each first numerical partition, each row element in the non-zero numerical block in each row of the second numerical block in each second numerical partition, the row number indicated by each column of the non-zero index block in each column of the first index block in each first index partition, and the column number indicated by each column of the non-zero index block in each row of the second index block in each second index partition into the sparse matrix operation unit.
12. The processor according to claim 11, characterized in that: The sparse matrix operation unit includes a sparse matrix outer product subunit and a sparse matrix accumulation subunit; The sparse matrix outer product subunit is used to perform an outer product operation on each column element of the first numerical block and each row element of the second numerical block input each time to obtain a first intermediate result matrix; The sparse matrix outer product subunit is used to combine the row number indicated by each column of the first index block input each time and the column number indicated by each column of the second index block to form a first intermediate index matrix corresponding to the first intermediate result matrix, wherein the first intermediate index matrix is used to indicate the position information of each element in the first intermediate result matrix in the result matrix; The sparse matrix accumulation subunit is used to accumulate multiple first intermediate result matrices based on multiple first intermediate index matrices to obtain a result matrix of performing matrix multiplication operation on the first sparse matrix and the second sparse matrix.
13. The processor according to claim 12, characterized in that The processor comprises a plurality of matrix registers, For the first first intermediate result matrix output by the sparse matrix outer product subunit and the corresponding first intermediate index matrix, the sparse matrix processing unit is used to store each row of intermediate results in the first intermediate result matrix and each row position information corresponding to the first intermediate index matrix in the plurality of matrix registers respectively; For the first intermediate result matrix output by the sparse matrix accumulation subunit and the corresponding first intermediate index matrix, the sparse matrix accumulation subunit is used to accumulate the intermediate results of each row in the first intermediate result matrix and the intermediate results respectively stored in the multiple matrix registers based on the position information of each row in the first intermediate index matrix and the position information respectively stored in the multiple matrix registers, and store the accumulated intermediate results and the corresponding position information back to the multiple matrix registers respectively; The sparse matrix accumulation subunit is used to determine the result matrix based on the intermediate results and position information stored in the multiple matrix registers.
14. The processor according to claim 13, characterized in that When the storage space of the matrix register reaches the storage space threshold, the sparse matrix processing unit is used to store the intermediate results and corresponding position information stored in the matrix register into the memory in the form of a matrix to obtain a second intermediate result matrix and a second intermediate index matrix; The sparse matrix processing unit is used to group the second intermediate result matrices according to the number of the matrix registers to obtain multiple groups of second intermediate result matrices; The sparse matrix accumulation subunit is used to accumulate the second intermediate result matrices based on the second intermediate index matrix corresponding to each group of second intermediate result matrices to obtain the result matrix.
15. A computing device, characterized in that: The computing device comprises a memory and a processor as described in any one of claims 8 to 14 above, wherein the memory stores at least one instruction, and the processor executes the at least one instruction to execute the method as described in any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program code, and when the computer program code is executed by a computing device, the computing device is caused to perform the method according to any one of claims 1 to 7.
17. A computer program product comprising instructions, characterized in that When the computer program product is executed on a computing device, the computing device is caused to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Variable format, variable sparsity matrix multiplication instruction
CN110580175A
Asynchronous sparse matrix outer product multiplier based on event-driven circuit design
CN115617305A
Sparse matrix compression and data matching method and device
CN115618183A
Sparse matrix multiplication method using GPU
KR101400577B1
Matrix calculation apparatus, method, system, circuit, and device, and chip
US20230342419A1
Cited By
Hardware accelerator facing sparse matrix vector multiplication, equipment and application method
CN121743654A
Sparse tensor processing method and device, electronic equipment and storage medium
CN122242571A