Matrix storage device and method and matrix operation device and method
By storing sparse matrices using row compression format and column compression format, and supporting sparse multiplication operations of sparse matrices, the problems of sparse matrices in the prior art are solved, and efficient sparse matrix calculation is realized.
Patent Information
- Application Number
- CN202510639261.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The existing sparse matrix calculation technology cannot effectively process the matrix of random sparse proportions and features in the real world, resulting in an increase in the error of the calculation result and the direct multiplication between sparse matrices cannot be realized.
The sparse matrix is stored in row compression format and column compression format, and the numerical array, index array and indication array are generated through the processing unit, supporting the sparse multiplication operation of sparse matrix.
It significantly reduces the storage space requirements of sparse matrices, realizes direct multiplication between sparse matrices, and reduces the calculation amount and power consumption.
Smart Images

Figure CN120179976A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly relates to a matrix storage, matrix operation device and method. Background Art
[0002] Matrix multiplication operation is the most frequently involved operation type in scenarios such as artificial intelligence, big data processing, and scientific computing. Sparse matrices widely exist in datasets of various types of applications. Accelerating the operation of sparse matrices can effectively reduce the power consumption of hardware and improve the operation performance. Existing sparse matrix operations can only achieve a fixed proportion of sparse calculations or calculations of a single input sparse matrix.
[0003] Fixed proportion sparse matrix operation requires that the sparsification ratio of the input sparse matrix is fixed and meets certain constraints. For example, a GPU sparse computing unit on the market requires that at least 2 out of every 4 adjacent elements in the input sparse matrix have a value of 0, that is, a sparsity rate of 50%. However, the data sparsity ratio and sparsity characteristics in the real world are random. Fixing the sparse ratio and sparsification characteristics usually causes an increase in the error of matrix representation, thereby increasing the error of the calculation result.
[0004] Existing sparse matrix calculations also have a solution that only supports 1 of the input matrices to be a sparse matrix, that is, to perform the operation of "sparse matrix" and "dense matrix", and still cannot achieve the operation of "sparse matrix" and "sparse matrix". Therefore, the calculation performance is limited, and there is also the problem of energy consumption waste.
[0005] This section aims to provide background or context for the embodiments of the present application stated in the claims. The descriptions herein are not admitted to be prior art just because they are included in this section. Summary of the Invention
[0006] In order to solve at least one of the above problems existing in the prior art, embodiments of the present application provide a matrix storage, matrix operation device and method.
[0007] Embodiments of the present application provide a matrix storage device, including:
[0008] A processing unit, configured to traverse a sparse matrix in row-major order, and respectively generate a numerical array and an index array of the sparse matrix according to the values and column indexes of all non-zero elements in the sparse matrix; when traversing each row of the sparse matrix, record the number of non-zero elements in each row of the sparse matrix, and generate an indication array of the sparse matrix according to the number of non-zero elements in each row of the sparse matrix; A storage unit, coupled to the processing unit, is configured to store the numerical array, index array and indication array of the sparse matrix, wherein the numerical array, index array and indication array are the row-compressed format representation of the sparse matrix;
[0009] Alternatively,
[0010] a processing unit, configured to traverse a sparse matrix in column-major order, generate a numerical array and an index array of the sparse matrix respectively according to values and row indices of all non-zero elements in the sparse matrix; record the number of non-zero elements in each column of the sparse matrix when traversing each column of the sparse matrix, and generate an indication array of the sparse matrix according to the number of non-zero elements in each column of the sparse matrix; a storage unit, coupled to the processing unit, configured to store the numerical array, the index array and the indication array of the sparse matrix, wherein the numerical array, the index array and the indication array are representations of a column-compressed format of the sparse matrix.
[0011] An embodiment of the present application provides a matrix operation device, including: a reading module, coupled to a first storage unit, where the first storage unit stores a first sparse matrix in row-compressed format and a second sparse matrix in column-compressed format, and the reading module is configured to read the first sparse matrix in row-compressed format and the second sparse matrix in column-compressed format from the first storage unit, wherein the first sparse matrix in row-compressed format is represented by a first numerical array, a first index array and a first indication array, the first numerical array stores values of non-zero elements of the first sparse matrix in row-major order, the first index array stores column indices of non-zero elements of the first sparse matrix in row-major order, the first indication array stores the number of non-zero elements in each row of the first sparse matrix in row-major order, the second sparse matrix in column-compressed format is represented by a second numerical array, a second index array and a second indication array, the second numerical array stores values of non-zero elements of the second sparse matrix in column-major order, the second index array stores row indices of non-zero elements of the second sparse matrix in column-major order, the second indication array stores the number of non-zero elements in each column of the second sparse matrix in column-major order; an operation module, coupled to the reading module, configured to perform a sparse multiplication operation on the first sparse matrix and the second sparse matrix based on the first sparse matrix in row-compressed format and the second sparse matrix in column-compressed format to obtain an operation result matrix in row-compressed format or column-compressed format.
[0012] In some embodiments, the reading module reads non-zero elements in the i-th row of the first sparse matrix, column indices of each non-zero element in the i-th row and the number of non-zero elements in the i-th row from the first storage unit each time, and reads non-zero elements in the j-th column of the second sparse matrix, row indices of each non-zero element in the j-th column and the number of non-zero elements in the j-th column.
[0013] In some embodiments, the operation module includes a control unit and a multiplier; wherein, the control unit, coupled to the reading module, is configured to determine, based on the non-zero elements in the i-th row of the first sparse matrix and the column indices of the non-zero elements in the i-th row, and the non-zero elements in the j-th column of the second sparse matrix and the row indices of the non-zero elements in the j-th column, target element pairs that need to perform dot product operations among the non-zero elements in the i-th row of the first sparse matrix and the non-zero elements in the j-th column of the second sparse matrix; the multiplier, coupled to the control unit and the reading module, is configured to perform dot product operations on each of the target element pairs under the control of the control unit to obtain the dot product operation results of each of the target element pairs.
[0014] In some embodiments, the operation module further includes a first comparator and an accumulator; wherein, the first comparator, coupled to the reading module and the control unit, is configured to compare the number of non-zero elements in the i-th row of the first sparse matrix and the number of non-zero elements in the j-th column of the second sparse matrix, and send the smaller value to the control unit; the control unit is further configured to determine an accumulation count based on the smaller value; the accumulator, coupled to the control unit and the multiplier, is configured to accumulate the dot product operation results according to the accumulation count to obtain the element in the i-th row and j-th column of the operation result matrix.
[0015] In some embodiments, a second storage unit coupled to the accumulator is further included, and the second storage unit is configured to sequentially store the non-zero elements output by the accumulator until the numerical array of the operation result matrix is obtained.
[0016] In some embodiments, the operation module further includes a counter; wherein,
[0017] the counter, coupled to the accumulator, is configured to count the number of non-zero elements in each row of the operation result matrix output by the accumulator; the second storage unit is further coupled to the counter and is configured to sequentially store the number of non-zero elements in each row of the operation result matrix output by the counter until the indication array of the operation result matrix is obtained; or, the counter, coupled to the accumulator, is configured to count the number of non-zero elements in each column of the operation result matrix output by the accumulator; the second storage unit is further coupled to the counter and is configured to sequentially store the number of non-zero elements in each column of the operation result matrix output by the counter until the indication array of the operation result matrix is obtained.
[0018] In some embodiments, the operation module further includes a second comparator; wherein, the second comparator is coupled to the reading module and is configured to compare the column indices of each non-zero element in each row of the first sparse matrix with the row indices of each non-zero element in each column of the second sparse matrix, and send the comparison result to the control unit; the control unit is configured to determine the column indices of each non-zero element in the operation result matrix according to the comparison result; the second storage unit is further coupled to the control unit, and the second storage unit is further configured to sequentially store the column indices of each non-zero element until the index array of the operation result matrix is obtained; alternatively, the control unit is further configured to determine the row indices of each non-zero element in the operation result matrix according to the comparison result; the second storage unit is further coupled to the control unit, and the second storage unit is further configured to sequentially store the row indices of each non-zero element until the index array of the operation result matrix is obtained.
[0019] In some embodiments, the reading module includes:
[0020] A first loading unit, coupled to the first storage unit, is configured to read the non-zero elements in the i-th row of the first sparse matrix and the column indices of the non-zero elements in the i-th row from the first storage unit, and store the non-zero elements in the i-th row of the first sparse matrix and the column indices of the non-zero elements in the i-th row into a first temporary storage unit;
[0021] A first temporary storage unit, coupled to the first loading unit, the second comparator, the control unit, and the multiplier, is configured to temporarily store the non-zero elements in the i-th row of the first sparse matrix and the column indices of the non-zero elements in the i-th row;
[0022] A second loading unit, coupled to the first storage unit, is configured to read the non-zero elements in the j-th column of the second sparse matrix and the row indices of the non-zero elements in each column from the first storage unit, and store the non-zero elements in the j-th column of the second sparse matrix and the row indices of the non-zero elements in the j-th column into a second temporary storage unit;
[0023] A second temporary storage unit, coupled to the second loading unit, the second comparator, the control unit, and the multiplier, is configured to temporarily store the non-zero elements in the j-th column of the second sparse matrix and the row indices of the non-zero elements in the j-th column;
[0024] A third loading unit, coupled to the first storage unit, is configured to read the number of non-zero elements in the i-th row of the first sparse matrix from the first storage unit, and store the number of non-zero elements in the i-th row of the first sparse matrix into a third temporary storage unit;
[0025] A third temporary storage unit, coupled to the third loading unit, the first comparator, and the control unit, is configured to temporarily store the number of non-zero elements in the i-th row of the first sparse matrix;
[0026] A fourth loading unit, coupled to the first storage unit, is configured to read the number of non-zero elements in the j-th column of the second sparse matrix from the first storage unit, and store the number of non-zero elements in the j-th column of the second sparse matrix into a fourth temporary storage unit;
[0027] A fourth temporary storage unit, coupled to the fourth loading unit, the first comparator, and the control unit, is configured to temporarily store the number of non-zero elements in the j-th column of the second sparse matrix.
[0028] An embodiment of the present application further provides a matrix storage method, including:
[0029] The processing unit traverses the sparse matrix in row-major order, and respectively generates a numerical array and an index array of the sparse matrix according to the values and column indexes of all non-zero elements in the sparse matrix; when traversing each row of the sparse matrix, record the number of non-zero elements in each row of the sparse matrix, and generate an indication array of the sparse matrix according to the number of non-zero elements in each row of the sparse matrix;
[0030] The storage unit stores the numerical array, the index array, and the indication array of the sparse matrix, where the numerical array, the index array, and the indication array are representations in a row-compressed format of the sparse matrix;
[0031] Or,
[0032] The processing unit traverses the sparse matrix in column-major order, and respectively generates a numerical array and an index array of the sparse matrix according to the values and row indexes of all non-zero elements in the sparse matrix; when traversing each column of the sparse matrix, record the number of non-zero elements in each column of the sparse matrix, and generate an indication array of the sparse matrix according to the number of non-zero elements in each column of the sparse matrix;
[0033] The storage unit stores the numerical array, the index array, and the indication array of the sparse matrix, where the numerical array, the index array, and the indication array are representations in a column-compressed format of the sparse matrix.
[0034] The embodiment of the present application also provides a matrix operation method, including: a reading module reads a first sparse matrix in row compressed format and a second sparse matrix in column compressed format stored in a first storage unit respectively, wherein the first storage unit stores the first sparse matrix in row compressed format and the second sparse matrix in column compressed format, the first sparse matrix in row compressed format is represented by a first numerical array, a first index array and a first indication array, the first numerical array stores the values of non-zero elements of the first sparse matrix in row-major order, the first index array stores the column indices of non-zero elements of the first sparse matrix in row-major order, the first indication array stores the number of non-zero elements in each row of the first sparse matrix in row-major order, the second sparse matrix in column compressed format is represented by a second numerical array, a second index array and a second indication array, the second numerical array stores the values of non-zero elements of the second sparse matrix in column-major order, the second index array stores the row indices of non-zero elements of the second sparse matrix in column-major order, the second indication array stores the number of non-zero elements in each column of the second sparse matrix in column-major order; an operation module performs a sparse multiplication operation on the first sparse matrix and the second sparse matrix based on the first sparse matrix in row compressed format and the second sparse matrix in column compressed format, to obtain an operation result matrix in row compressed format or column compressed format.
[0035] In some embodiments, the reading module reads the non-zero elements in the i-th row of the first sparse matrix, the column indices of the non-zero elements in the i-th row, and the number of non-zero elements in the i-th row from the first storage unit each time, and reads the non-zero elements in the j-th column of the second sparse matrix, the row indices of the non-zero elements in the j-th column, and the number of non-zero elements in the j-th column.
[0036] In some embodiments, a control unit of the operation module determines target element pairs that need to perform dot product operations among the non-zero elements in the i-th row of the first sparse matrix and the non-zero elements in the j-th column of the second sparse matrix based on the non-zero elements in the i-th row of the first sparse matrix and the column indices of the non-zero elements in the i-th row, and the non-zero elements in the j-th column of the second sparse matrix and the row indices of the non-zero elements in the j-th column; a multiplier of the operation module performs dot product operations on each of the target element pairs under the control of the control unit to obtain dot product operation results of each of the target element pairs.
[0037] In some embodiments, a first comparator of the operation module compares the number of non-zero elements in the i-th row of the first sparse matrix with the number of non-zero elements in the j-th column of the second sparse matrix, and sends the smaller value to the control unit; the control unit determines an accumulation count based on the smaller value; an accumulator of the operation module accumulates the dot product operation results according to the accumulation count to obtain the element at the i-th row and j-th column of the operation result matrix.
[0038] In some embodiments, the method further includes: a second storage unit sequentially stores the non-zero elements output by the accumulator until a numerical array of the operation result matrix is obtained.
[0039] In some embodiments, a counter of the operation module counts the number of non-zero elements in each row of the operation result matrix output by the accumulator; the second storage unit sequentially stores the number of non-zero elements in each row of the operation result matrix output by the counter until an indication array of the operation result matrix is obtained;
[0040] Or
[0041] the counter counts the number of non-zero elements in each column of the operation result matrix output by the accumulator; the second storage unit sequentially stores the number of non-zero elements in each column of the operation result matrix output by the counter until an indication array of the operation result matrix is obtained.
[0042] In some embodiments, a second comparator of the operation module; compares the column index of each non-zero element in each row of the first sparse matrix with the row index of each non-zero element in each column of the second sparse matrix, and sends the comparison result to the control unit; the control unit further determines the column index of each non-zero element in the operation result matrix according to the comparison result, and the second storage unit further sequentially stores the column index of each non-zero element until an index array of the operation result matrix is obtained; or, the control unit further determines the row index of each non-zero element in the operation result matrix according to the comparison result, and the second storage unit further sequentially stores the row index of each non-zero element until an index array of the operation result matrix is obtained.
[0043] In some embodiments, a first loading unit of the reading module reads the non-zero elements in the i-th row of the first sparse matrix and the column indices of the non-zero elements in the i-th row from the first storage unit, and stores the non-zero elements in the i-th row of the first sparse matrix and the column indices of the non-zero elements in the i-th row into a first temporary storage unit;
[0044] The first temporary storage unit of the reading module temporarily stores the non-zero elements in the i-th row of the first sparse matrix and the column indices of the non-zero elements in the i-th row;
[0045] The second loading unit of the reading module reads the non-zero elements of the j-th column of the second sparse matrix and the row indices of the non-zero elements of each column from the first storage unit, and stores the non-zero elements of the j-th column of the second sparse matrix and the row indices of the non-zero elements of the j-th column into the second temporary storage unit;
[0046] The second temporary storage unit of the reading module temporarily stores the non-zero elements of the j-th column of the second sparse matrix and the row indices of the non-zero elements of the j-th column;
[0047] The third loading unit of the reading module reads the number of non-zero elements of the i-th row of the first sparse matrix from the first storage unit, and stores the number of non-zero elements of the i-th row of the first sparse matrix into the third temporary storage unit;
[0048] The third temporary storage unit of the reading module temporarily stores the number of non-zero elements of the i-th row of the first sparse matrix;
[0049] The fourth loading unit of the reading module reads the number of non-zero elements of the j-th column of the second sparse matrix from the first storage unit, and stores the number of non-zero elements of the j-th column of the second sparse matrix into the fourth temporary storage unit;
[0050] The fourth temporary storage unit of the reading module temporarily stores the number of non-zero elements of the j-th column of the second sparse matrix.
[0051] The embodiment of the present application also provides an electronic device, which includes the matrix storage device described in the above embodiment and / or the matrix operation device described in any of the above embodiments.
[0052] The matrix storage, matrix operation device and method proposed by the present application store sparse matrices in a new row compression format or column compression format, which can significantly reduce the storage space for storing sparse matrices, support direct sparse multiplication operations on sparse matrices in row compression format and column compression format, and can directly obtain the operation result matrix in row compression format or column compression format. During the operation process, there is no need to decompress the compression format, realizing the multiplication operation of two input sparse matrices that are both in compression format, reducing the amount of operation, operation power consumption and storage capacity requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:
[0054] Figure 1 This is a schematic structural diagram of a matrix storage device provided by the present application.
[0055] Figure 2 This is a schematic flowchart of a matrix storage method provided by an embodiment of the present application.
[0056] Figure 3 This is a schematic flowchart of another matrix storage method provided by an embodiment of the present application.
[0057] Figure 4 This is a schematic structural diagram of a matrix operation device provided by an embodiment of the present application.
[0058] Figure 5 This is a schematic structural diagram of a matrix operation device provided by an embodiment of the present application.
[0059] Figure 6 This is a schematic structural diagram of a matrix operation device provided by an embodiment of the present application.
[0060] Figure 7 This is a schematic structural diagram of a matrix operation device provided by an embodiment of the present application.
[0061] Figure 8 This is a schematic structural diagram of a matrix operation device provided by an embodiment of the present application.
[0062] Figure 9 This is a schematic structural diagram of a matrix operation device provided by an embodiment of the present application.
[0063] Figure 10 This is a schematic structural diagram of a matrix operation device provided by an embodiment of the present application.
[0064] Figure 11 This is a schematic structural diagram of a matrix operation device provided by an embodiment of the present application.
[0065] Figure 12 This is a schematic flowchart of a matrix operation method provided by an embodiment of the present application.
[0066] Figure 13 This is a schematic physical structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0067] Now, reference will be made in detail to the exemplary embodiments of the present invention, and examples of the exemplary embodiments are illustrated in the accompanying drawings. As long as possible, the same component symbols are used in the drawings and the description to represent the same or similar parts.
[0068] As used throughout the specification (including the claims) of this case, the term "coupled (or connected)" may refer to any direct or indirect connection means. For example, if it is described in the text that a first device is coupled (or connected) to a second device, it should be interpreted that the first device can be directly connected to the second device, or the first device can be indirectly connected to the second device through other devices or certain connection means. The terms "first", "second", etc. mentioned throughout the specification (including the claims) of this case are used to name components (elements), rather than to limit the upper or lower limit of the number of components, nor to limit the order of components. Additionally, wherever possible, components / elements / steps with the same reference numerals in the drawings and embodiments represent the same or similar parts. Components / elements / steps with the same reference numerals or the same terms used in different embodiments can be cross-referred to the relevant descriptions.
[0069] Figure 1 is a schematic structural diagram of a matrix storage device provided by this application, as Figure 1 shown, this application provides a matrix storage device 100, including:
[0070] A processing unit 11, configured to traverse a sparse matrix in row-major order, and respectively generate a numerical array and an index array of the sparse matrix according to the values and column indices of all non-zero elements in the sparse matrix; when traversing each row of the sparse matrix, record the number of non-zero elements in each row of the sparse matrix, and generate an indication array of the sparse matrix according to the number of non-zero elements in each row of the sparse matrix;
[0071] A storage unit 12, coupled to the processing unit 11, configured to store the numerical array, index array, and indication array of the sparse matrix, wherein the numerical array, index array, and indication array are representations in row-compressed format of the sparse matrix;
[0072] Or,
[0073] A processing unit 11, configured to traverse a sparse matrix in column-major order, and respectively generate a numerical array and an index array of the sparse matrix according to the values and row indices of all non-zero elements in the sparse matrix; when traversing each column of the sparse matrix, record the number of non-zero elements in each column of the sparse matrix, and generate an indication array of the sparse matrix according to the number of non-zero elements in each column of the sparse matrix;
[0074] A storage unit 12, coupled to the processing unit 11, configured to store the numerical array, index array, and indication array of the sparse matrix, wherein the numerical array, index array, and indication array are representations in column-compressed format of the sparse matrix.
[0075] Specifically, this embodiment proposes a compression format for sparse matrices, representing the sparse matrix in the compression format using a numerical array, an index array, and an indicator array; wherein, the compression format includes a row compression format and a column compression format; wherein, in the row compression format of the sparse matrix, the numerical array stores all non-zero elements of the sparse matrix in row order, the index array stores the column indices where the non-zero elements are located in row order, and the indicator array stores the number of non-zero elements in each row in row order.
[0076] Similarly, in the column compression format of the sparse matrix, the numerical array stores all non-zero elements of the matrix in column order, the index array stores the row indices where the non-zero elements are located in column order, and the indicator array stores the number of non-zero elements in each column in column order.
[0077] Since sparse matrices usually contain multiple zero elements, retaining only the non-zero elements of the sparse matrix through the above row compression format and column compression format can effectively reduce memory occupancy.
[0078] For example, the following is a 3x3 sparse matrix.
[0079]
[0080] Representing the above sparse matrix in the row compression format, we get:
[0081] RN = [2, 2, 2];
[0082] V = [1, 2, 5, 3, 4, 2];
[0083] CI = [1, 2, 0, 1, 0, 2].
[0084] Among them, in the row compression format,
[0085] RN: represents the number of non-zero elements in each row;
[0086] V: stores all non-zero elements of the matrix in row order;
[0087] CI: stores the column indices where the non-zero elements are located in row order.
[0088] Representing the above sparse matrix in the column compression format, we get:
[0089] CN = [2, 2, 2];
[0090] V = [5, 4, 1, 3, 2, 2];
[0091] RI = [1, 2, 0, 1, 0, 3].
[0092] Among them, in the column compression format,
[0093] CN: represents the number of non-zero elements in each column;
[0094] V: stores all non-zero elements of the matrix in column order;
[0095] RI: stores the row indices where non-zero elements are located in column order.
[0096] It can be seen that storing the sparse matrix in the row compression format or column compression format proposed in this embodiment can significantly reduce the storage space of the sparse matrix, and the sparse matrices in row compression format and column compression format can directly perform multiplication operations.
[0097] Figure 2 is a flowchart of a matrix storage method provided by an embodiment of the present application. As Figure 2 shown, an embodiment of the present application provides a matrix storage method, including:
[0098] S11. The processing unit traverses the sparse matrix in row-major order, and respectively generates a numerical array and an index array of the sparse matrix according to the values and column indices of all non-zero elements in the sparse matrix; when traversing each row of the sparse matrix, record the number of non-zero elements in each row of the sparse matrix, and generate an indication array of the sparse matrix according to the number of non-zero elements in each row of the sparse matrix;
[0099] S12. The storage unit stores the numerical array, index array, and indication array of the sparse matrix, where the numerical array, index array, and indication array are the row compression format representation of the sparse matrix.
[0100] The specific process of the matrix storage method provided by the embodiment of the present application can be seen in the introduction content of the matrix storage device in the above embodiment, and will not be elaborated here. Using the matrix storage method provided by the embodiment of the present application can significantly reduce the storage space of the sparse matrix, and the obtained sparse matrices in row compression format and column compression format support direct multiplication operations.
[0101] Figure 3 is a flowchart of another matrix storage method provided by an embodiment of the present application. As Figure 3 shown, an embodiment of the present application provides a matrix storage method, including:
[0102] S21. The processing unit traverses the sparse matrix in column-major order, and respectively generates a numerical array and an index array of the sparse matrix according to the values and row indices of all non-zero elements in the sparse matrix; when traversing each column of the sparse matrix, record the number of non-zero elements in each column of the sparse matrix, and generate an indication array of the sparse matrix according to the number of non-zero elements in each column of the sparse matrix;
[0103] S22. The storage unit stores the numerical array, index array, and indication array of the sparse matrix, where the numerical array, index array, and indication array are representations in the column-compressed format of the sparse matrix.
[0104] For the specific process of the matrix storage method provided in the embodiments of this application, refer to the description of the matrix storage device in the above embodiments, which will not be elaborated here. By using the matrix storage method provided in the embodiments of this application, the storage space for storing the sparse matrix can be significantly reduced, and the row-compressed format and column-compressed format sparse matrices obtained support direct multiplication operations.
[0105] Figure 4 It is a schematic structural diagram of a matrix operation device provided in the embodiments of this application. As Figure 4 shown, the embodiments of this application provide a matrix operation device 200, including:
[0106] A reading module 21, coupled to the first storage unit. The first storage unit stores a first sparse matrix in row-compressed format and a second sparse matrix in column-compressed format. The reading module 21 is configured to read the first sparse matrix in row-compressed format and the second sparse matrix in column-compressed format from the first storage unit. Among them, the first sparse matrix in row-compressed format is represented by a first numerical array, a first index array, and a first indication array. The first numerical array stores the values of the non-zero elements of the first sparse matrix in row-major order. The first index array stores the column indices of the non-zero elements of the first sparse matrix in row-major order. The first indication array stores the number of non-zero elements in each row of the first sparse matrix in row-major order. The second sparse matrix in column-compressed format is represented by a second numerical array, a second index array, and a second indication array. The second numerical array stores the values of the non-zero elements of the second sparse matrix in column-major order. The second index array stores the row indices of the non-zero elements of the second sparse matrix in column-major order. The second indication array stores the number of non-zero elements in each column of the second sparse matrix in column-major order;
[0107] An operation module 22, coupled to the reading module 21, is configured to perform a sparse multiplication operation on the first sparse matrix and the second sparse matrix based on the first sparse matrix in row-compressed format and the second sparse matrix in column-compressed format, to obtain an operation result matrix in row-compressed format or column-compressed format.
[0108] Specifically, the first storage unit and the storage unit in the above matrix storage device may be the same storage unit. Of course, they may also not be the same storage unit, and this embodiment does not limit this. The reading module 21 directly reads the first sparse matrix in row-compressed format and the second sparse matrix in column-compressed format, and sends the first sparse matrix in row-compressed format and the second sparse matrix in column-compressed format to the operation module 22. The operation module 22 directly performs a sparse multiplication operation on the first sparse matrix in row-compressed format and the second sparse matrix in column-compressed format to obtain an operation result matrix in row-compressed format or column-compressed format.
[0109] It can be seen that the matrix operation device 200 provided by the embodiment of the present application directly performs a sparse multiplication operation on the sparse matrices in row-compressed format and column-compressed format, and can directly obtain an operation result matrix in row-compressed format or column-compressed format. During the operation process, it is not necessary to decompress the compression format, realizing the multiplication operation of two input sparse matrices both in compressed format, reducing the amount of operation, operation power consumption, and storage capacity requirements.
[0110] In some embodiments, the reading module 21 reads the non-zero elements of the i-th row of the first sparse matrix, the column indexes of the non-zero elements of the i-th row, and the number of non-zero elements of the i-th row from the first storage unit each time, and reads the non-zero elements of the j-th column of the second sparse matrix, the row indexes of the non-zero elements of the j-th column, and the number of non-zero elements of the j-th column.
[0111] Specifically, the reading module 21 reads the non-zero elements of one row of the first sparse matrix, the column indexes of the non-zero elements of this row, and the number of non-zero elements of this row from the first storage unit each time, and at the same time reads the non-zero elements of one column of the second sparse matrix, the row indexes of the non-zero elements of this column, and the number of non-zero elements of this column from the first storage unit. It should be understood that the reading module 21 reads the non-zero elements of one row of the first sparse matrix from the first numerical array stored in the first storage unit. Similarly, the reading module 21 reads the column indexes of the non-zero elements of this row from the first index array stored in the first storage unit and reads the number of non-zero elements of this row from the first indication array stored in the first storage unit. The same applies to the second sparse matrix and will not be elaborated here.
[0112] Such as Figure 5As shown, in some embodiments, the operation module 22 includes a control unit 221 and a multiplier 222; wherein, the control unit 221, coupled to the reading module 21, is configured to determine, based on the non-zero elements of the i-th row of the first sparse matrix and the column indices of the non-zero elements of the i-th row, and the non-zero elements of the j-th column of the second sparse matrix and the row indices of the non-zero elements of the j-th column, target element pairs among the non-zero elements of the i-th row of the first sparse matrix and the non-zero elements of the j-th column of the second sparse matrix that need to perform dot product operations; the multiplier 222, coupled to the control unit 221 and the reading module 21, is configured to perform dot product operations on each of the target element pairs under the control of the control unit 221 to obtain the dot product operation results of each of the target element pairs.
[0113] Specifically, each time, the control unit 221 obtains, in the reading module 21, the non-zero elements of a row of the first sparse matrix and the column indices of the non-zero elements of that row, and the non-zero elements of a column of the second sparse matrix and the row indices of the non-zero elements of that column, and determines target element pairs among the non-zero elements of that row of the first sparse matrix and the non-zero elements of that column of the second sparse matrix that need to perform dot product operations. For example, the non-zero elements of the first row of a first sparse matrix A are {1, 2}, and the column indices of the non-zero elements of the first row are {1, 2}; the non-zero elements of the first column of a second sparse matrix B are {1, 2, 5}, and the row indices of the non-zero elements of the first column are {0, 1, 2}; then the target element pairs among the non-zero elements of the first row of the first sparse matrix A and the non-zero elements of the first column of the second sparse matrix B that need to perform dot product operations are {1,2} and {2, 5}.
[0114] The multiplier 222 performs dot product operations on each of the target element pairs determined by the control unit 221 in this time under the control of the control unit 221 to obtain the dot product operation results of each of the target element pairs. Continuing with the above first sparse matrix A and second sparse matrix B as an example, if the target element pairs determined by the control unit 221 in this time are {1,2} and {2, 5}, then the multiplier 222 performs a dot product operation on 1 and 2 in the target element pair {1, 2} to obtain a dot product operation result of 2, and performs a dot product operation on 2 and 5 in the target element pair {2, 5} to obtain a dot product operation result of 10.
[0115] Such as Figure 6As shown, in some embodiments, the operation module 22 further includes a first comparator 223 and an accumulator 224; wherein, the first comparator 223, coupled to the reading module 21 and the control unit 221, is configured to compare the number of non-zero elements in the i-th row of the first sparse matrix and the number of non-zero elements in the j-th column of the second sparse matrix, and send the smaller value to the control unit 221; the control unit 221 is further configured to determine an accumulation count based on the smaller value; the accumulator 224, coupled to the control unit 221 and the multiplier 222, is configured to accumulate the dot product operation results according to the accumulation count to obtain the element at the i-th row and j-th column of the operation result matrix.
[0116] Specifically, the first comparator 223 obtains the number of non-zero elements in the row of the first sparse matrix and the number of non-zero elements in the column of the second sparse matrix from the reading module 21, and sends the smaller value to the control unit 221. The control unit 221 determines an accumulation count based on the smaller value; the accumulator 224 accumulates the dot product operation results generated by the multiplier 222 according to the accumulation count to obtain the element at the corresponding row and corresponding column of the operation result matrix. Then, taking the first sparse matrix A and the second sparse matrix B as an example above, the number of non-zero elements in the first row of the first sparse matrix A is 2, and the number of non-zero elements in the first column of the second sparse matrix B is 3. The first comparator 223 compares the sizes of 2 and 3, obtains the smaller value of 2, and sends 2 to the control unit 221. The control unit 221 calculates the accumulation count = 2 - 1 = 1 based on the smaller value of 2; then, the control unit 221 controls the accumulator 224 to accumulate the dot product operation results 2 and 10 generated by the multiplier 222 according to the accumulation count of 1 to obtain the element 12 at the first row and first column of the operation result matrix.
[0117] As Figure 7 shown, in some embodiments, a second storage unit 23 is further included. The second storage unit 23 is coupled to the accumulator 224 and is configured to sequentially store the non-zero elements output by the accumulator 224 until the numerical array of the operation result matrix is obtained. The second storage unit and the first storage unit may refer to the same storage unit. Of course, they may also be different storage units. Optionally, the second storage unit 23 may be either a non-volatile memory (such as NAND Flash, NOR Flash, EEPROM, etc.) or a volatile memory (such as SRAM, DRAM, etc.).
[0118] As Figure 8As shown, in some embodiments, the operation module 22 further includes a counter 225; wherein, the counter 225 is coupled to the accumulator 224 and is configured to count the number of non-zero elements in each row of the operation result matrix output by the accumulator 224. The second storage unit 23 is also coupled to the counter 225 and is configured to sequentially store the number of non-zero elements in each row of the operation result matrix output by the counter 225 until an indication array of the operation result matrix is obtained; or, the counter 225 is coupled to the accumulator 224 and is configured to count the number of non-zero elements in each column of the operation result matrix output by the accumulator 224. The second storage unit 23 is also coupled to the counter 225 and is configured to sequentially store the number of non-zero elements in each column of the operation result matrix output by the counter 225 until an indication array of the operation result matrix is obtained.
[0119] As Figure 9 shown, in some embodiments, the operation module 22 further includes a second comparator 226; wherein,
[0120] The second comparator 226 is coupled to the reading module 21 and is configured to compare the column indices of each non-zero element in each row of the first sparse matrix with the row indices of each non-zero element in each column of the second sparse matrix, and send the comparison result to the control unit 221;
[0121] The control unit 221 is configured to determine the column index of each non-zero element in the operation result matrix according to the comparison result; the second storage unit 23 is also coupled to the control unit 221, and the second storage unit 23 is further configured to sequentially store the column indices of each non-zero element until an index array of the operation result matrix is obtained; or, the control unit 221 is further configured to determine the row index of each non-zero element in the operation result matrix according to the comparison result; the second storage unit 23 is also coupled to the control unit 221, and the second storage unit 23 is further configured to sequentially store the row indices of each non-zero element until an index array of the operation result matrix is obtained.
[0122] As Figure 10 shown, in some embodiments, the reading module 21 includes:
[0123] A first loading unit 211, coupled to the first storage unit, is configured to read the non-zero elements in the i-th row of the first sparse matrix and the column indices of the non-zero elements in the i-th row from the first storage unit, and store the non-zero elements in the i-th row of the first sparse matrix and the column indices of the non-zero elements in the i-th row into a first temporary storage unit 212;
[0124] The first temporary storage unit 212, coupled to the first loading unit 211, the second comparator 226, the control unit 221, and the multiplier 222, is configured to temporarily store the non-zero elements of the i-th row of the first sparse matrix and the column indices of the non-zero elements of each row of the i-th row;
[0125] The second loading unit 213, coupled to the first storage unit, is configured to read the non-zero elements of the j-th column of the second sparse matrix and the row indices of the non-zero elements of each column from the first storage unit, and store the non-zero elements of the j-th column of the second sparse matrix and the row indices of the non-zero elements of each column of the j-th column into the second temporary storage unit 214;
[0126] The second temporary storage unit 214, coupled to the second loading unit 213, the second comparator 226, the control unit 221, and the multiplier 222, is configured to temporarily store the non-zero elements of the j-th column of the second sparse matrix and the row indices of the non-zero elements of each column of the j-th column;
[0127] The third loading unit 215, coupled to the first storage unit, is configured to read the number of non-zero elements of the i-th row of the first sparse matrix from the first storage unit, and store the number of non-zero elements of the i-th row of the first sparse matrix into the third temporary storage unit 216;
[0128] The third temporary storage unit 216, coupled to the third loading unit 215, the first comparator 223, and the control unit 221, is configured to temporarily store the number of non-zero elements of the i-th row of the first sparse matrix;
[0129] The fourth loading unit 217, coupled to the first storage unit, is configured to read the number of non-zero elements of the j-th column of the second sparse matrix from the first storage unit, and store the number of non-zero elements of the j-th column of the second sparse matrix into the fourth temporary storage unit 218;
[0130] The fourth temporary storage unit 218, coupled to the fourth loading unit 217, the first comparator 223, and the control unit 221, is configured to temporarily store the number of non-zero elements of the j-th column of the second sparse matrix.
[0131] For a better understanding of the present application, the matrix operation device provided by the present application will be described in detail below through a specific embodiment.
[0132] The following is a matrix multiplication operation example provided by this embodiment.
[0133]
[0134] Store the above sparse matrix A in row-compressed format:
[0135] RN = [2, 2, 2];
[0136] V = [1, 2, 5, 3, 4, 2];
[0137] CI = [1, 2, 0, 1, 0, 2].
[0138] Store the above sparse matrix B in column-compressed format:
[0139] CN = [3, 1, 1];
[0140] V = [1, 2, 5, 5, 3];
[0141] RI = [0, 1, 2, 1, 1].
[0142] Figure 11 It is a schematic structural diagram of a matrix operation device provided by an embodiment of the present application. Figure 11 The descriptions of the respective functional modules are as follows:
[0143] MEM: Storage unit, used to store the compressed format data of the matrix;
[0144] LD: Data loading unit, used to load matrix data from storage;
[0145] Comparison: Comparison unit, used to compare whether the input data is equal, and output a control signal to the control unit for judgment;
[0146] FIFO: Temporary storage unit, used to temporarily store calculation data;
[0147] Control: Control unit, used to control the dequeue of FIFO data and control the calculations of the multiplier and accumulator;
[0148] Mul: Multiplier, implementing multiplication operations;
[0149] Acc: Accumulator, implementing accumulation operations;
[0150] NonZero Counter: Non-zero counter, used to detect the number of non-zero values in the output result.
[0151] In one embodiment, using the matrix operation device 200 as shown in Figure 11 the row-compressed format of the product operation result matrix C of matrix A and matrix B can be directly calculated. The working logic of the matrix operation device 200 is as follows:
[0152] Step 1: Store matrix A and matrix B in the storage unit (MEM) in row-compressed format and column-compressed format respectively;
[0153] Step 2: The data loading unit (LD) reads the elements of the numerical array V and the index array CI of matrix A in pairs from the storage unit (MEM), and reads the elements of the numerical array V and the index array RI of matrix B in pairs, and stores them in the corresponding FIFOs;
[0154] Step 3: The data loading unit (LD) reads RN of matrix A and CN of matrix B from the storage unit (MEM) respectively, and stores them in the corresponding FIFOs;
[0155] Step 4: Each clock cycle, the FIFO reads out data under the action of the deq signal (a control signal) of the control unit;
[0156] Step 5: Compare the sizes of RN and CN, and send the smaller value acc_number to the control unit. The control unit controls the accumulation times of the ACC accumulator based on this value;
[0157] Step 6: The comparator compares whether the RI and CI values are equal, and sends the comparison signal to the control unit. If the RI and CI values are equal, the control signal is output to control the corresponding V to perform multiplication. If they are not equal, the multiplier is not enabled, and the corresponding V value is operated, and the multiplier directly outputs a 0 value;
[0158] Step 7: The non-zero counter counts the number of non-zero values in each column of the ACC result and stores it in the output RN FIFO.
[0159] Step 8: The control unit judges the column number of the non-zero result according to the comparison result of the input CI and RI, outputs the CI value, and stores it in the corresponding FIFO;
[0160] Step 9: The ACC accumulator stores the calculated non-zero element values in V;
[0161] After the calculation is completed, {RN, V, CI} in the three output FIFOs is the output operation result matrix, and the matrix is directly presented in row compression format.
[0162] Step 10: Store the values of the three output FIFOs in the storage unit MEM, that is, complete the operation process.
[0163] The following gives the detailed operation process of this embodiment:
[0164] First, calculate C 00 (the element in the first row and first column of matrix C):
[0165] The FIFO reads A 0j ={1, 2}, reads the indication A in CI 0j of {1, 2}, reads the indication A in RN 0j of {2}; where j = {1, 2, 3}, A 0jRepresents the data of the first row of matrix A.
[0166] FIFO reads B i0 ={1, 2, 5}, reads the indication of B in RI i0 of {0, 1, 2}, reads the indication of B in CN i0 of {3}; where i = {1, 2, 3}, B i0 Represents the data of the first column of matrix B.
[0167] Compares RN and CN, the smaller value is 2, then the control unit determines that the accumulation count is 2 - 1 = 1.
[0168] Compares CI and RI, the control unit controls the multiplier to multiply the corresponding equal values of CI and RI in A 0j and B i0 : 1 * 2 = 2, 2 * 5 = 10; at the same time, the control unit gets the column index 0 of the non - zero result.
[0169] The accumulator adds the results output by the multiplier according to the accumulation count 1: C 00 = 2 + 10 = 12.
[0170] Stores the value 12 of C 00 into V.
[0171] Secondly, calculates C 01 (the element in the first row and second column of matrix C):
[0172] FIFO takes out A 0j ={1, 2}, takes out the indication of A in CI 0j of {1, 2}, takes out the indication of A in RN 0j of {2}.
[0173] FIFO takes out B i1 ={5}, takes out the indication of B in RI i1 of {2}, takes out the indication of B in CN i1 of {1}; where B i1 Represents the data of the second column of matrix B.
[0174] Compares RN and CN, the smaller value is 1, then the control unit determines that the accumulation count is 1 - 1 = 0.
[0175] Compares CI and RI, the control unit controls the multiplier to multiply the corresponding equal values of CI and RI in A 0k and B k0 : 2 * 5 = 10; at the same time, the control unit gets the column number 1 of the non - zero result.
[0176] The accumulator gets according to the accumulation count 0: C = 10.
[0177] Store the value 10 of C 01 into V.
[0178] Next, calculate C 02 (the element in the first row and the third column of matrix C):
[0179] FIFO reads A 0j = {1, 2}, reads the indication of A in CI 0j of {1, 2}, reads the indication of A in RN 0j of {2};
[0180] FIFO reads B i2 = {3}, reads the indication of B in RI i2 of {1}, reads the indication of B in CN i2 of {1}; where B i2 represents the data in the second column of matrix B;
[0181] Compare RN and CN, the smaller value is 1, then the controller determines that the accumulation times is 1 - 1 = 0;
[0182] Compare CI and RI, the control unit controls the multiplier to multiply the corresponding equal values of A 0k and B k0 : 1 * 3 = 3; At the same time, the control unit obtains the column number 2 of the non - zero result;
[0183] The accumulator obtains according to the accumulation times 0: C 02 = 3;
[0184] Store the value 3 of C 02 into V.
[0185] So far, C 00 、C 01 、C 02 The calculation is completed, the non - zero counter obtains the first element of RN as 3, and stores 3 into RN.
[0186] And so on, until the V, RN, and CI of matrix C are obtained, that is, the row - compressed format of matrix C.
[0187] In another embodiment, a matrix operation device 200 as shown in Figure 11 is adopted. Similar to the calculation logic of the above - mentioned embodiment, by changing the calculation sequence of the input data and the judgment logic of the control unit, the output matrix result in column - compressed format can also be obtained. Specifically, in this embodiment, the working logic of the matrix operation device 200 is as follows:
[0188] Step 1: Store matrix A and matrix B in the storage unit (MEM) in row - compressed format and column - compressed format respectively;
[0189] Step 2: The data loading unit (LD) reads the elements in the numerical array V of matrix A and the index array CI, and the elements in the numerical array V of matrix B and the index array RI in pairs from the storage unit (MEM), and stores them in the corresponding FIFOs;
[0190] Step 3: The data loading unit (LD) reads RN of matrix A and CN of matrix B from the storage unit (MEM) respectively, and stores them in the corresponding FIFOs;
[0191] Step 4: Each clock cycle, the FIFO reads out data under the action of the deq signal (a control signal) of the control unit;
[0192] Step 5: Compare the sizes of RN and CN, and send the smaller value acc_number to the control unit. The control unit controls the accumulation times of the ACC accumulator based on this value;
[0193] Step 6: The comparator compares whether the RI and CI values are equal, and sends the comparison signal to the control unit. If the RI and CI values are equal, the control signal is output to control the corresponding V to perform multiplication. If they are not equal, the multiplier is not enabled, and the corresponding V value is operated, and the multiplier directly outputs a 0 value;
[0194] Step 7: The non-zero counter counts the number of non-zero values in each column of the ACC result and stores it in the output RN FIFO.
[0195] Step 8: The control unit judges the row number of the non-zero result according to the comparison result of the input CI and RI, outputs the RI value, and stores it in the corresponding FIFO;
[0196] Step 9: The ACC accumulator stores the calculated non-zero element value in V;
[0197] After the calculation is completed, {CN, V, RI} in the three output FIFOs is the output operation result matrix, and the matrix is directly presented in column compression format.
[0198] Step 10: Store the values of the three output FIFOs in the storage unit MEM, and the operation process is completed.
[0199] The following gives the detailed calculation process of this embodiment:
[0200] First, calculate C 00 (the element in the first row and first column of matrix C):
[0201] The FIFO reads A 0j ={1, 2}, reads {1, 2} indicating A in CI 0j and reads {2} indicating A in RN 0j ; where j = {1, 2, 3}, A0j Represents the data of the first row of matrix A.
[0202] FIFO reads B i0 ={1, 2, 5}, reads the indication of B in RI i0 of {0, 1, 2}, reads the indication of B in CN i0 of {3}; where i = {1, 2, 3}, B i0 Represents the data of the first column of matrix B.
[0203] Compares RN and CN, the smaller value is 2, then the control unit determines that the accumulation times is 2 - 1 = 1.
[0204] Compares CI and RI, the control unit controls the multiplier to multiply the corresponding equal values of CI and RI in A 0j and B i0 : 1 * 2 = 2, 2 * 5 = 10; at the same time, the control unit gets the row index 0 of the non - zero result.
[0205] The accumulator adds the results output by the multiplier according to the accumulation times 1: C 00 = 2 + 10 = 12.
[0206] Stores the value 12 of C 00 into V.
[0207] Secondly, calculates C 10 (the element in the second row and the first column of matrix C):
[0208] FIFO takes out A 1j ={5, 3}, takes out the indication of A in CI 1j of {0, 1}, takes out the indication of A in RN 1j of {2}.
[0209] FIFO takes out B i0 ={1, 2, 5}, takes out the indication of B in RI i0 of {0, 1, 2}, takes out the indication of B in CN i0 of {3}; where B i0 Represents the data of the first column of matrix B.
[0210] Compares RN and CN, the smaller value is 2, then the control unit determines that the accumulation times is 2 - 1 = 1.
[0211] Compares CI and RI, the control unit controls the multiplier to multiply the corresponding equal values of CI and RI in A 1j and B i0 : 1 * 5 = 5, 3 * 2 = 6; at the same time, the control unit gets the row number 1 of the non - zero result.
[0212] The accumulator gets according to the accumulation times 1: C10 = 5 + 6 = 11.
[0213] Store the value 11 of C 10 into V.
[0214] Next, calculate C 20 (the element in the 3rd row and 1st column of matrix C):
[0215] FIFO reads A 2j = {4, 2}, reads the indication of A in CI 2j as {0, 2}, reads the indication of A in RN 2j as {2};
[0216] FIFO reads B i2 = {1, 2, 5}, reads the indication of B in RI i2 as {0, 1, 2}, reads the indication of B in CN i2 as {3}; where B i2 represents the 3rd column data of matrix B.
[0217] Compare RN and CN, the smaller value is 2, then the controller determines that the accumulation times is 2 - 1 = 1;
[0218] Compare CI and RI, the control unit controls the multiplier to multiply the corresponding equal values of A 2j and B i0 : 4 * 1 = 4, 2 * 5 = 10; At the same time, the control unit obtains the row number 2 of the non - zero result;
[0219] The accumulator obtains according to the accumulation times 1: C 20 = 4 + 10 = 14;
[0220] Store the value 14 of C 20 into V.
[0221] So far, the calculations of C 00 , C 10 , C 20 are completed. The non - zero counter obtains the first element of CN as 3 and stores 3 into RN.
[0222] And so on, until the V, CN, and RI of matrix C are obtained, that is, the column - compressed format of matrix C.
[0223] Figure 12 is a flowchart of a matrix operation method provided by an embodiment of the present application. As Figure 12 shown, the present application provides a matrix operation method, including:
[0224] S31. The reading module reads the first sparse matrix in row-compressed format and the second sparse matrix in column-compressed format stored in the first storage unit respectively. Wherein, the first storage unit stores the first sparse matrix in row-compressed format and the second sparse matrix in column-compressed format. The first sparse matrix in row-compressed format is represented by a first numerical array, a first index array and a first indication array. The first numerical array stores the values of the non-zero elements of the first sparse matrix in row-major order. The first index array stores the column indices of the non-zero elements of the first sparse matrix in row-major order. The first indication array stores the number of non-zero elements in each row of the first sparse matrix in row-major order. The second sparse matrix in column-compressed format is represented by a second numerical array, a second index array and a second indication array. The second numerical array stores the values of the non-zero elements of the second sparse matrix in column-major order. The second index array stores the row indices of the non-zero elements of the second sparse matrix in column-major order. The second indication array stores the number of non-zero elements in each column of the second sparse matrix in column-major order;
[0225] S32. The operation module performs a sparse multiplication operation on the first sparse matrix and the second sparse matrix based on the first sparse matrix in row-compressed format and the second sparse matrix in column-compressed format, and obtains an operation result matrix in row-compressed format or column-compressed format.
[0226] For the specific calculation process of the matrix operation method provided in the embodiments of the present application, see the introduction content of the matrix operation device in the above embodiments, and details are not described herein again. By using the matrix operation method provided in the embodiments of the present application, a sparse multiplication operation is directly performed on the sparse matrices in row-compressed format and column-compressed format, and an operation result matrix in row-compressed format or column-compressed format can be directly obtained. During the operation process, there is no need to decompress the compressed format, realizing the multiplication operation of two sparse matrices both in compressed format as input, reducing the operation amount, operation power consumption and storage capacity requirements.
[0227] In some embodiments, the reading module reads the non-zero elements of the i-th row of the first sparse matrix, the column indices of the non-zero elements of the i-th row, and the number of non-zero elements of the i-th row from the storage unit each time, and reads the non-zero elements of the j-th column of the second sparse matrix, the row indices of the non-zero elements of the j-th column, and the number of non-zero elements of the j-th column.
[0228] In some embodiments, the control unit of the operation module determines target element pairs that need to perform dot product operations among the non-zero elements in the i-th row of the first sparse matrix and the column indices of each non-zero element in the i-th row, and among the non-zero elements in the j-th column of the second sparse matrix and the row indices of each non-zero element in the j-th column; the multiplier of the operation module performs dot product operations on each of the target element pairs under the control of the control unit to obtain the dot product operation results of each of the target element pairs.
[0229] In some embodiments, the first comparator of the operation module compares the number of non-zero elements in the i-th row of the first sparse matrix with the number of non-zero elements in the j-th column of the second sparse matrix, and sends the smaller value to the control unit; the control unit determines an accumulation count based on the smaller value; the accumulator of the operation module accumulates the dot product operation results according to the accumulation count to obtain the element at the i-th row and j-th column of the operation result matrix.
[0230] In some embodiments, the storage unit also sequentially stores the non-zero elements output by the accumulator until the numerical array of the operation result matrix is obtained.
[0231] In some embodiments, the counter of the operation module counts the number of non-zero elements in each row of the operation result matrix output by the accumulator; the storage unit sequentially stores the number of non-zero elements in each row of the operation result matrix output by the counter until the indication array of the operation result matrix is obtained; or, the counter counts the number of non-zero elements in each column of the operation result matrix output by the accumulator; the storage unit sequentially stores the number of non-zero elements in each column of the operation result matrix output by the counter until the indication array of the operation result matrix is obtained.
[0232] In some embodiments, the second comparator of the operation module compares the column index of each non-zero element in each row of the first sparse matrix with the row index of each non-zero element in each column of the second sparse matrix, and sends the comparison result to the control unit; the control unit also determines the column index of each non-zero element in the operation result matrix according to the comparison result, and the storage unit also sequentially stores the column index of each non-zero element until the index array of the operation result matrix is obtained; or, the control unit also determines the row index of each non-zero element in the operation result matrix according to the comparison result, and the storage unit also sequentially stores the row index of each non-zero element until the index array of the operation result matrix is obtained.
[0233] In some embodiments, the first loading unit of the reading module reads the non-zero elements of the i-th row of the first sparse matrix and the column indices of the non-zero elements of the i-th row from the storage unit, and stores the non-zero elements of the i-th row of the first sparse matrix and the column indices of the non-zero elements of the i-th row into the first temporary storage unit; the first temporary storage unit of the reading module temporarily stores the non-zero elements of the i-th row of the first sparse matrix and the column indices of the non-zero elements of the i-th row; the second loading unit of the reading module reads the non-zero elements of the j-th column of the second sparse matrix and the row indices of the non-zero elements of each column from the storage unit, and stores the non-zero elements of the j-th column of the second sparse matrix and the row indices of the non-zero elements of the j-th column into the second temporary storage unit; the second temporary storage unit of the reading module temporarily stores the non-zero elements of the j-th column of the second sparse matrix and the row indices of the non-zero elements of the j-th column; the third loading unit of the reading module reads the number of non-zero elements of the i-th row of the first sparse matrix from the storage unit, and stores the number of non-zero elements of the i-th row of the first sparse matrix into the third temporary storage unit; the third temporary storage unit of the reading module temporarily stores the number of non-zero elements of the i-th row of the first sparse matrix; the fourth loading unit of the reading module reads the number of non-zero elements of the j-th column of the second sparse matrix from the storage unit, and stores the number of non-zero elements of the j-th column of the second sparse matrix into the fourth temporary storage unit; the fourth temporary storage unit of the reading module temporarily stores the number of non-zero elements of the j-th column of the second sparse matrix.
[0234] An embodiment of the present application further provides an electronic device, which includes the matrix operation device or the matrix storage device described in any of the above embodiments.
[0235] Since the electronic device provided by the embodiment of the present application includes the matrix operation device or the matrix storage device, it can achieve the same technical effects as the matrix operation device or the matrix storage device, which will not be elaborated here.
[0236] Figure 13 It is a schematic physical structure diagram of an electronic device provided by an embodiment of the present application. As Figure 13 shown, the electronic device 300 may include: a processor 31, a communication interface 32, a memory 33, and a communication bus 34. Among them, the processor 31, the communication interface 32, and the memory 33 communicate with each other through the communication bus 34. The processor 31 may include the matrix operation device or the matrix storage device described in any of the above embodiments, and the processor 31 may call the logical instructions in the memory 33 to execute the method described in any of the above embodiments.
[0237] In addition, when the logical instructions in the above-mentioned memory 33 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0238] This embodiment claims a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided by the above-mentioned method embodiments.
[0239] This embodiment provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. The computer program causes the computer to execute the methods provided by the above-mentioned method embodiments.
[0240] Those skilled in the art should understand that the embodiments of this application can be provided as a method, a system, or a computer program product. Therefore, this application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program codes.
[0241] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0242] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0243] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0244] In the description of this specification, the descriptions with reference to the terms "one embodiment", "a specific embodiment", "some embodiments", "for example", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0245] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above are only specific embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A matrix storage device, characterized in that: include: A processing unit is configured to traverse the sparse matrix in a row-first order, and generate a value array and an index array of the sparse matrix according to the values and column indices of all non-zero elements in the sparse matrix; When traversing each row of the sparse matrix, recording the number of non-zero elements in each row of the sparse matrix, and generating an indication array of the sparse matrix according to the number of non-zero elements in each row of the sparse matrix; A storage unit, coupled to the processing unit, configured to store a value array, an index array, and an indicator array of the sparse matrix, wherein the value array, the index array, and the indicator array are row-compressed format representations of the sparse matrix; or, A processing unit is configured to traverse the sparse matrix in column priority order, and generate a value array and an index array of the sparse matrix according to the values and row indices of all non-zero elements in the sparse matrix; when traversing each column of the sparse matrix, record the number of non-zero elements in each column of the sparse matrix, and generate an indication array of the sparse matrix according to the number of non-zero elements in each column of the sparse matrix; A storage unit, coupled to the processing unit, is configured to store a value array, an index array and an indicator array of the sparse matrix, wherein the value array, the index array and the indicator array are column compressed format representations of the sparse matrix.
2. A matrix operation device, characterized in that: include: A reading module, coupled to a first storage unit, wherein the first storage unit stores a first sparse matrix in a row compression format and a second sparse matrix in a column compression format, and the reading module is configured to read the first sparse matrix in the row compression format and the second sparse matrix in the column compression format from the first storage unit, wherein the first sparse matrix in the row compression format is represented by a first numerical array, a first index array, and a first indicator array, wherein the first numerical array stores the values of non-0 elements of the first sparse matrix in a row priority order, the first index array stores the column indexes of the non-0 elements of the first sparse matrix in a row priority order, the first indicator array stores the number of non-0 elements of each row of the first sparse matrix in a row priority order, and the second sparse matrix in the column compression format is represented by a second numerical array, a second index array, and a second indicator array, wherein the second numerical array stores the values of non-0 elements of the second sparse matrix in a column priority order, the second index array stores the row indexes of the non-0 elements of the second sparse matrix in a column priority order, and the second indicator array stores the number of non-0 elements of each column of the second sparse matrix in a column priority order; The operation module is coupled to the reading module and is configured to perform sparse multiplication operations on the first sparse matrix in the row compression format and the second sparse matrix in the column compression format to obtain an operation result matrix in the row compression format or the column compression format.
3. The matrix operation device according to claim 2, characterized in that: The reading module reads the non-0 elements in the i-th row of the first sparse matrix, the column index of each non-0 element in the i-th row, and the number of non-0 elements in the i-th row from the first storage unit each time, and reads the non-0 elements in the j-th column of the second sparse matrix, the row index of each non-0 element in the j-th column, and the number of non-0 elements in the j-th column.
4. The matrix operation device according to claim 3, characterized in that: The operation module includes a control unit and a multiplier; wherein, The control unit is coupled to the reading module and is configured to determine a target element pair that needs to be subjected to a dot product operation among the non-0 elements in the i-th row of the first sparse matrix and the non-0 elements in the j-th column of the second sparse matrix based on the column index of the non-0 elements in the i-th row and each non-0 element in the i-th row of the first sparse matrix, and the row index of the non-0 elements in the j-th column and each non-0 element in the j-th column of the second sparse matrix; The multiplier is coupled to the control unit and the reading module, and is configured to perform a dot product operation on each of the target element pairs under the control of the control unit to obtain a dot product operation result of each of the target element pairs.
5. The matrix operation device according to claim 4, characterized in that: The operation module also includes a first comparator and an accumulator; wherein, The first comparator is coupled to the reading module and the control unit, and is configured to compare the number of non-0 elements in the i-th row of the first sparse matrix and the number of non-0 elements in the j-th column of the second sparse matrix, and send the smaller value to the control unit; The control unit is further configured to determine an accumulation number based on the smaller value; The accumulator is coupled to the control unit and the multiplier, and is configured to accumulate the dot product operation results according to the number of accumulations to obtain the element of the i-th row and j-th column of the operation result matrix.
6. The matrix operation device according to claim 5, characterized in that: It also includes a second storage unit coupled to the accumulator, and the second storage unit is configured to sequentially store non-zero elements output by the accumulator until a numerical array of the operation result matrix is obtained.
7. The matrix operation device according to claim 6, characterized in that: The operation module also includes a counter; wherein, The counter is coupled to the accumulator and configured to count the number of non-zero elements in each row of the operation result matrix output by the accumulator; the second storage unit is also coupled to the counter and configured to sequentially store the number of non-zero elements in each row of the operation result matrix output by the counter until an indication array of the operation result matrix is obtained; or, The counter is coupled to the accumulator and is configured to count the number of non-zero elements in each column of the operation result matrix output by the accumulator; the second storage unit is also coupled to the counter and is configured to sequentially store the number of non-zero elements in each column of the operation result matrix output by the counter until an indication array of the operation result matrix is obtained.
8. The matrix operation device according to claim 7, characterized in that: The operation module also includes a second comparator; wherein, The second comparator is coupled to the reading module and is configured to compare the column index of each non-0 element in each row of the first sparse matrix with the row index of each non-0 element in each column of the second sparse matrix, and send the comparison result to the control unit; The control unit is configured to determine the column index of each non-0 element in the operation result matrix according to the comparison result; the second storage unit is also coupled to the control unit, and the second storage unit is also configured to store the column index of each non-0 element in sequence until the index array of the operation result matrix is obtained; or, the control unit is also configured to determine the row index of each non-0 element in the operation result matrix according to the comparison result; the second storage unit is also coupled to the control unit, and the second storage unit is also configured to store the row index of each non-0 element in sequence until the index array of the operation result matrix is obtained.
9. The matrix operation device according to claim 8, characterized in that: The reading module comprises: A first loading unit, coupled to the first storage unit, is configured to read the non-0 elements in the i-th row of the first sparse matrix and the column index of each non-0 element in the i-th row from the first storage unit, and store the non-0 elements in the i-th row of the first sparse matrix and the column index of each non-0 element in the i-th row into a first temporary storage unit; A first temporary storage unit, coupled to the first loading unit, the second comparator, the control unit and the multiplier, configured to temporarily store the non-zero elements of the i-th row of the first sparse matrix and the column index of each non-zero element of the i-th row; A second loading unit, coupled to the first storage unit, is configured to read the non-0 elements of the j-th column of the second sparse matrix and the row index of each non-0 element of each column from the first storage unit, and store the non-0 elements of the j-th column of the second sparse matrix and the row index of each non-0 element of the j-th column into a second temporary storage unit; A second temporary storage unit, coupled to the second loading unit, the second comparator, the control unit and the multiplier, configured to temporarily store the non-zero elements of the j-th column of the second sparse matrix and the row index of each non-zero element of the j-th column; a third loading unit, coupled to the first storage unit, configured to read the number of non-0 elements in the i-th row of the first sparse matrix from the first storage unit, and store the number of non-0 elements in the i-th row of the first sparse matrix into a third temporary storage unit; A third temporary storage unit, coupled to the third loading unit, the first comparator and the control unit, and configured to temporarily store the number of non-zero elements in the i-th row of the first sparse matrix; a fourth loading unit, coupled to the first storage unit, configured to read the number of non-0 elements in the j-th column of the second sparse matrix from the first storage unit, and store the number of non-0 elements in the j-th column of the second sparse matrix into a fourth temporary storage unit; A fourth temporary storage unit is coupled to the fourth loading unit, the first comparator and the control unit, and is configured to temporarily store the number of non-zero elements in the j-th column of the second sparse matrix.
10. A matrix storage method, characterized in that: include: The processing unit traverses the sparse matrix in a row-priority order, and generates a value array and an index array of the sparse matrix according to the values and column indices of all non-zero elements in the sparse matrix; When traversing each row of the sparse matrix, recording the number of non-zero elements in each row of the sparse matrix, and generating an indication array of the sparse matrix according to the number of non-zero elements in each row of the sparse matrix; The storage unit stores a value array, an index array and an indicator array of the sparse matrix, wherein the value array, the index array and the indicator array are row compressed format representations of the sparse matrix; or, The processing unit traverses the sparse matrix in column priority order, and generates a value array and an index array of the sparse matrix according to the values and row indexes of all non-zero elements in the sparse matrix; when traversing each column of the sparse matrix, the number of non-zero elements in each column of the sparse matrix is recorded, and an indication array of the sparse matrix is generated according to the number of non-zero elements in each column of the sparse matrix; The storage unit stores a value array, an index array and an indicator array of the sparse matrix, wherein the value array, the index array and the indicator array are represented in a column compression format of the sparse matrix.
11. A matrix operation method, characterized in that: include: The reading module reads the first sparse matrix in row compression format and the second sparse matrix in column compression format stored in the first storage unit respectively, wherein the first storage unit stores the first sparse matrix in row compression format and the second sparse matrix in column compression format, the first sparse matrix in row compression format is represented by a first numerical array, a first index array and a first indicator array, the first numerical array stores the values of non-0 elements of the first sparse matrix in row priority order, the first index array stores the column indexes of the non-0 elements of the first sparse matrix in row priority order, the first indicator array stores the number of non-0 elements in each row of the first sparse matrix in row priority order, the second sparse matrix in column compression format is represented by a second numerical array, a second index array and a second indicator array, the second numerical array stores the values of non-0 elements of the second sparse matrix in column priority order, the second index array stores the row indexes of the non-0 elements of the second sparse matrix in column priority order, and the second indicator array stores the number of non-0 elements in each column of the second sparse matrix in column priority order; The operation module performs a sparse multiplication operation on the first sparse matrix in the row compression format and the second sparse matrix in the column compression format to obtain an operation result matrix in the row compression format or the column compression format.
12. The matrix operation method according to claim 11, characterized in that: The reading module reads the non-0 elements of the i-th row of the first sparse matrix, the column index of each non-0 element of the i-th row, and the number of non-0 elements of the i-th row from the first storage unit each time, and reads the non-0 elements of the j-th column of the second sparse matrix, the row index of each non-0 element of the j-th column, and the number of non-0 elements of the j-th column; The control unit of the operation module determines a target element pair that needs to perform a dot product operation on the non-0 elements in the i-th row of the first sparse matrix and the non-0 elements in the j-th column of the second sparse matrix based on the column index of the non-0 elements in the i-th row and each non-0 element in the i-th row of the first sparse matrix, and the row index of the non-0 elements in the j-th column and each non-0 element in the j-th column of the second sparse matrix; The multiplier of the operation module performs a dot product operation on each of the target element pairs under the control of the control unit to obtain a dot product operation result of each of the target element pairs; The first comparator of the operation module compares the number of non-0 elements in the i-th row of the first sparse matrix and the number of non-0 elements in the j-th column of the second sparse matrix, and sends the smaller value to the control unit; The control unit determines an accumulation number based on the smaller value; The accumulator of the operation module accumulates the dot product operation results according to the number of accumulations to obtain the element of the i-th row and j-th column of the operation result matrix.
13. The matrix operation method according to claim 12, characterized in that: The method further comprises: The second storage unit sequentially stores the non-zero elements output by the accumulator until a numerical array of the result matrix is obtained; The counter of the operation module counts the number of non-zero elements in each row of the operation result matrix output by the accumulator, and the second storage unit sequentially stores the number of non-zero elements in each row of the operation result matrix output by the counter until an indication array of the operation result matrix is obtained, or the counter counts the number of non-zero elements in each column of the operation result matrix output by the accumulator, and the second storage unit sequentially stores the number of non-zero elements in each column of the operation result matrix output by the counter until an indication array of the operation result matrix is obtained.
14. The matrix operation method according to claim 13, characterized in that: The second comparator of the operation module compares the column index of each non-0 element in each row of the first sparse matrix and the row index of each non-0 element in each column of the second sparse matrix, and sends the comparison result to the control unit; The control unit also determines the column index of each non-0 element in the operation result matrix according to the comparison result, and the second storage unit also stores the column index of each non-0 element in sequence until the index array of the operation result matrix is obtained, or the control unit also determines the row index of each non-0 element in the operation result matrix according to the comparison result, and the second storage unit also stores the row index of each non-0 element in sequence until the index array of the operation result matrix is obtained.
15. An electronic device, characterized in that: It comprises the matrix storage device as described in claim 1 or the matrix operation device as described in any one of claims 2 to 9.
Citation Information
Patent Citations
Method and device for realizing sparse matrix multiplication on reconfigurable processor array
CN112507284A
In-memory sparse matrix multiplication operation method, equation solving method and solver
CN113870918A
Universal sparse matrix multiplication implementation method and device based on 2D systolic array
CN115328440A
Calculation device, calculation method and related product
CN117235424A
Sparse matrix calculations untilizing ightly coupled memory and gather / scatter engine
US20220019430A1