Hardware acceleration method and device for sparse matrix multiplication operation of convolutional neural layer
Through the Img2col operation and CSC encoding module, the sparse matrix of the convolutional neural network is quickly processed in the edge device, which solves the problems of space resource waste and computing time overhead, and realizes efficient computing and storage.
Patent Information
- Application Number
- CN202510250163.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
When deploying convolutional neural networks at edge devices, sparse parameter matrix results in waste of space resources and increased computational time overhead.
The original image and convolution kernel are converted into sparse matrices through Img2col operation, and encoded using the CSC encoding module to quickly obtain non-0 elements, and realize fast computing and compression of storage space.
It greatly compresses the data storage space, reduces computing time, is small overhead, and improves the computing efficiency of edge devices.
Smart Images

Figure CN120179975A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of integrated circuit design, and particularly to a hardware acceleration method and device for sparse matrix multiplication operations for convolutional neural layers. Background Art
[0002] In recent years, the deployment of artificial intelligence models on edge devices has gradually become a popular research direction. Convolutional neural networks are a relatively common mode in the application of artificial intelligence. Compared with directly implementing neural network operations in hardware, a more general approach is to implement matrix multiplication operations. After years of development, this technology is relatively mature, and the img2col technology can be used to achieve this conversion.
[0003] The patent application with the publication number CN118966293A discloses a neural network data operation acceleration device, method, equipment, and medium. By first obtaining the input data of the neural network operation and Cin convolutional kernel data, the feature matrix SMF is first divided into N feature sub-matrices MF in the C direction. Then, for each feature sub-matrix MF, an index table for the img2col operation after padding is generated. Then, the index table for the img2col operation is mapped back to the index table of the original feature matrix SMF. Finally, the localization of the data for the neural network operation is achieved and the operation process is entered, and multiple sliding windows can be calculated at one time, that is, multiple data can be read in at one time, thus greatly reducing the number of data accesses and significantly improving the operation efficiency. Compared with traditional technologies, the localization operation of the neural network is realized, effectively reducing the cache occupancy of the computer while reducing the number of accesses to the feature data, and can be applied to the operation acceleration of multiple neural networks.
[0004] The patent application with the publication number CN117785031A discloses a data processing method, device, electronic device, and storage medium. The method includes: in response to a data processing instruction for a target input feature map, obtaining a target input data block from the input data block arrangement corresponding to the target input feature map in the memory; the input data block arrangement is obtained by dividing the target input feature map into multiple input data blocks according to the data throughput of the general matrix processing engine in each clock cycle and storing them in the memory according to a preset arrangement method; caching the target input data block; and based on the cached target input data block, performing matrix multiplication operation processing using the general matrix processing engine to obtain a data processing result. The embodiments of the present disclosure improve the data processing efficiency based on Img2Col, and further improve the data processing efficiency of the electronic device based on the deep learning network.
[0005] Generally, before deploying edge devices, the model is pruned and quantized, resulting in a relatively sparse parameter matrix with many zero elements. However, in actual operation, only non-zero elements can participate in the calculation. The methods disclosed in the above two patent applications have two problems when deploying the matrix converted by img2col to edge devices. First, the edge device must store the matrix completely, but the zero elements themselves do not participate in the calculation, which will cause waste of a large amount of space resources. Second, even if only non-zero elements participate in the operation, the edge device still needs to access all addresses in sequence to traverse the matrix, which will cause unnecessary time overhead. Summary of the Invention
[0006] The present invention provides a hardware acceleration method for sparse matrix multiplication operations for convolutional neural layers. This hardware acceleration method can quickly obtain non-zero elements and perform fast operations on non-zero elements with relatively small overhead.
[0007] A specific embodiment of the present invention provides a hardware acceleration method for sparse matrix multiplication operations for convolutional neural layers, including:
[0008] The original image and the convolution kernel are respectively obtained as a pixel calculation matrix and a convolution kernel calculation matrix through the Img2col operation. The pixel calculation matrix and the convolution kernel calculation matrix are respectively encoded through a CSC encoding module to obtain a first encoding result and a second encoding result. The CSC encoding module includes a one-dimensional encoding array D1, a one-dimensional encoding array D2, and a one-dimensional encoding array D3. The one-dimensional encoding array D1 is used to describe the non-zero elements in the sparse matrix from top to bottom and from left to right. The one-dimensional encoding array D2 is used to describe the row numbers corresponding to the non-zero elements. The one-dimensional encoding array D3 is used to describe the index of the first non-zero element in each column other than the first column in the sparse matrix in the one-dimensional encoding array D1;
[0009] The pixel calculation matrix and the convolution kernel calculation matrix are respectively cut into multiple pixel calculation matrix units and multiple convolution kernel calculation matrix units. Based on the first and second encoding results, the non-zero elements traversed from the pixel calculation matrix units and the corresponding convolution kernel calculation matrix units are operated to obtain matrix unit operation results. The multiple matrix unit operation results are combined to obtain the operation results of the pixel calculation matrix and the convolution kernel calculation matrix.
[0010] Preferably, operating the non-zero elements traversed from the pixel calculation matrix units and the corresponding convolution kernel calculation matrix units based on the first and second encoding results to obtain matrix unit operation results includes:
[0011] S31. Initialize the partial sum storage space;
[0012] S32. Traverse the non-zero element B(m, n) in the first non-empty column of the convolution kernel calculation matrix through the one-dimensional encoding arrays D3 and D1 of the second encoding result, and record the column number colNum;
[0013] S33. Traverse the non-zero element in the first non-empty column of the pixel calculation matrix through the one-dimensional encoding arrays D3 and D1 of the first encoding result, and then traverse the non-zero elements of the pixel calculation matrix through the one-dimensional encoding array D1 to obtain the element A(k, m) whose column number is equal to m;
[0014] S34. Traverse the pixel calculation matrix through the one-dimensional encoding array D2 of the first encoding result to obtain the row number of the element A(k, m), record the row number rowNum, and calculate the storage location site through the recorded column number colNum and row number rowNum;
[0015] S35. Access the value in the partial sum storage space at the address site, add the value to the calculation result of A(k, m) × B(m, n), and write it back to the partial sum storage space at the storage location site;
[0016] S36. According to the D3 and D1 encoding arrays of the first encoding result, traverse to the next non-zero element in the pixel calculation matrix whose column number is equal to m, and repeat steps S34 - S36 until the pixel calculation matrix traversal is completed;
[0017] S37. Traverse to the position of the non-zero element in the next non-empty column of the convolution kernel calculation matrix through the D3 and D1 encoding arrays of the second encoding result, and record the column number of the non-zero element in the next non-empty column, and repeat steps S33 - S36 until the convolution kernel calculation matrix traversal is completed to obtain the matrix unit operation result, and the operation is completed.
[0018] Preferably, calculate the storage location site = rowNum + colNum * n, where n is the number of rows of the pixel calculation matrix.
[0019] Preferably, merge the operation results of multiple matrix units to obtain the operation results of the pixel calculation matrix and the convolution kernel calculation matrix, including:
[0020] After the operation results of all pixel calculation matrix units and the corresponding convolution kernel calculation matrix units are stored in the partial sum storage space, perform addition to obtain the operation results of the pixel calculation matrix and the convolution kernel calculation matrix. After outputting the operation results of the pixel calculation matrix and the convolution kernel calculation matrix, initialize the partial sum storage space, set the content therein to 0, and perform the next round of operation.
[0021] Preferably, obtain the pixel calculation matrix and the convolution kernel calculation matrix by passing the feature map of the original image and the convolution kernel through the Img2col operation respectively, including:
[0022] S11. Starting from the upper left corner of the feature map according to the convolution kernel size, obtain a set C of M*M elements covered by the convolution kernel;
[0023] S12. Expand the element set C into a one-dimensional array CTI with a length of M*M;
[0024] S13. Move the convolution kernel in the order from left to right and from top to bottom according to the convolution stride s, and repeat the processes of S11 and S12 to obtain multiple one-dimensional arrays CTI;
[0025] S14. Combine the multiple one-dimensional arrays CTI obtained in step S13 row by row to obtain a matrix CTM, that is, the pixel calculation matrix;
[0026] S15. Expand each convolution kernel WI into columns to obtain multiple one-dimensional arrays WTI, and combine the multiple one-dimensional arrays WTI column by column to obtain a matrix WTM, that is, the convolution kernel calculation matrix.
[0027] Preferably, encode the pixel calculation matrix and the convolution kernel calculation matrix through a CSC encoding module respectively to obtain a first encoding result and a second encoding result, including:
[0028] Traverse the array of the pixel calculation matrix column by column to obtain all non-zero elements, and store all non-zero elements in a one-dimensional encoding array D1 in the order of appearance. Traverse the array of the pixel calculation matrix column by column to obtain the row numbers of the rows where all non-zero elements are located, and store them in a one-dimensional encoding array D2 in the order corresponding to the elements in the one-dimensional encoding array D1. Traverse the array of the pixel calculation matrix column by column to obtain the positions of the first non-zero element in each column starting from the second column in the one-dimensional encoding array D1, and store the information in a one-dimensional encoding array D3, thereby obtaining the first encoding result;
[0029] Traverse the array of the convolution kernel calculation matrix column by column to obtain all non-zero elements, and store all non-zero elements in a one-dimensional encoding array D1 in the order of appearance. Traverse the array of the convolution kernel calculation matrix column by column to obtain the row numbers of the rows where all non-zero elements are located, and store them in a one-dimensional encoding array D2 in the order corresponding to the elements in the one-dimensional encoding array D1. Traverse the array of the convolution kernel calculation matrix column by column to obtain the positions of the first non-zero element in each column starting from the second column in the one-dimensional encoding array D1, and store the information in a one-dimensional encoding array D3, thereby obtaining the second encoding result.
[0030] Preferably, cut the pixel calculation matrix and the convolution kernel calculation matrix into multiple pixel calculation matrix units and multiple convolution kernel calculation matrix units respectively, including:
[0031] According to the number of computing units in the computing cluster, the pixel computing matrix and the convolution kernel computing matrix are respectively cut into multiple pixel computing matrix units and multiple convolution kernel computing matrix units, and each pixel computing matrix unit and the corresponding convolution kernel computing matrix unit are sent to the computing unit, and the multiplication of each pixel computing matrix unit and the corresponding convolution kernel computing matrix unit is calculated through the computing unit.
[0032] The present invention also provides a hardware acceleration device for sparse matrix multiplication operation for a convolutional neural layer, including:
[0033] A matrix conversion structure for obtaining a pixel computing matrix and a convolution kernel computing matrix by respectively performing an Img2col operation on the feature map of the original image and the convolution kernel;
[0034] An encoding generation structure for respectively encoding the pixel computing matrix and the convolution kernel computing matrix through a CSC encoding module to obtain a first encoding result and a second encoding result. The CSC encoding module includes a one-dimensional encoding array D1, a one-dimensional encoding array D2, and a one-dimensional encoding array D3. The one-dimensional encoding array D1 is used to describe the non-zero elements in the sparse matrix from top to bottom and from left to right. The one-dimensional encoding array D2 is used to describe the row numbers corresponding to the non-zero elements, and the one-dimensional encoding array D3 is used to describe the index of the first non-zero element in each column other than the first column in the sparse matrix in the one-dimensional encoding array D1;
[0035] A matrix splitting controller for respectively cutting the pixel computing matrix and the convolution kernel computing matrix into multiple pixel computing matrix units and multiple convolution kernel computing matrix units;
[0036] An operation structure for performing operations on the non-zero elements traversed from the pixel computing matrix unit and the corresponding convolution kernel computing matrix unit based on the first and second encoding results to obtain a matrix unit operation result;
[0037] A merging structure for merging multiple matrix unit operation results to obtain the operation results of the pixel computing matrix and the convolution kernel computing matrix.
[0038] Compared with the prior art, the beneficial effects of the present invention are:
[0039] Through the constructed CSC encoding module, the present invention can quickly obtain non-zero elements from the sparse matrix, the data storage space can be greatly compressed, and at the same time, it helps the computing unit to only traverse non-zero elements, which can greatly reduce the computing time and has a small overhead. Description of the Drawings
[0040] Figure 1 It is a flowchart of a hardware acceleration method for sparse matrix multiplication operation for a convolutional neural layer provided by a specific embodiment of the present invention;
[0041] Figure 2 Schematic diagram of Img2col conversion provided by a specific embodiment of the present invention;
[0042] Figure 3 Schematic diagram of the sparse matrix coding effect provided by a specific embodiment of the present invention;
[0043] Figure 4 Schematic diagram of the matrix splitting and merging effect provided by a specific embodiment of the present invention;
[0044] Figure 5 Schematic diagram of the matrix operation process provided by a specific embodiment of the present invention;
[0045] Figure 6 Schematic diagram of the hardware acceleration device for sparse matrix multiplication operation for convolutional neural layers provided by a specific embodiment of the present invention;
[0046] Figure 7 Schematic diagram of the data flow of the computing unit provided by a specific embodiment of the present invention. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] A specific embodiment of the present invention provides a hardware acceleration method for sparse matrix multiplication operation for convolutional neural layers, as Figure 1 shown, including:
[0049] S1. Respectively obtain a pixel calculation matrix and a convolution kernel calculation matrix by performing Img2col operations on the feature map of the original image and the convolution kernel.
[0050] Specifically, as Figure 2 shown in a and b of
[0051] In a specific embodiment, the method of inputting an original image into Img2col to obtain a pixel calculation matrix includes: The original image extracts an element set from the original image in the order from left to right and from top to bottom according to the size of the convolution kernel, and expands it into a one-dimensional array row by row. Finally, the obtained one-dimensional arrays are integrated into an input matrix row by row in the order of acquisition, that is, the pixel calculation matrix, and the Img2col conversion process is completed.
[0052] Specifically, the feature map of the original image and the convolution kernel are respectively operated by Img2col to obtain a pixel calculation matrix and a convolution kernel calculation matrix, including:
[0053] S11. Starting from the upper left corner of the feature map according to the convolution kernel size, obtain M*M element sets C of the convolution kernel coverage size;
[0054] S12. Expand the element set C into a one-dimensional array CTI with a length of M*M;
[0055] S13. Move the convolution kernel in the order from left to right and from top to bottom according to the convolution step s, and repeat the processes of S11 and S12 to obtain multiple one-dimensional arrays CTI;
[0056] S14. Combine the multiple one-dimensional arrays CTI obtained in step S13 row by row to obtain an input matrix CTM, that is, the pixel calculation matrix;
[0057] S15. Expand each convolution kernel WI into columns to obtain multiple one-dimensional arrays WTI, and combine the multiple one-dimensional arrays WTI column by column to obtain a convolution kernel matrix WTM, that is, the convolution kernel calculation matrix.
[0058] S2. In a specific embodiment of the present invention, the pixel calculation matrix and the convolution kernel calculation matrix are respectively encoded by a CSC encoding module to obtain a first encoding result and a second encoding result. The CSC encoding module includes a one-dimensional encoding array D1, a one-dimensional encoding array D2, and a one-dimensional encoding array D3. The one-dimensional encoding array D1 is used to describe the non-zero elements from top to bottom and from left to right in the sparse matrix. The one-dimensional encoding array D2 is used to describe the row numbers corresponding to the non-zero elements. The one-dimensional encoding array D3 is used to describe the index of the first non-zero element in each column other than the first column in the one-dimensional encoding array D1 of the sparse matrix. Through this encoding method, the non-zero elements of the sparse matrix can be quickly found.
[0059] Compared with storing all matrix information, in the specific embodiments of the present invention, it is necessary to store three vectors, namely the one-dimensional encoded arrays D1, D2, and D3 generated after encoding the sparse matrix. Generally, the number of elements in the one-dimensional encoded array D1 is the number of columns of the original matrix minus 1. The number of elements in the one-dimensional encoded arrays D2 and D3 is related to the number of non-zero elements in the matrix, and the bit width of the elements in the one-dimensional encoded array D2 is generally 4 bits. For a sparse matrix with a sparsity of 40%, the data storage space can be compressed to 50 - 60% of the original.
[0060] In a specific embodiment, in this embodiment, the pixel calculation matrix and the convolution kernel calculation matrix are respectively encoded through the CSC encoding module to obtain the first encoding result and the second encoding result, including:
[0061] In this embodiment, the pixel calculation matrix is traversed column by column to obtain all non-zero elements, and all non-zero elements are stored in the one-dimensional encoded array D1 in the order of appearance. The pixel calculation matrix is traversed column by column to obtain the row numbers of all non-zero elements, and they are stored in the one-dimensional encoded array D2 in the order corresponding to the elements in the one-dimensional encoded array D1. The pixel calculation matrix is traversed column by column to obtain the positions of the first non-zero element in each column starting from the second column in the one-dimensional encoded array D1, and the information is stored in the one-dimensional encoded array D3, thereby obtaining the first encoding result.
[0062] In this embodiment, the convolution kernel calculation matrix is traversed column by column to obtain all non-zero elements, and all non-zero elements are stored in the one-dimensional encoded array D1 in the order of appearance. The convolution kernel calculation matrix is traversed column by column to obtain the row numbers of all non-zero elements, and they are stored in the one-dimensional encoded array D2 in the order corresponding to the elements in the one-dimensional encoded array D1. The convolution kernel calculation matrix is traversed column by column to obtain the positions of the first non-zero element in each column starting from the second column in the one-dimensional encoded array D1, and the information is stored in the one-dimensional encoded array D3, thereby obtaining the second encoding result.
[0063] In one embodiment, as Figure 3 shown, in this embodiment, the non-zero elements start from 1 and end at 9. The data content of the one-dimensional encoded array D1 corresponds to 1 - 9. The one-dimensional encoded array D2 identifies each non-zero element, that is, the row numbers of 1 - 9 in the original matrix. Since the second column is all 0 columns in the one-dimensional encoded array D3, the first element is 31. The 4-bit binary number represents D3, and the maximum value is 31. 31 does not represent a meaningful number and indicates an empty column. The subsequent data represents the index values of the first non-zero element in each subsequent column in the one-dimensional encoded array D1.
[0064] S3. In specific embodiments of the present invention, the pixel calculation matrix and the convolution kernel calculation matrix are respectively cut into multiple pixel calculation matrix units and multiple convolution kernel calculation matrix units, and matrix unit operation results are obtained by performing operations on non-zero elements traversed from the pixel calculation matrix units and the corresponding convolution kernel calculation matrix units based on the first and second encoding results.
[0065] Specifically, according to the number of calculation units in the calculation cluster, the pixel calculation matrix and the convolution kernel calculation matrix are respectively cut into multiple pixel calculation matrix units and multiple convolution kernel calculation matrix units, and each pixel calculation matrix unit and convolution kernel calculation matrix unit are distributed to the calculation units, and the multiplication of each pixel calculation matrix unit and the corresponding convolution kernel calculation matrix unit is calculated by the calculation unit.
[0066] Specifically, in specific embodiments of the present invention, matrix unit operation results are obtained by performing operations on non-zero elements traversed from the pixel calculation matrix units and the corresponding convolution kernel calculation matrix units based on the first and second encoding results, including:
[0067] S31. Initialize the partial sum storage space, which is used to temporarily store intermediate results of the operations. The intermediate results include the operation results of the calculation units and the intermediate results during the operations of the calculation units.
[0068] S32. Traverse to the non-zero element B(m, n) in the first non-empty column of the convolution kernel calculation matrix through the one-dimensional encoding arrays D3 and D1 of the second encoding result, that is, the element in the m-th row and the n-th column, and record the column number colNum.
[0069] S33. Traverse to the non-zero element in the first non-empty column of the pixel calculation matrix through the one-dimensional encoding arrays D3 and D1 of the first encoding result, and then traverse the non-zero elements of the pixel calculation matrix through the one-dimensional encoding array D1 to obtain the element A(k, m) whose column number is equal to m, that is, the element in the m-th row and the k-th column.
[0070] S34. Traverse to the row number of the element A(k, m) in the pixel calculation matrix through the one-dimensional encoding array D2 of the first encoding result, record the row number rowNum, and calculate the storage location site through the recorded column number colNum and row number rowNum, that is, calculate the storage location site =
[0071] rowNum + colNum * n, where n is the number of rows of the pixel calculation matrix. Since the storage location of the operation result is one-dimensional in the physical space, it is necessary to convert it into one-dimensional data through the above formula to obtain the storage location.
[0072] S35. Access the partial sum storage space at site and the value stored therein, add the value to the calculation result of A(k, m)×B(m, n), and write the result back to the partial sum storage space at site.
[0073] S36. Traverse the non-zero elements in the pixel calculation matrix where the next column number is equal to m according to the D3 and D1 coding arrays of the first coding result, and repeat steps S34 - S36 until the traversal of the pixel calculation matrix is completed.
[0074] S37. Traverse to the position of the non-zero element in the next non-empty column of the convolution kernel calculation matrix through the D3 and D1 coding arrays of the second coding result, and record the column number of the non-zero element in the next non-empty column. Repeat steps S33 - S36 until the traversal of the convolution kernel calculation matrix is completed to obtain the operation result of the matrix unit. After the operation is completed, send a Cal_fin signal to the overall control unit to indicate that the multiplication operation of the current calculation unit is completed. After the control unit receives the multiplication operation results of all calculation units, it starts the addition operation, adds up all the multiplication operation results to obtain the multiplication operation result of the pixel calculation matrix and the convolution kernel calculation matrix, initializes the partial sum storage space, sets the content therein to 0, and starts a new round of operations.
[0075] Specifically, in the specific embodiments of the present invention, the operation results of multiple matrix units are combined to obtain the operation results of the pixel calculation matrix and the convolution kernel calculation matrix, including:
[0076] After the controller receives the signal that the operation results of the pixel calculation matrix units and the corresponding convolution kernel calculation matrix units in all calculation units are stored in the partial sum storage space, that is, after receiving all Cal_fin signals, perform matrix addition on the operation results of the calculation units to obtain the operation results of the output pixel calculation matrix and the convolution kernel calculation matrix, initialize the partial sum storage space, set the content therein to 0, and perform the next round of operations.
[0077] In one embodiment, as Figure 4 shown, when multiplying matrix A(3*3) and matrix B(3*1), it is necessary to split the matrix according to the number of calculation units in the calculation array and merge the intermediate results into the final result. At the first moment, the three values A11, A12, and A13 are multiplied by B11, B21, and B31 respectively, and then the three results are accumulated. At the second moment, the three values A21, A22, and A23 are multiplied and accumulated with B11, B21, and B31 respectively. At the third moment, A31, A32, and A33 are multiplied and accumulated with B11, B21, and B31 respectively. Thus, the multiplication of matrix A and matrix B is completed. Figure 4In the example, since the number of columns of matrix A and the number of rows of matrix B are only 3, the multiply-accumulate operation of 3 data at a time can be completed simultaneously in the operation array. If a large matrix is split, the results of the operation units need to be used as partial sums and accumulated multiple times to obtain the final result.
[0078] Among them, as Figure 5 shown, when performing sparse matrix multiplication, only the multiplication of non-zero data needs to be performed. Therefore, the positions of non-zero data in matrix A and matrix B need to be recorded, and only then can the positions in the result matrix C be determined. Therefore, it is necessary to rely on Figure 3 the 3 vectors in the sparse matrix encoding method mentioned to determine the position of the operation result in the matrix.
[0079] Compared with traversing all matrix data for calculation, the setting of encoding provided by the specific embodiment of the present invention can help the calculation unit to only traverse non-zero elements. For a sparse matrix with a sparsity rate of 40%, the calculation time can be reduced by more than 60%.
[0080] The specific embodiment of the present invention also provides a hardware acceleration device for sparse matrix multiplication operation for a convolutional neural layer, as Figure 6 shown, including a matrix conversion structure T1, an encoding generation structure T2, a matrix segmentation controller P1, an operation structure C1, and a merging structure N1.
[0081] The matrix conversion structure provided in this embodiment is used to obtain a pixel calculation matrix and a convolution kernel calculation matrix from the feature map of the original image and the convolution kernel respectively through the Img2col operation.
[0082] The encoding generation structure provided in this embodiment is used to encode the pixel calculation matrix and the convolution kernel calculation matrix respectively through the CSC encoding module to obtain a first encoding result and a second encoding result. The CSC encoding module includes a one-dimensional encoding array D1, a one-dimensional encoding array D2, and a one-dimensional encoding array D3. The one-dimensional encoding array D1 is used to describe the non-0 elements in the sparse matrix from top to bottom and from left to right. The one-dimensional encoding array D2 is used to describe the row numbers corresponding to the non-0 elements. The one-dimensional encoding array D3 is used to describe the index of the first non-zero element in each column other than the first column in the sparse matrix in the one-dimensional encoding array D1.
[0083] The matrix segmentation controller provided in this embodiment is used to cut the pixel calculation matrix and the convolution kernel calculation matrix into multiple pixel calculation matrix units and multiple convolution kernel calculation matrix units respectively.
[0084] The operation structure provided in this embodiment is used to perform operations on the non-zero elements traversed from the pixel calculation matrix unit and the corresponding convolution kernel calculation matrix unit based on the first and second encoding results to obtain a matrix unit operation result.
[0085] The merging structure provided in this embodiment is used to merge the operation results of multiple matrix units to obtain the operation results of the pixel calculation matrix and the convolution kernel calculation matrix.
[0086] As Figure 7 shown, a dedicated sparse matrix multiplication operation unit is proposed for the sparse matrix non-zero data encoding method in Figure 3 . In the sparse matrix encoding, each non-zero data corresponds to an element in the D1 and D2 vectors. Therefore, the D1 and D2 vectors are concatenated and sent together to the local memory of the operation unit. In the dedicated operation unit, there are a multiply-accumulate operation unit and a position index unit. The multiply-accumulate operation unit calculates the size of the partial sum result through the values of the non-zero data of the two matrices, and the position index unit determines the position index of the calculation result through the position index of the non-zero data. In addition, there are 5 memories to implement near-memory computing. The 5 memories are the concatenated vector memory of matrix A, the D3 vector memory of matrix A, the concatenated vector memory of matrix B, the D3 vector memory of matrix B, and the partial sum result memory. Matrix A is set as the weight matrix and matrix B is set as the input data matrix. According to the operation data stream in the convolution operation, the capacity of the concatenated vector memory of matrix B needs to be set larger, so this memory is implemented using SRAM, and the other 4 memories have smaller storage capacities and are set as register files for data storage. A Booth multiplier is used in the multiply-accumulate operation unit to complete the matrix multiplication operation. The position index unit identifies the D2 and D3 vectors to determine the row index and column index of the non-zero value in the sparse matrix, and determines the row index and column index of the operation result in the result matrix. After the large matrix multiplication is divided into small matrix multiplications, the lower operation unit will pass the partial sum result upward. Therefore, an adder is set in the operation unit to complete the accumulation of the partial sum result. A FIFO is provided for data caching in the data input stage, and the data stream control is implemented through the finite state machine in the control module inside the operation unit.
Claims
1. A hardware acceleration method for sparse matrix multiplication operations in convolutional neural layers, characterized in that: include: The original image and the convolution kernel are respectively subjected to the Img2col operation to obtain a pixel calculation matrix and a convolution kernel calculation matrix, and the pixel calculation matrix and the convolution kernel calculation matrix are respectively encoded by the CSC encoding module to obtain a first encoding result and a second encoding result, wherein the CSC encoding module includes a one-dimensional encoding array D1, a one-dimensional encoding array D2 and a one-dimensional encoding array D3, wherein the one-dimensional encoding array D1 is used to describe the non-zero elements from top to bottom and from left to right in the sparse matrix, the one-dimensional encoding array D2 is used to describe the row numbers corresponding to the non-zero elements, and the one-dimensional encoding array D3 is used to describe the index of the first non-zero element of each column other than the first column in the sparse matrix in the one-dimensional encoding array D1; The pixel calculation matrix and the convolution kernel calculation matrix are respectively cut into multiple pixel calculation matrix units and multiple convolution kernel calculation matrix units, and based on the first and second encoding results, the non-zero elements traversed from the pixel calculation matrix unit and the corresponding convolution kernel calculation matrix unit are operated to obtain the matrix unit operation results, and the multiple matrix unit operation results are merged to obtain the pixel calculation matrix and the convolution kernel calculation matrix operation results.
2. The hardware acceleration method for sparse matrix multiplication operations for convolutional neural layers according to claim 1, characterized in that: Based on the first and second encoding results, the non-zero elements traversed from the pixel calculation matrix unit and the corresponding convolution kernel calculation matrix unit are operated to obtain the matrix unit operation result, including: S31, initialization part and storage space; S32, traverse the non-zero element B(m, n) of the first non-empty column of the convolution kernel calculation matrix through the one-dimensional encoding arrays D3 and D1 of the second encoding result, and record the column number colNum; S33, traversing the non-0 elements of the first non-empty column of the pixel calculation matrix through the one-dimensional encoding arrays D3 and D1 of the first encoding result, and then traversing the non-0 elements of the pixel calculation matrix through the one-dimensional encoding array D1 to obtain an element A(k, m) whose column number is equal to m; S34, traverse the pixel calculation matrix through the one-dimensional encoding array D2 of the first encoding result to obtain the row number of the element A(k, m), record the row number rowNum, and calculate the storage location site through the recorded column number colNum and row number rowNum; S35, access the value of the portion and storage space of the address site, add the value to the calculation result of A(k, m)×B(m, n), and write it back to the portion and storage space of the storage location site; S36, traversing the D3 and D1 encoding arrays of the first encoding result to the next non-zero element of the pixel calculation matrix whose column number is equal to m, repeating steps S34-S36 until the pixel calculation matrix traversal is completed; S37. Traverse the D3 and D1 encoding arrays of the second encoding result to the position of the non-0 element of the next non-empty column of the convolution kernel calculation matrix, and record the column number of the non-0 element of the next non-empty column, repeat steps S33-S36 until the convolution kernel calculation matrix traversal is completed to obtain the matrix unit operation result, and the operation is completed.
3. The hardware acceleration method for sparse matrix multiplication operation for convolutional neural layer according to claim 2, characterized in that: The storage location site=rowNum+colNum*n is calculated, where n is the number of rows in the pixel calculation matrix.
4. The hardware acceleration method for sparse matrix multiplication operation for convolutional neural layer according to claim 1, characterized in that: The operation results of multiple matrix units are combined to obtain the operation results of the pixel calculation matrix and the convolution kernel calculation matrix, including: After the calculation results of all pixel calculation matrix units and the corresponding convolution kernel calculation matrix units are stored in the partial sum storage space, the sum is obtained by adding the calculation results of the pixel calculation matrix and the convolution kernel calculation matrix. After the calculation results of the pixel calculation matrix and the convolution kernel calculation matrix are output, the partial sum storage space is initialized and the contents therein are set to 0 for the next round of calculation.
5. The hardware acceleration method for sparse matrix multiplication operation for convolutional neural layer according to claim 1, characterized in that: The feature map and convolution kernel of the original image are respectively subjected to the Img2col operation to obtain the pixel calculation matrix and the convolution kernel calculation matrix, including: S11. Starting from the upper left corner of the feature map according to the convolution kernel size, obtain a set C of M*M elements of the convolution kernel coverage size; S12, expanding the element set C into a one-dimensional array CTI with a length of M*M; S13, moving the convolution kernel from left to right and from top to bottom according to the convolution step s, repeating the processes of S11 and S12 to obtain multiple one-dimensional arrays CTI; S14, merging the multiple one-dimensional arrays CTI obtained in step S13 row by row to obtain a matrix CTM, i.e., a pixel calculation matrix; S15. Expand each convolution kernel WI into columns to obtain multiple one-dimensional arrays WTI. Merge the multiple one-dimensional arrays WTI by columns to obtain a matrix WTM, that is, the convolution kernel calculation matrix.
6. The hardware acceleration method for sparse matrix multiplication operation for convolutional neural layer according to claim 1, characterized in that: The pixel calculation matrix and the convolution kernel calculation matrix are respectively encoded by the CSC encoding module to obtain a first encoding result and a second encoding result, including: Traverse the array by columns for the pixel calculation matrix to obtain all non-zero elements, and store all non-zero elements in the one-dimensional coding array D1 in the order of appearance; traverse the array by columns for the pixel calculation matrix to obtain the row numbers of the rows where all non-zero elements are located, and store them in the one-dimensional coding array D2 in the order corresponding to the elements in the one-dimensional coding array D1; traverse the array by columns for the pixel calculation matrix to obtain the position of the first non-zero element in each column starting from the second column in the one-dimensional coding array D1, and store the information in the one-dimensional coding array D3, thereby obtaining a first coding result; The convolution kernel calculation matrix is traversed through the array column by column to obtain all non-zero elements, and all non-zero elements are stored in the one-dimensional coding array D1 in the order of appearance. The convolution kernel calculation matrix is traversed through the array column by column to obtain the row numbers of the rows where all non-zero elements are located, and stored in the one-dimensional coding array D2 in the order corresponding to the elements in the one-dimensional coding array D1. The convolution kernel calculation matrix is traversed through the array column by column to obtain the position of the first non-zero element in each column starting from the second column in the one-dimensional coding array D1, and the information is stored in the one-dimensional coding array D3, thereby obtaining the second coding result.
7. The hardware acceleration method for sparse matrix multiplication operations for convolutional neural layers according to claim 1, characterized in that: The pixel calculation matrix and the convolution kernel calculation matrix are respectively cut into a plurality of pixel calculation matrix units and a plurality of convolution kernel calculation matrix units, including: According to the number of computing units in the computing cluster, the pixel computing matrix and the convolution kernel computing matrix are cut into multiple pixel computing matrix units and multiple convolution kernel computing matrix units respectively, and each pixel computing matrix unit and the corresponding convolution kernel computing matrix unit are sent to the computing unit, and the multiplication of each pixel computing matrix unit and the corresponding convolution kernel computing matrix unit is calculated by the computing unit.
8. A hardware acceleration device for sparse matrix multiplication operations for convolutional neural layers, characterized in that: include: The matrix conversion structure is used to obtain the pixel calculation matrix and the convolution kernel calculation matrix of the original image through the Img2col operation respectively; A coding generation structure, used for respectively encoding a pixel calculation matrix and a convolution kernel calculation matrix through a CSC coding module to obtain a first coding result and a second coding result, wherein the CSC coding module includes a one-dimensional coding array D1, a one-dimensional coding array D2 and a one-dimensional coding array D3, wherein the one-dimensional coding array D1 is used to describe non-zero elements from top to bottom and from left to right in a sparse matrix, the one-dimensional coding array D2 is used to describe row numbers corresponding to non-zero elements, and the one-dimensional coding array D3 is used to describe the index of the first non-zero element of each column other than the first column in the sparse matrix in the one-dimensional coding array D1; A matrix splitting controller, used for splitting the pixel calculation matrix and the convolution kernel calculation matrix into a plurality of pixel calculation matrix units and a plurality of convolution kernel calculation matrix units respectively; An operation structure, used for operating the non-zero elements traversed from the pixel calculation matrix unit and the corresponding convolution kernel calculation matrix unit based on the first and second encoding results to obtain a matrix unit operation result; The merging structure is used to merge the operation results of multiple matrix units to obtain the operation results of the pixel calculation matrix and the convolution kernel calculation matrix.
Citation Information
Patent Citations
Data processing method and device, electronic equipment and storage medium
CN117785031A
Neural network data operation acceleration device and method, equipment and medium
CN118966293A