Neural network model sparse calculation method and device

By sparse the weight of the convolutional layer of the deep neural network model and divide the input matrix into operator matrix, the problem of low computing efficiency and inability to meet the real-time requirements in the existing technology is solved, and efficient sparse computing performance is achieved.

CN119990199APending Publication Date: 2025-05-13HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311503148.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

After the scale of the existing deep neural network model is expanded, it is difficult to efficiently process using existing accelerators, and the existing sparse method still requires more parameters to participate in complex matrix operations, which cannot meet the real-time computing requirements.

Method used

By sparse the weights of the convolution layer of the neural network model and divide the input matrix into multiple operator matrices, the column index of the weight parameters in each row of the weight matrix is ​​used to obtain the corresponding rows from the operator matrix, and simple vector calculation and accumulation operations are performed to complete the sparse operation of the neural network.

Benefits of technology

It effectively reduces the storage space required by the model, reduces the demand for hardware computing resources, improves the computing speed of sparse computing in neural networks, can meet the needs of real-time computing, and reduces the situation of cache misses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990199A_ABST
    Figure CN119990199A_ABST
Patent Text Reader

Abstract

The invention provides a neural network model rarefaction calculation method and device, and the method comprises the steps: carrying out the weight rarefaction of each convolution layer of a neural network model, so as to obtain a rarefaction neural network model; converting weight parameters of each convolutional layer of the sparse neural network model into a weight matrix, and converting input data into an input matrix; dividing the input matrix into a plurality of operation sub-matrixes according to the cache capacity; acquiring a column index of a weight parameter in each row of the weight matrix, and acquiring a corresponding row from the operation sub-matrix according to the column index; multiplying each weight parameter in the same row by a corresponding row in the operation sub-matrix to obtain a one-dimensional to-be-processed matrix, and accumulating the one-dimensional to-be-processed matrix corresponding to each weight parameter in the same row to obtain a one-dimensional intermediate matrix; and determining a convolution result according to the weight parameter of each row and the one-dimensional intermediate matrix obtained by each operation sub-matrix. According to the neural network model sparse calculation method provided by the invention, the sparse calculation performance can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to a method and device for calculating sparseness of a neural network model. Background Art

[0002] As the scale of deep convolutional neural networks increases, existing accelerators cannot efficiently process the increasingly large deep neural networks. Sparsification of neural network models is one of the ways to solve this problem. Studies have shown that sparse neural networks can effectively reduce redundant weights in neural network parameters without losing model accuracy. Many research works optimize existing neural network models through pruning, quantization, and compression. However, existing sparsification methods usually rely on sparse encoding of input data or converting sparse convolutions into regular convolutions for processing. This processing method still requires more parameters to participate in complex matrix operations or complex logical judgments, and still has high requirements for hardware resources, which cannot meet real-time computing requirements. Summary of the invention

[0003] The neural network model sparse calculation method provided by the present invention can effectively improve the sparse calculation performance.

[0004] The present invention provides a neural network model sparse calculation method, the method comprising:

[0005] The weights of each convolutional layer of the neural network model are sparsed to obtain a sparse neural network model;

[0006] Convert the weight parameters of each convolutional layer of the sparse neural network model into a weight matrix, and convert the input data into an input matrix;

[0007] According to the cache capacity, the input matrix is ​​divided into a plurality of operator matrices;

[0008] Obtaining a column index of a weight parameter in each row of the weight matrix, and obtaining a corresponding row from the operator matrix according to the column index;

[0009] Multiply each weight parameter in the same row by the corresponding row in the operator matrix to obtain a one-dimensional matrix to be processed, and accumulate the one-dimensional matrix to be processed corresponding to each weight parameter in the same row to obtain a one-dimensional intermediate matrix;

[0010] The convolution result is determined based on the one-dimensional intermediate matrix obtained by each row of weight parameters and each operator matrix.

[0011] Optionally, the input matrix is ​​divided into a plurality of operator matrices according to the cache capacity, including:

[0012] According to the cache capacity, the maximum number of columns of the operator matrix is ​​determined to be n columns; wherein n≥1, and the cache capacity required for n columns does not exceed the maximum capacity of the cache that can currently be used to store the matrix.

[0013] Optionally, after determining the maximum number of columns of the operator matrix to be n columns according to the cache capacity, the method further includes:

[0014] When the number of columns of the input matrix is ​​a non-integer multiple of n, the input matrix is ​​divided into an integer number of operator matrices, and the remaining columns less than n are divided according to The columns are recursively divided into multiple operator matrices in descending order; where k ≥ 1;

[0015] When the number of columns of the input matrix is ​​an integer multiple of n, the input matrix is ​​divided into an integer number of operator matrices.

[0016] Optionally, performing weight sparsification on each convolutional layer of the neural network model to obtain a sparse neural network model includes:

[0017] According to the first weight parameter threshold, the convolution layer of the neural network model is subjected to sparse processing to form a first intermediate model;

[0018] According to a preset first clipping threshold, clipping the channels in the convolutional layer of the first intermediate model to form a second intermediate model;

[0019] According to a preset first judgment threshold, the convolution mode of each convolution layer of the second intermediate model is determined and labeled to form a sparse neural network model.

[0020] Optionally, the step of performing a sparse processing on the convolution layer of the neural network model according to the first weight parameter threshold to form a first intermediate model includes:

[0021] Counting the weight parameter distribution histogram in the convolutional layer of the neural network model;

[0022] Determining a first weight parameter threshold value according to a preset first model clipping ratio and the weight parameter distribution histogram;

[0023] The weight parameters in the convolution layer of the neural network model that are lower than the first weight parameter threshold are set to zero to complete the sparsification process.

[0024] Optionally, the step of trimming channels in a convolutional layer of the first intermediate model according to a preset first trimming threshold to form a second intermediate model includes:

[0025] Comparing the sparsity rate of each channel in the convolutional layer of the first intermediate model with the first clipping threshold;

[0026] When the sparsity rate of the channel is greater than the first clipping threshold, the channel is removed so that the channel does not participate in the convolution calculation.

[0027] Optionally, before comparing the sparsity rate of each channel in the convolutional layer of the first intermediate model with the first clipping threshold, the method further includes:

[0028] The weight parameters of each channel in the convolution layer of the first intermediate model are subjected to sparse rate statistics, and a first trimming threshold is determined according to the sparse rate statistics result.

[0029] Optionally, judging the convolution mode of each convolution layer of the second intermediate model and marking it according to a preset first judgment threshold to form a sparse neural network model includes:

[0030] Comparing the sparsity rate of each convolutional layer of the second intermediate model with the first judgment threshold;

[0031] When the sparsity rate of the convolution layer is greater than the first judgment threshold, marking the convolution layer as sparse calculation;

[0032] When the sparsity rate of the convolution layer is not greater than the first judgment threshold, the convolution layer is marked as dense calculation.

[0033] Optionally, after judging the convolution mode of each convolution layer of the second intermediate model according to the preset first judgment threshold and marking it, the method further includes:

[0034] For the convolutional layer marked as sparse computing, the performance comparison between sparse computing and dense computing is performed;

[0035] When the dense computing performance of the convolutional layer is higher than the sparse computing performance, the computing mode of the convolutional layer is transformed into dense computing and marked.

[0036] Optionally, after trimming the channels in the convolution layer of the first intermediate model according to the preset first trimming threshold to form the second intermediate model, it also includes: numerically compressing and storing the second intermediate model in an indexed storage form.

[0037] In a second aspect, the present invention provides a neural network model sparse calculation device, comprising:

[0038] The sparse module is used to sparse the weights of each convolutional layer of the neural network model to obtain a sparse neural network model;

[0039] A conversion module, used to convert the weight parameters of each convolutional layer of the sparse neural network model into a weight matrix, and convert the input data into an input matrix;

[0040] A partitioning module, used for partitioning the input matrix into a plurality of operator matrices according to a cache capacity;

[0041] An acquisition module, used to acquire a column index of a weight parameter in each row of the weight matrix, and acquire a corresponding row from the operator matrix according to the column index;

[0042] An operation module, used for multiplying each weight parameter in the same row with the corresponding row in the operator matrix to obtain a one-dimensional matrix to be processed, and accumulating the one-dimensional matrix to be processed corresponding to each weight parameter in the same row to obtain a one-dimensional intermediate matrix;

[0043] The result module is used to determine the convolution result based on the one-dimensional intermediate matrix obtained by each row of weight parameters and each operator matrix.

[0044] The neural network model sparse calculation method provided by the present invention can effectively reduce the storage space required by the model by sparsely weighting the weights. In the process of calculation, by dividing the input matrix into multiple operator matrices, only simple vector calculation and accumulation operation are required to complete the sparse calculation of the neural network, which greatly reduces the demand for hardware computing resources, improves the calculation speed of the neural network sparse calculation, and can effectively meet the needs of real-time calculation. At the same time, since the demand for hardware cache resources of the operator matrix is ​​greatly reduced, the occurrence of cache misses can be effectively reduced. The technical solution provided by the present invention can greatly improve the speed of real-time calculation under lower hardware requirements, avoiding the need to increase hardware computing resources and hardware cache resources separately for neural network calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A flowchart of a neural network model sparsification calculation method according to an embodiment of the present invention;

[0046] Figure 2 A flowchart of partitioning operator matrices of a neural network model sparse calculation method according to another embodiment of the present invention;

[0047] Figure 3 A schematic diagram of a partitioning operator matrix of a neural network model sparse calculation method according to another embodiment of the present invention;

[0048] Figure 4 A schematic diagram of a calculation method for a weight matrix and an operator matrix of a neural network model sparsification calculation method according to another embodiment of the present invention;

[0049] Figure 5A flowchart of weight sparseness of a neural network model sparseness calculation method according to another embodiment of the present invention;

[0050] Figure 6 A flowchart of a convolutional layer sparse processing method for a neural network model sparse calculation method according to another embodiment of the present invention;

[0051] Figure 7 A flowchart of a convolutional layer channel pruning process of a neural network model sparse calculation method according to another embodiment of the present invention;

[0052] Figure 8 A flowchart of determining a convolutional layer calculation method in a neural network model sparsification calculation method according to another embodiment of the present invention;

[0053] Fig. 9 The present invention is a flowchart for determining a convolutional layer calculation method for sparse calculation of a neural network model according to another embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0055] The present invention provides a neural network model sparse calculation method, such as Figure 1 As shown, the method includes:

[0056] Step 100, performing weight sparsification on each convolutional layer of the neural network model to obtain a sparse neural network model;

[0057] In some embodiments, weight sparsification of each convolutional layer of the neural network model refers to screening the weight parameters in each convolutional layer of the neural network model according to pre-set rules, and setting the screened weight parameters to zero. Preferably, each convolutional layer of the neural network model can be further pruned to obtain a sparse neural network model.

[0058] Step 200, converting the weight parameters of each convolutional layer of the sparse neural network model into a weight matrix, and converting the input data into an input matrix;

[0059] In some embodiments, converting the weight parameters of each convolution layer into a weight matrix means that the convolution kernel in each convolution layer is set as a one-dimensional row matrix, and the one-dimensional row matrices corresponding to multiple output channels are arranged in the corresponding order of the channels to form a weight matrix. Converting the input data into an input matrix means that the area corresponding to each sliding of the convolution kernel of the input data is set as a one-dimensional column matrix, and then the multiple one-dimensional column matrices corresponding to each sliding of the convolution kernel are arranged in the corresponding order to form an input matrix. For example, the parameters of the input data are height Ih, width Iw, input channel Ic, the convolution kernel parameters are height Kh, width Kw, padding Kp (pad), sliding step S (strid), the number of output channels is oc, and the weight matrix A formed is M*K, where M=oc, K=Kh*Kw*ic; the input matrix B formed is K*N, where K=Kh*Kw*Ic, N=((Ih-Kh+2Kp) / S+1)*((Iw-Kw+2Kp) / S+1).

[0060] Step 300, dividing the input matrix into a plurality of operator matrices according to the cache capacity;

[0061] In some embodiments, since the input matrix size is usually large, in order to avoid the problem of cache misses caused by loading all at once, the input matrix is ​​divided into multiple operator matrices, thereby effectively avoiding performance degradation caused by frequent cache misses. Step 400, obtaining the column index of the weight parameter in each row of the weight matrix, and obtaining the corresponding row from the operator matrix according to the column index;

[0062] In some embodiments, during matrix operations, in each row of the weight matrix, each parameter corresponds to a row of parameters of the operator matrix. Through the column index of the weight parameter, all the corresponding rows in the operator matrix are obtained at one time, thereby forming an operation method of numerical values ​​and row matrices, rather than using matrix operations, which can effectively reduce the resources required for operations.

[0063] Step 500, multiplying each weight parameter in the same row with the corresponding row in the operator matrix to obtain a one-dimensional matrix to be processed, and accumulating the one-dimensional matrix to be processed corresponding to each weight parameter in the same row to obtain a one-dimensional intermediate matrix;

[0064] In some embodiments, a numerical value is multiplied by a row matrix, and a single parameter in a row of the weight matrix is ​​multiplied by the corresponding row to form a one-dimensional matrix to be processed. Then, the one-dimensional matrix to be processed corresponding to each parameter in a row is accumulated by the row matrix, thereby completing the multiplication of a single row of the weight matrix and the entire operator matrix. This processing method only involves the multiplication of a numerical value by a row matrix and the accumulation calculation of a single row matrix. The calculation method is simple, requires fewer resources, and can effectively improve the computing performance.

[0065] Step 600, determining the convolution result according to the one-dimensional intermediate matrix obtained by each row of weight parameters and each operator matrix.

[0066] In some embodiments, by arranging the one-dimensional intermediate matrix corresponding to each row of weight parameters in the column direction, the convolution sub-result corresponding to the operator matrix can be obtained, and then the convolution sub-results corresponding to multiple operator matrices are arranged in the segmentation order to obtain the convolution result.

[0067] The sparse calculation method of the neural network model provided by the embodiment of the present invention can effectively reduce the storage space required for the model by sparsely calculating the weights. In the calculation process, by dividing the input matrix into multiple operator matrices, only simple vector calculations and accumulation operations are required to complete the sparse calculation of the neural network, which greatly reduces the demand for hardware computing resources, improves the calculation speed of the neural network sparse calculation, and can effectively meet the needs of real-time calculation. At the same time, since the demand for hardware cache resources of the operator matrix is ​​greatly reduced, the occurrence of cache misses can be effectively reduced. The technical solution provided by the embodiment of the present invention can greatly improve the speed of real-time calculation under lower hardware requirements, avoiding the need to increase hardware computing resources and hardware cache resources separately for neural network calculations.

[0068] As an optional implementation, Figure 2 As shown, in step 300, the input matrix is ​​divided into a plurality of operator matrices according to the cache capacity, including:

[0069] Step 310, according to the cache capacity, determine the maximum number of columns of the operator matrix to be n columns; wherein n≥1, and the cache capacity required for n columns does not exceed the maximum capacity of the cache currently available for storing the matrix.

[0070] In some embodiments, since the size of the input matrix is ​​large, in order to avoid cache overflow, resulting in cache misses during the calculation process and thus causing a decrease in computing performance, the operator matrix is ​​divided so that each operator matrix can be loaded into the cache for calculation.

[0071] As an optional implementation, continue as Figure 2 As shown, in step 310, after determining the maximum number of columns of the operator matrix to be n columns according to the cache capacity, the following steps are performed:

[0072] Step 311: when the number of columns of the input matrix is ​​a non-integer multiple of n, the input matrix is ​​divided into an integer number of operator matrices, and the remaining columns less than n are divided into The columns are recursively divided into multiple operator matrices in descending order; where k ≥ 1;

[0073] In some embodiments, since the number of columns of the input matrix may not be an integer multiple of n, when the remaining columns are less than n, the remaining columns are recursively divided, which is conducive to SIMD instruction acceleration. For example, when n is 48, when the remaining columns are less than 48, the remaining columns can be divided according to the number of columns of 32, 16, 8, 4, 2, and 1.

[0074] Step 312: When the number of columns of the input matrix is ​​an integer multiple of n, divide the input matrix into an integer number of operator matrices.

[0075] In some embodiments, when the number of columns of the input matrix is ​​an integer multiple of n, directly dividing the data matrix into an integer number of operator matrices can effectively reduce the number of calculations.

[0076] like Figure 2-3 As shown, the division method of the operator matrix of the present invention and the method of participating in the calculation in the matrix are exemplarily shown, which are as follows:

[0077] like Figure 2 As shown, the parameters of the input data are height Ih, width Iw, input channel Ic, the parameters of the convolution kernel are height Kh, width Kw, padding Kp (pad), sliding step S (strid), the number of output channels is oc, and the weight matrix A formed is M*K, where M=oc, K=Kh*Kw*ic; the input matrix B formed is K*N, where K=Kh*Kw*Ic, N=((Ih-Kh+2Kp) / S+1)*((Iw-Kw+2Kp) / S+1). When dividing the operator matrix, the matrix B is divided into blocks in the N (column direction) dimension, and the block size is the maximum number of columns is n columns. The size of n can be determined according to the cache size (for example, 48). N is divided into n blocks, and the remainder columns form blocks of 32, 16, 8, 4, 2, and 1 respectively, which is convenient for SIMD instruction acceleration within the block. In the overall calculation process, the weight matrix is ​​calculated row by row along the M dimension, and the input matrix N dimension is calculated according to the above block calculation, and finally it will be converted into a series of small matrix calculations such as 1*K and K*n, 1*K and K*32. The internal calculation method of the small matrix is ​​as follows Figure 3 As shown, loop traverse along the K dimension of the weight matrix to obtain the Wj weight, determine the corresponding data row in the corresponding operator matrix according to the weight parameter index j, and load n columns of data in the corresponding data row. The loaded n columns of data are multiplied and accumulated with Wj to obtain a one-dimensional intermediate matrix, that is, the K dimension traversal is completed to obtain the corresponding output features of one row of the weight matrix. After completing the N dimension and oc dimension loop, the calculation result of the entire convolution can be obtained. Figure 2 and Figure 3In the example shown, the calculation process when the convolution kernel is Kh=1, Kw=1, pad=0, strid=1, at this time, the weight matrix is ​​M*K: M=oc, K=ic, and the input matrix is ​​K*N: K=ic, N=ih*iw.

[0078] As an optional implementation, Figure 5 As shown, in step 100, the weights of each convolutional layer of the neural network model are sparsely distributed to obtain a sparse neural network model, including:

[0079] Step 110, performing a sparse processing on the convolution layer of the neural network model according to the first weight parameter threshold to form a first intermediate model;

[0080] In some embodiments, the sparsification process refers to comparing the weight parameter in the convolution layer of the neural network model with a first weight parameter threshold. When the weight parameter is lower than the first weight parameter threshold, it indicates that the weight parameter plays a very small role in the entire convolution process. At this time, the weight parameter can be directly set to zero.

[0081] Step 120, trimming the channels in the convolutional layer of the first intermediate model according to a preset first trimming threshold to form a second intermediate model;

[0082] In some embodiments, clipping means that when the sparsity rate in a channel in the convolution layer is greater than a first clipping threshold, it indicates that the channel plays a very small role in the convolution process, and therefore, the channel can be directly clipped. In some preferred embodiments, the preset first clipping threshold can be, for example, above 90%.

[0083] Step 130, based on a preset first judgment threshold, determine the convolution mode of each convolution layer of the second intermediate model and mark it to form a sparse neural network model.

[0084] In some embodiments, after completing the aforementioned processing, the sparsity rate of each convolution layer is compared with the first judgment threshold. When the sparsity rate is greater than the first judgment threshold, it is determined to be a sparse computing mode. When the sparsity rate is less than the first judgment threshold, it is determined to be a dense computing mode. In some preferred embodiments, the preset judgment threshold may be, for example, 65%±5%.

[0085] As an optional implementation, Figure 6 As shown, in step 110, the convolution layer of the neural network model is subjected to sparse processing according to the first weight parameter threshold to form a first intermediate model, including:

[0086] Step 111, counting the weight parameter distribution histogram in the convolutional layer of the neural network model;

[0087] In some embodiments, the weight parameter distribution histogram refers to counting weight parameters with the same value and representing them in the form of a histogram. The weight distribution histogram can represent the number of weight parameters corresponding to each value.

[0088] Step 112, determining a first weight parameter threshold value according to a preset first model clipping ratio and the weight parameter distribution histogram;

[0089] In some embodiments, since the weight parameter distribution histogram characterizes the number of weight parameters corresponding to each value, at this time, it is possible to determine which values ​​of weight parameters to trim from the lower weight parameter part according to the first model trimming ratio, and then determine the first weight parameter threshold. In some preferred embodiments, the preset first model trimming ratio can be, for example, less than 60%.

[0090] Step 113, setting the weight parameters in the convolution layer of the neural network model that are lower than the first weight parameter threshold to zero to complete the sparsification process.

[0091] In some embodiments, the weight parameters below the first weight parameter threshold play a very small role in the convolution process, and therefore, they can be directly set to zero.

[0092] As an optional implementation, Figure 7 As shown, in step 120, the channels in the convolution layer of the first intermediate model are trimmed according to a preset first trimming threshold to form a second intermediate model, including:

[0093] Step 121, comparing the sparsity rate of each channel in the convolutional layer of the first intermediate model with the first trimming threshold;

[0094] Step 122: When the sparsity rate of the channel is greater than the first clipping threshold, the channel is removed so that the channel does not participate in the convolution calculation.

[0095] In some embodiments, for channels with very high sparsity rates, their contribution to the convolution process is very low, and therefore, they can be pruned. Comparing the sparsity rate of each channel with the first pruned threshold can determine which channels have a low contribution to the convolution process.

[0096] As an optional implementation, in step 121, before comparing the sparsity rate of each channel in the convolutional layer of the first intermediate model with the first clipping threshold, the step further includes:

[0097] The weight parameters of each channel in the convolution layer of the first intermediate model are subjected to sparse rate statistics, and a first trimming threshold is determined according to the sparse rate statistics result.

[0098] In some embodiments, the sparsity rate statistics of the weight parameters of each channel in the convolutional layer can characterize the distribution law of the sparsity rate. Based on this law, the first clipping threshold can be determined more accurately to avoid accuracy not meeting the requirements due to clipping.

[0099] As an optional implementation, Figure 8 As shown, in step 130, the convolution mode of each convolution layer of the second intermediate model is determined and labeled according to a preset first judgment threshold to form a sparse neural network model, including:

[0100] Step 131, comparing the sparsity rate of each convolutional layer of the second intermediate model with the first judgment threshold;

[0101] Step 132: when the sparsity rate of the convolutional layer is greater than the first judgment threshold, marking the convolutional layer as sparse calculation;

[0102] Step 133: when the sparsity rate of the convolutional layer is not greater than the first judgment threshold, mark the convolutional layer as dense calculation.

[0103] In some embodiments, the first judgment threshold is used to determine which convolutional layers use sparse computing and which convolutional layers use dense computing. For convolutional layers with higher sparsity, that is, convolutional layers higher than the first judgment threshold, sparse computing can be used, and for convolutional layers with lower sparsity, that is, convolutional layers lower than the first judgment threshold, dense computing can be used.

[0104] As an optional implementation, Fig. 9 As shown, in step 130, after judging the convolution mode of each convolution layer of the second intermediate model according to the preset first judgment threshold and marking it, it also includes:

[0105] Step 140, performing a performance comparison between sparse computing and dense computing on the convolutional layer marked as sparse computing;

[0106] Step 150: When the dense computing performance of the convolutional layer is higher than the sparse computing performance, the computing method of the convolutional layer is transformed into dense computing and marked.

[0107] In some embodiments, for a convolutional layer marked as sparse computing, the performance of using a sparse computing method is not necessarily higher than that of a dense computing method. Therefore, in this embodiment, the performance of sparse computing is compared with the performance of dense computing, and then a computing method with better performance is selected.

[0108] As an optional implementation, after the channels in the convolution layer of the first intermediate model are trimmed according to the preset first trimming threshold to form the second intermediate model, it also includes: numerically compressing and storing the second intermediate model in the form of index storage. In some embodiments, when storing, the storage precision is divided into two types, one is single precision or half precision, and the other is integer. Numerical compression storage means that the weight parameters that are set to zero can be discarded and not stored. When using the weight parameters, the index is used to obtain their position. For example, the data before compression can be as follows:

[0109] <![CDATA[0 1,1 ]]> <![CDATA[1 1,2 ]]> <![CDATA[0 1,3 ]]> <![CDATA[2 1,4 ]]> <![CDATA[3 1,5 ]]> <![CDATA[0 1,6 ]]> <![CDATA[4 1,7 ]]> <![CDATA[5 1,8 ]]> <![CDATA[0 1,9 ]]> <![CDATA[6 1,10 ]]> <![CDATA[7 1,11 ]]> <![CDATA[0 1,12 ]]> <![CDATA[0 2,1 ]]> <![CDATA[0 2,2 ]]> <![CDATA[8 2,3 ]]> <![CDATA[0 2,4 ]]> <![CDATA[0 2,5 ]]> <![CDATA[9 2,6 ]]> <![CDATA[0 2,7 ]]> <![CDATA[10 2,8 ]]> <![CDATA[0 2,9 ]]> <![CDATA[0 2,10 ]]> <![CDATA[0 2,11 ]]> <![CDATA[11 2,12 ]]> <![CDATA[12 3,1 ]]> <![CDATA[0 3,2 ]]> <![CDATA[13 3,3 ]]> <![CDATA[0 3,4 ]]> <![CDATA[14 3,5 ]]> <![CDATA[15 3,6 ]]> <![CDATA[0 3,7 ]]> <![CDATA[16 3,8 ]]> <![CDATA[17 3,9 ]]> <![CDATA[18 3,10 ]]> <![CDATA[19 3,11 ]]> <![CDATA[20 3,12 ]]>

[0110] The compressed data is as follows:

[0111]

[0112] Those skilled in the art can understand that all or part of the processes in the above method embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above method embodiments. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0113] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A method for calculating sparseness of a neural network model, characterized in that: The method comprises: The weights of each convolutional layer of the neural network model are sparsed to obtain a sparse neural network model; Convert the weight parameters of each convolutional layer of the sparse neural network model into a weight matrix, and convert the input data into an input matrix; According to the cache capacity, the input matrix is ​​divided into a plurality of operator matrices; Obtaining a column index of a weight parameter in each row of the weight matrix, and obtaining a corresponding row from the operator matrix according to the column index; Multiply each weight parameter in the same row by the corresponding row in the operator matrix to obtain a one-dimensional matrix to be processed, and accumulate the one-dimensional matrix to be processed corresponding to each weight parameter in the same row to obtain a one-dimensional intermediate matrix; The convolution result is determined based on the one-dimensional intermediate matrix obtained by each row of weight parameters and each operator matrix.

2. The method according to claim 1, characterized in that: The input matrix is ​​divided into a plurality of operator matrices according to the cache capacity, including: According to the cache capacity, the maximum number of columns of the operator matrix is ​​determined to be n columns; wherein n≥1, and the cache capacity required for n columns does not exceed the maximum capacity of the cache that can currently be used to store the matrix.

3. The method according to claim 2, characterized in that After determining the maximum number of columns of the operator matrix to be n columns according to the cache capacity, the method further includes: When the number of columns of the input matrix is ​​a non-integer multiple of n, the input matrix is ​​divided into an integer number of operator matrices, and the remaining columns less than n are divided according to The columns are recursively divided into multiple operator matrices in descending order; where k ≥ 1; When the number of columns of the input matrix is ​​an integer multiple of n, the input matrix is ​​divided into an integer number of operator matrices.

4. The method according to claim 1, characterized in that: The step of performing weight sparse processing on each convolutional layer of the neural network model to obtain a sparse neural network model includes: According to the first weight parameter threshold, the convolution layer of the neural network model is subjected to sparse processing to form a first intermediate model; According to a preset first clipping threshold, clipping the channels in the convolutional layer of the first intermediate model to form a second intermediate model; According to a preset first judgment threshold, the convolution mode of each convolution layer of the second intermediate model is determined and labeled to form a sparse neural network model.

5. The method according to claim 4, characterized in that The step of performing a sparse processing on the convolution layer of the neural network model according to the first weight parameter threshold to form a first intermediate model includes: Counting the weight parameter distribution histogram in the convolutional layer of the neural network model; Determining a first weight parameter threshold value according to a preset first model clipping ratio and the weight parameter distribution histogram; The weight parameters in the convolution layer of the neural network model that are lower than the first weight parameter threshold are set to zero to complete the sparsification process.

6. The method according to claim 4, characterized in that The method of trimming the channels in the convolution layer of the first intermediate model according to a preset first trimming threshold to form a second intermediate model includes: Comparing the sparsity rate of each channel in the convolutional layer of the first intermediate model with the first clipping threshold; When the sparsity rate of the channel is greater than the first clipping threshold, the channel is removed so that the channel does not participate in the convolution calculation.

7. The method according to claim 6, characterized in that Before comparing the sparsity rate of each channel in the convolutional layer of the first intermediate model with the first clipping threshold, the method further includes: The weight parameters of each channel in the convolution layer of the first intermediate model are subjected to sparse rate statistics, and a first trimming threshold is determined according to the sparse rate statistics result.

8. The method according to claim 4, characterized in that The method of determining the convolution mode of each convolution layer of the second intermediate model and marking the convolution mode according to the preset first judgment threshold to form a sparse neural network model includes: Comparing the sparsity rate of each convolutional layer of the second intermediate model with the first judgment threshold; When the sparsity rate of the convolution layer is greater than the first judgment threshold, marking the convolution layer as sparse calculation; When the sparsity rate of the convolution layer is not greater than the first judgment threshold, the convolution layer is marked as dense calculation.

9. The method according to claim 4, characterized in that After judging the convolution mode of each convolution layer of the second intermediate model according to the preset first judgment threshold and marking it, the method further includes: For the convolutional layer marked as sparse computing, the performance comparison between sparse computing and dense computing is performed; When the dense computing performance of the convolutional layer is higher than the sparse computing performance, the computing mode of the convolutional layer is transformed into dense computing and marked.

10. The method according to claim 4, characterized in that After the channels in the convolution layer of the first intermediate model are trimmed according to the preset first trimming threshold to form the second intermediate model, the method further includes: The second intermediate model is numerically compressed and stored in an indexed storage format.

11. A neural network model sparse calculation device, characterized in that: include: The sparse module is used to sparse the weights of each convolutional layer of the neural network model to obtain a sparse neural network model; A conversion module, used to convert the weight parameters of each convolutional layer of the sparse neural network model into a weight matrix, and convert the input data into an input matrix; A partitioning module, used for partitioning the input matrix into a plurality of operator matrices according to a cache capacity; An acquisition module, used to acquire a column index of a weight parameter in each row of the weight matrix, and acquire a corresponding row from the operator matrix according to the column index; An operation module, used for multiplying each weight parameter in the same row with the corresponding row in the operator matrix to obtain a one-dimensional matrix to be processed, and accumulating the one-dimensional matrix to be processed corresponding to each weight parameter in the same row to obtain a one-dimensional intermediate matrix; The result module is used to determine the convolution result based on the one-dimensional intermediate matrix obtained by each row of weight parameters and each operator matrix.