An acceleration method and system for neural network model compression
Through mask-aware convolution calculation method, effective filters and input channels in neural network models are identified and reorganized, and invalid calculations are used to skip invalid calculations, solving the problem of waste and optimization time-consuming resources in the existing technology, and achieving the acceleration of the model compression process and the improvement of resource utilization.
Patent Information
- Application Number
- CN202210587507.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-25
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-05-25
AI Technical Summary
The existing neural network model compression technology has problems such as wasting resources, low resource utilization, and long optimization in the compression process, and it cannot effectively accelerate the model compression process.
The mask-aware convolution calculation method is adopted to identify and reorganize effective filters and input channels by perceiving mask information, and use GEMM convolution to skip invalid calculations to reduce redundant calculations.
The neural network model compression process is accelerated, the calculation amount and data movement are reduced, and resource utilization and optimization efficiency are improved.
Smart Images

Figure CN115018049B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of neural network model compression, and in particular to a neural network model compression acceleration method and system combined with mask information. Background Art
[0002] In recent years, neural network models have achieved significant accuracy improvements in various intelligent tasks such as image classification and object detection. However, in order to obtain better prediction accuracy, the neural network structure has become deeper and wider. The reasoning of these complex neural network models has high requirements for computing, storage, and energy consumption, and is difficult to use in some scenarios. With the increasing popularity of the Internet of Things, more and more intelligent tasks are being migrated to mobile and edge devices. In these resource-limited environments, such as embedded systems, it is challenging to deploy neural network models. Compression and acceleration of neural network models is a commonly used technical means.
[0003] Neural network filter pruning is one of the most representative model compression techniques. It compresses the model into a smaller model by deleting the weight parameters of unimportant filters in the convolution layer, thereby reducing the size of the model and the complexity of calculation at the same time. Generally, pruning and retraining are two important stages of filter pruning. In the pruning stage, the weight parameters of the filters selected for pruning will be completely set to zero. In the retraining stage, the pruned model is updated to restore the accuracy of the model as much as possible. Through the pruning and retraining stages, the neural network model can be compressed and accelerated without obvious loss of accuracy. The focus of the existing technology is usually on the final model compression and acceleration effect, such as designing a sophisticated pruning algorithm, but the efficiency of the model compression process is not considered. This leads to the problems of resource waste, low resource utilization, and long optimization time in the compression process of the existing technology.
[0004] Mainstream neural network model compression technology uses a mask-based approach to flexibly control the compression algorithm. During the compression process, the weight parameters of the filter to be pruned or not pruned are updated using mask 0 / 1, respectively, and the actual number of weight parameters remains unchanged. In the mask-based compression process, the compressed weight parameters still have learning capabilities, and can be restored at any time once it is found to be important, which can better approach the limit of model compression. Most existing research work focuses on the acceleration of the final compressed model, while ignoring the acceleration of the model compression process, which is also very time-consuming.
[0005] Figure 1 The principle of mask-based model compression is shown. The boxes and circles in the figure represent the number of effective filters and the number of effective channels, respectively. The actual shape is 32×32×R l ×Sl The solid arrows represent masks. Four filters are pruned for the l-1th convolutional layer, and the dotted arrows indicate the masks. 16 filters are pruned for the lth convolutional layer. The weight parameters of the l-1th convolutional layer are Masked After updating, it contains 28 effective filters, and the lth convolution layer is masked After updating, it contains 16 effective filters. The l-1th convolution is After the update, the input tensor of the lth convolutional layer is and weight parameters The effective number of channels is 28. Similarly, due to The impact of The number of effective input channels is 16. During the back propagation process, the output tensor Gradient Contains 16 valid input channels. When propagating along the computation graph, Gradient Contains 16 effective filters and 28 effective channels, Gradient Contains 28 valid input channels. Then, as it propagates downward, Gradient Contains 28 valid filters.
[0006] During the compression process, the weight parameters of these invalid filters are simply masked and updated to zero instead of being removed. When using the existing convolution calculation method, the calculation of the filter weight parameters that are set to zero and the input channels that are all zero cannot be skipped. In other words, the amount of calculation of the compressed convolution layer is the same as the original uncompressed convolution layer. Therefore, the existing mask-based model compression optimization process cannot be accelerated. Summary of the invention
[0007] In view of the problems existing in the prior art, the present invention constructs a new convolution calculation method, which reduces redundant calculations in the compressed neural network model by perceiving mask information, thereby accelerating the compression process of the neural network model. One difficulty in accelerating the compression process of the neural network model lies in the identification of effective filter weight parameters or effective input channels. In each training cycle of the model compression process, the retraining stage will update the weight parameters of the pruned filters, and the pruning stage will recalculate the importance of the weight parameters of all filters in the convolution layer according to the compression algorithm. The index of unimportant filters is not a fixed sequence. In order to solve this difficulty, the present invention proposes a mask-aware recognition method. This recognition method uses the mask information of the compression algorithm to identify and reorganize the effective channels and filters of the input tensors and weight parameters.
[0008] Another difficulty in accelerating the compression process of the neural network model is that in the feedforward propagation and backward propagation of the convolution layer, only the weight parameters of the valid channels / filters are calculated, while the redundant calculation of the weight parameters of the invalid channels / filters is skipped. To this end, the present invention adopts matrix multiplication (General Matrix Multiply, GEMM) to implement convolution calculations. By skipping the calculation of invalid rows or columns in the matrix, the calculation related to the weight parameters of the invalid channels / filters is avoided. In addition, using GEMM convolution, the existing high-performance library can be used to accelerate the calculation.
[0009] Specifically, the present invention proposes an acceleration method for neural network model compression, which includes:
[0010] Step 1: Get the training dataset Training cycle epoch max , pruning rate Weight parameters of the neural network model L is the convolutional layer in the neural network model The number of
[0011] Step 2: traverse the convolutional layer of the neural network model, sort the importance of the weight parameters of the filters in the convolutional layer, and according to the sorting results, In front of Lieutenant General The element value corresponding to the weight parameter of the least important filter is set to 0, where K l for The number of filters in
[0012] Step 3: Traverse all convolutional layers and execute and Element-wise multiplication of
[0013] Step 4: Based on and training dataset Calculate the output of the neural network model through mask-aware convolution calculation;
[0014] Step 5: Based on the output and training data set The label of the neural network model is obtained.
[0015] Step 6: Based on the mask information, use the mask-aware convolution calculation method to calculate the gradient of the loss function with respect to the input tensor and weight parameters;
[0016] Step 7: Use the gradient to update the weight parameters
[0017] Step 8: Execute steps 2-7 for epoch max After cycles, remove the current neural network model The weight parameters marked as zero result in a compressed model.
[0018] The acceleration method for neural network model compression, wherein the mask-aware convolution calculation method in step 4 includes:
[0019] Step 41: According to the mask information of the lth layer in the neural network model and the mask information of layer l-1 Organize the weight parameters in order The weight parameter of the effective filter in is obtained The corresponding (K′ l )×(C′ l ·R l ·S l )
[0020] Step 42: Based on the image to matrix algorithm, through the mask Identify the input tensor As the convolution kernel moves on the effective channel of the input tensor, the elements on the effective channel are vectorized into a column of the matrix to obtain the input tensor The dense matrix The shape is (C′ l ·R l ·S l )×(H l+1 ·W l+1 );
[0021] Step 43: Perform matrix multiplication Get a (K ′ l )×(H l+1 ·W l+1 )’s output matrix
[0022] Step 44: Using the mask right Perform an inverse expansion operation and output the matrix Restore the original convolution output tensor
[0023] The acceleration method for neural network model compression, wherein the mask-aware convolution calculation method in step 6 includes:
[0024] Step 61: Using the mask and right Rearrange the valid weight parameters in and skip the invalid weight parameters to get a (K′ l )×(C′ l ·R l ·S l )
[0025] Step 62: Using the mask The gradient of the output tensor of the lth convolutional layer in the neural network model Convert, rearrange its valid channels in order, skip invalid channels, and obtain (K′ l )×(H l+1 ·W l+1 )
[0026] Step 63: Perform matrix multiplication Get a (C ′ l ·R l ·S l )×(H l+1 ·W l+1 )’s output matrix
[0027] Step 64: Output matrix Based on the matrix to image algorithm, according to the mask Only valid channels are converted. Map to C l ×H l ×W l on the data layout.
[0028] The acceleration method for neural network model compression, wherein the neural network model is an image classification model, the training data set Including multiple pictures, the label is the category corresponding to the picture, and step 8 includes inputting the picture to be classified into the compression model to obtain the category of the picture to be classified.
[0029] The present invention also proposes an acceleration system for neural network model compression, which includes:
[0030] Initial module, used to obtain training data set Training cycle epoch max , pruning rate Weight parameters of the neural network model L is the convolutional layer in the neural network model The number of
[0031] The weight sorting module is used to traverse the convolutional layer of the neural network model, sort the importance of the weight parameters of the filters of the convolutional layer, and according to the sorting results, In front of Lieutenant General The element value corresponding to the weight parameter of the least important filter is set to 0, where K l for The number of filters in
[0032] Matrix multiplication module, used to traverse all convolutional layers and perform and Element-wise multiplication of
[0033] Feedforward propagation module, used for and training dataset Calculate the output of the neural network model through mask-aware convolution calculation;
[0034] The loss calculation module is used to calculate the loss based on the output and the training data set. The label of the neural network model is obtained.
[0035] The back-propagation module is used to calculate the gradient of the loss function with respect to the input tensor and weight parameters based on the mask information using a mask-aware convolution calculation method;
[0036] Weight update module, used to update weight parameters with this gradient
[0037] The output module is used to call the weight sorting module again and execute epoch max After cycles, remove the current neural network model The weight parameters marked as zero result in a compressed model.
[0038] The acceleration system for neural network model compression, wherein the mask-aware convolution calculation method in the feedforward propagation module includes:
[0039] According to the mask information of the lth layer in the neural network model and the mask information of layer l-1 Organize the weight parameters in order The weight parameter of the effective filter in is obtained The corresponding (K′ l )×(C′ l ·R l ·S l )
[0040] Based on the image to matrix algorithm, through the mask Identify the input tensor As the convolution kernel moves on the effective channel of the input tensor, the elements on the effective channel are vectorized into a column of the matrix to obtain the input tensor The dense matrix The shape is (C′ l ·R l ·S l )×(H l+1 ·W l+1 );
[0041] Performing matrix multiplication Get a (K ′ l )×(H l+1 ·W l+1 )’s output matrix
[0042] Utilizing Masks right Perform an inverse expansion operation and output the matrix Restore the original convolution output tensor
[0043] The acceleration system for neural network model compression, wherein the mask-aware convolution calculation method in the back-propagation module includes:
[0044] Utilizing Masks and right Rearrange the valid weight parameters in and skip the invalid weight parameters to get a (K′ l )×(C′ l ·R l ·S l )
[0045] Utilizing Masks The gradient of the output tensor of the lth convolutional layer in the neural network model Convert, rearrange its valid channels in order, skip invalid channels, and obtain (K′ l )×(H l+1 ·W l+1 )
[0046] Performing matrix multiplication Get a (C ′ l ·R l ·S l )×(H l+1 ·W l+1 )’s output matrix
[0047] For the output matrix Based on the matrix to image algorithm, according to the mask Only valid channels are converted. Map to C l ×H l ×W l on the data layout.
[0048] The acceleration system for neural network model compression, wherein the neural network model is an image classification model, the training data set It includes multiple pictures, the label is the category corresponding to the picture, and the output module is used to input the picture to be classified into the compression model to obtain the category of the picture to be classified.
[0049] The present invention also proposes a storage medium for storing a program for executing any one of the acceleration methods for neural network model compression.
[0050] The present invention also proposes a client for use in any of the acceleration systems for neural network model compression.
[0051] The acceleration strategy designed by the present invention for the compression process of the neural network model can utilize the mask information of the compression algorithm to perform calculations and data conversions only on the weight parameters of the effective filters and the effective input channels through the GEMM convolution method, thereby avoiding calculations and data movement related to the weight parameters of the invalid filters and the invalid input channels. This mask-aware convolution calculation method can reduce the amount of feedforward calculation and reverse calculation in the model compression process, thereby accelerating the entire model compression process and reducing the time cost and labor cost of development. In addition, the acceleration method and system of the present invention can adapt to different hardware platforms and have versatility and extensibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic diagram of the principle of mask-based model compression in the prior art;
[0053] Figure 2 A schematic diagram of a feedforward calculation process using a mask-aware convolution calculation method according to the present invention;
[0054] Figure 3 The figure is a schematic diagram of the gradient calculation process using the mask-aware convolution calculation method of the present invention. DETAILED DESCRIPTION
[0055] In order to make the above features and effects of the present invention more clearly understood, embodiments are given below and described in detail with reference to the accompanying drawings.
[0056] The present invention aims at the problems of resource waste, low resource utilization and long optimization time in the current neural network compression process, and designs a mask-aware convolution calculation method. It rearranges and organizes the weight parameters and channels of effective filters by efficiently utilizing mask information, and uses GEMM convolution to avoid the calculation of weight parameters and channels of invalid filters. The present invention proposes a mask-aware convolution calculation method, including:
[0057] A: Using mask information to speed up the calculation process of feedforward propagation of convolutional layer
[0058] In the feedforward propagation of mask-based model compression, the traditional GEMM convolution transforms all channels of the input tensor and then performs multiplication on two sparse matrices including invalid columns / rows. In contrast, the present invention transforms the weight parameters of the valid filters in the convolution layer (1 is valid and 0 is invalid) and the valid channels of the input tensor, thereby performing multiplication on two compact dense matrices. The mask-aware convolution calculation method proposed in the present invention avoids the redundant calculations generated by the weight parameters and channels of the invalid filters, effectively reduces the calculation cost, and thus reduces the feedforward propagation time of the model compression process.
[0059] B: Using mask information to speed up the calculation process of convolutional layer back propagation
[0060] In the back-propagation of mask-based model compression, the existing techniques are in the process of computing the gradient of the input tensor. and the gradient of the weight parameters , the sparse weight parameters and input tensors of the convolutional layer are used directly. The gradient of the output tensor of the lth convolutional layer is The sparsity of is uncertain and is determined by the gradient function of the l+1th layer. The non-zero gradient can be calculated for the zeroed weight parameters, so the compressed weight parameters can be updated to non-zero, thereby maintaining the learning ability of the model. Therefore, the calculation of the gradient of the weight parameters must maintain a sparse mode and cannot use mask information to obtain acceleration.
[0061] Gradient with respect to the input tensor In the calculation, the traditional GEMM convolution performs matrix multiplication on a sparse matrix and a matrix with uncertain sparsity, and then transforms the result of the matrix multiplication to finally obtain the gradient of the input tensor. Different from this, when calculating the gradient of the input tensor, the present invention uses mask information to convert the sparse weight parameter matrix into a dense matrix, and only performs data movement conversion on the valid channels of the input tensor, thereby reducing redundant calculations and data movement in the gradient calculation, thereby accelerating the back propagation process in the model compression process.
[0062] C: An acceleration strategy for the neural network model compression process
[0063] Combining the two technical points A and B, the present invention proposes an acceleration strategy for the compression process of neural network models. This strategy accelerates the time-consuming model compression process in a mask-aware manner. In the model output calculation of feedforward propagation, the weight parameters of the effective filters and the effective channels of the input tensor are reorganized according to the mask information of the compression algorithm, and the redundant calculation of invalid weight parameters and channels and data movement are avoided through GEMM convolution. In back propagation, the gradient calculation of the input tensor is accelerated. According to the calculation mode of the gradient of the input tensor, the weight parameters of the effective filters of the convolution layer are rearranged, so that invalid rows and columns are skipped in the GEMM calculation and subsequent data conversion operations.
[0064] 1. Use mask-aware convolution to speed up feedforward computation
[0065] In the process of model compression, the use of mask-aware convolution calculation can accelerate the feedforward propagation of the model. Figure 2 As shown in the figure, the mask-aware convolution calculation method uses mask information to dynamically convert the weight parameters of the filter, expand the input tensor, and restore the output tensor. The feedforward calculation of mask-aware convolution mainly includes the following steps:
[0066] ·Conversion filter weight parameters: according to the mask information of the lth layer and the mask information of layer l-1 Sparse weight parameters The weight parameters of the effective filters in are reorganized in an orderly manner, that is, they are accessed in sequence The elements in the array are stored in order. After the transformation, we can get a (K′ l)×(C′ l ·R l ·S l ) Among them, K′ l Represents the number of effective filters, C′ l Indicates the number of effective channels in each filter, R l Indicates the height of the filter, S l Indicates the width of the filter.
[0067] Expand the input tensor: Convert the input tensor in a way similar to the Image to Column (im2col) algorithm Through the mask The identified valid channels are expanded, where the valid channels are the input channels corresponding to the filters that have not been pruned, that is, the channels whose element values are not zero. As the convolution kernel moves on the valid channels of the input tensor, the elements on the valid channels are vectorized into a column of the matrix (the matrix here refers to the matrix in the image to matrix algorithm), and finally a compact dense matrix is obtained. Its shape is (C′ l ·R l ·S l )×(H l+1 ·W l+1 ).
[0068] GEMM: performs matrix multiplication Get a (K ′ l )×(H l+1 ·W l+1 )’s output matrix Compared with the traditional GEMM convolution, the computational complexity of the matrix multiplication in the present invention is smaller because the mask information of the compression algorithm is utilized to generate a smaller matrix as input for the GEMM convolution.
[0069] Restore output tensor: To meet the needs of neural network compression processes such as pruning, the output of matrix multiplication is Restore the original convolution output tensor To maintain the learning ability of the compressed model. To achieve this requirement, the mask right Perform the inverse operation of unfolding.
[0070] 2. Use mask-aware convolution to speed up reverse calculation
[0071] In the back propagation of model compression, the gradient of the input tensor of the convolutional layer is calculated using a mask-aware approach. The calculation of Figure 3As shown, the present invention utilizes mask information to dynamically convert the weight parameters of the filter and the gradient of the output tensor. It mainly includes the following steps:
[0072] ·Convert filter weight parameters: Similar to the conversion method of filter weight parameters in feedforward calculation, use mask and right The valid weight parameters in are rearranged, and invalid weight parameters are skipped. Finally, we get a (K′ l )×(C′ l ·R l ·S l )
[0073] Conversion Utilizing Masks The gradient of the output tensor of the lth convolutional layer of the neural network model After the transformation, the valid channels are rearranged in order and the invalid channels are skipped. After rearrangement, a (K′ l )×(H l+1 ·W l+1 )
[0074] GEMM: performs matrix multiplication Get a (C ′ l ·R l ·S l )×(H l+1 ·W i+1 )’s output matrix Compared with the original GEMM convolution, the matrix multiplication in the present invention reduces the calculations associated with invalid filter weight parameters and invalid channels.
[0075] Conversion Output for GEMM In a similar way to the matrix to image algorithm (Column toImage, col2im), according to the mask Only the valid channels in the GEMM output matrix are converted. Map to C l ×H l ×W l on the data layout.
[0076] 3. An acceleration strategy for the neural network model compression process
[0077] Aiming at the compression process of neural network, an acceleration system is designed, which can effectively improve the efficiency of model compression. It mainly consists of two parts:
[0078] Acceleration of feedforward propagation: For the calculation mode of feedforward propagation, the input tensors and weight parameters are reorganized using mask information, and then the high-performance GEMM library is used to implement convolution calculations, and finally the matrix results are restored to the original output tensor data layout.
[0079] · Acceleration of back propagation: By analyzing the calculation mode of gradients in back propagation, the original calculation method of the gradient calculation of weight parameters is retained, and the gradient calculation of input tensors is optimized. The gradients of weight parameters and output tensors are rearranged using mask information, and then GEMM calculations are performed, and finally the output of GEMM is mapped to the data layout of the original gradient of the input tensor.
[0080] Step 1: Prepare input for the acceleration system for model compression
[0081] The input of the acceleration system includes: training data set Training cycle epoch max , pruning rate The weight parameters of the original model Here, L is the number of convolutional layers in the model.
[0082] Step 2: Generate mask information
[0083] Traverse all convolutional layers and sort the importance of the weight parameters of the filters of each convolutional layer according to the compression algorithm. In the middle, the front The values of the elements corresponding to the weight parameters of the unimportant filters are set to 0, where K l express The number of filters in .
[0084] Step 3: Update weight parameters using mask information
[0085] Traverse all convolutional layers and execute and Element-wise multiplication of Thus The weight parameters of unimportant filters are updated to 0.
[0086] Step 4: Perform feed-forward computation using mask-aware convolution
[0087] based on and training dataset The output of the model is calculated using mask-aware convolution.
[0088] Step 5: Calculate the loss function
[0089] Calculate the loss function of the neural network model based on the labels of the training dataset and the output of the model.
[0090] Step 6: Calculate gradients using mask-aware convolution
[0091] Based on the mask information, the gradient of the loss function with respect to the input tensor and weight parameters is calculated using the mask-aware convolution calculation method.
[0092] Step 7: Update model parameters
[0093] Update the weight parameters according to the calculation results of the gradient in back propagation
[0094] Step 8: Return the compressed model
[0095] Perform steps 2-7 for epoch max cycles, then remove The weight parameters marked as zero in are used to obtain a final compressed model.
[0096] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in conjunction with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. In order to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiment.
[0097] The present invention also proposes an acceleration system for neural network model compression, which includes:
[0098] Initial module, used to obtain training data set Training cycle epoch max , pruning rate Weight parameters of the neural network model L is the convolutional layer in the neural network model The number of
[0099] The weight sorting module is used to traverse the convolutional layer of the neural network model, sort the importance of the weight parameters of the filters of the convolutional layer, and according to the sorting results, In front of Lieutenant General The element value corresponding to the weight parameter of the least important filter is set to 0, where K l for The number of filters in
[0100] Element-by-element multiplication module, used to traverse all convolutional layers and perform and Element-wise multiplication of
[0101] Feedforward propagation module, used for and training dataset Calculate the output of the neural network model through mask-aware convolution calculation;
[0102] The loss calculation module is used to calculate the loss based on the output and the training data set. The label of the neural network model is obtained.
[0103] The back-propagation module is used to calculate the gradient of the loss function with respect to the input tensor and weight parameters based on the mask information using a mask-aware convolution calculation method;
[0104] Weight update module, used to update weight parameters with this gradient
[0105] The output module is used to call the weight sorting module again to execute epoch max After cycles, remove the current neural network model The weight parameters marked as zero result in a compressed model.
[0106] The acceleration system for neural network model compression, wherein the mask-aware convolution calculation method in the feedforward propagation module includes:
[0107] According to the mask information of the lth layer in the neural network model and the mask information of layer l-1 Organize the weight parameters in order The weight parameter of the effective filter in is obtained The corresponding (K′ l )×(C′ l ·R l ·S l )
[0108] Based on the image to matrix algorithm, through the mask Identify the input tensor As the convolution kernel moves on the effective channel of the input tensor, the elements on the effective channel are vectorized into a column of the matrix to obtain the input tensor The dense matrix The shape is (C′ l ·R l ·S l )×(H l+1 ·W l+1 );
[0109] Performing matrix multiplication Get a (K ′ l )×(H l+1 ·W l+1 )’s output matrix
[0110] Utilizing Masks right Perform an inverse expansion operation and output the matrix Restore the original convolution output tensor
[0111] The acceleration system for neural network model compression, wherein the mask-aware convolution calculation method in the back-propagation module includes:
[0112] Utilizing Masks and right Rearrange the valid weight parameters in and skip the invalid weight parameters to get a (K′ l )×(C′ l ·R l ·S l )
[0113] Utilizing Masks The gradient of the output tensor of the lth convolutional layer in the neural network model Convert, rearrange its valid channels in order, skip invalid channels, and obtain (K′ l )×(H l+1 ·W l+1 )
[0114] Performing matrix multiplication Get a (C ′ l ·R l ·S l )×(H l+1 ·W l+1 )’s output matrix
[0115] For the output matrix Based on the matrix to image algorithm, according to the mask Only valid channels are converted. Map to C l ×H l ×W l on the data layout.
[0116] The acceleration system for neural network model compression, wherein the neural network model is an image classification model, the training data set It includes multiple pictures, the label is the category corresponding to the picture, and the output module is used to input the picture to be classified into the compression model to obtain the category of the picture to be classified.
[0117] The present invention also proposes a storage medium for storing a program for executing any one of the acceleration methods for neural network model compression.
[0118] The present invention also proposes a client for use in any of the acceleration systems for neural network model compression.
Claims
1. An acceleration method for neural network model compression, characterized in that: include: Step 1: Get the training dataset Training cycle epoch max , pruning rate Weight parameters of the neural network model L is the convolutional layer in the neural network model The number of Step 2: traverse the convolutional layer of the neural network model, sort the importance of the weight parameters of the filters in the convolutional layer, and according to the sorting results, In front of Lieutenant General The element value corresponding to the weight parameter of the least important filter is set to 0, where K l for The number of filters in Step 3: Traverse all convolutional layers and execute and Element-wise multiplication of Step 4: Based on and training dataset Calculate the output of the neural network model through mask-aware convolution calculation; Step 5: Based on the output and training data set The label of the neural network model is obtained. Step 6: Based on the mask information, use the mask-aware convolution calculation method to calculate the gradient of the loss function with respect to the input tensor and weight parameters; Step 7: Use the gradient to update the weight parameters Step 8: Execute steps 2-7 for epoch max After cycles, remove the current neural network model The weight parameters marked as zero result in a compressed model; The mask-aware convolution calculation method in step 4 includes: Step 41: According to the mask information of the lth layer in the neural network model and the mask information of layer l–1 Organize the weight parameters in order The weight parameter of the effective filter in is obtained The corresponding (K′ l )×(C′ l ·R l ·S l ) K′ l Represents the number of effective filters, C′ l Indicates the number of effective channels in each filter, R l Indicates the height of the filter, S l Indicates the width of the filter; Step 42: Based on the image to matrix algorithm, through the mask Identify the input tensor As the convolution kernel moves on the effective channel of the input tensor, the elements on the effective channel are vectorized into a column of the matrix to obtain the input tensor The dense matrix The shape is (C′ l ·R l ·S l )×(H l+1 ·W l+1 ); Step 43: Perform matrix multiplication Get a (K ′ l )×(H l+1 ·W l+1 )’s output matrix Step 44: Using the mask right Perform an inverse expansion operation and output the matrix Restore the original convolution output tensor The mask-aware convolution calculation method in step 6 includes: Step 61: Using the mask and right Rearrange the valid weight parameters in and skip the invalid weight parameters to get a (K′ l )×(C′ l ·R l ·S l ) Step 62: Using the mask The gradient of the output tensor of the lth convolutional layer in the neural network model Convert, rearrange its valid channels in order, skip invalid channels, and obtain (K′ l )×(H l+1 ·W l+1 ) Step 63: Perform matrix multiplication Get a (C ′ l ·R l ·S l )×(H l+1 ·W l+1 )’s output matrix Step 64: Output matrix Based on the matrix to image algorithm, according to the mask Only valid channels are converted. Map to C l ×H l ×W l On the data layout; The neural network model is a picture classification model. The training data set Including multiple pictures, the label is the category corresponding to the picture, and step 8 includes inputting the picture to be classified into the compression model to obtain the category of the picture to be classified.
2. An acceleration system for neural network model compression, characterized in that: include: Initial module, used to obtain training data set Training cycle epoch max , pruning rate Weight parameters of the neural network model L is the convolutional layer in the neural network model The number of The weight sorting module is used to traverse the convolutional layer of the neural network model, sort the importance of the weight parameters of the filters of the convolutional layer, and according to the sorting results, In front of Lieutenant General The element value corresponding to the weight parameter of the least important filter is set to 0, where K l for The number of filters in Element-by-element multiplication module, used to traverse all convolutional layers and perform and Element-wise multiplication of Feedforward propagation module, used based on and training dataset Calculate the output of the neural network model through mask-aware convolution calculation; The loss calculation module is used to calculate the loss based on the output and the training data set. The label of the neural network model is obtained. The back-propagation module is used to calculate the gradient of the loss function with respect to the input tensor and weight parameters based on the mask information using a mask-aware convolution calculation method; Weight update module, used to update weight parameters with this gradient The output module is used to call the weight sorting module again to execute epoch max After cycles, remove the current neural network model The weight parameters marked as zero result in a compressed model; The mask-aware convolution calculation method in the feedforward propagation module includes: According to the mask information of the lth layer in the neural network model and the mask information of layer l–1 Organize the weight parameters in order The weight parameter of the effective filter in is obtained The corresponding (K′ l )×(C′ l ·R l ·S l ) K′ l Represents the number of effective filters, C′ l Indicates the number of effective channels in each filter, R l Indicates the height of the filter, S l Indicates the width of the filter; Based on the image to matrix algorithm, through the mask Identify the input tensor As the convolution kernel moves on the effective channel of the input tensor, the elements on the effective channel are vectorized into a column of the matrix to obtain the input tensor The dense matrix The shape is (C′ l ·R l ·S l )×(H l+1 ·W l+1 ); Performing matrix multiplication Get a (K ′ l )×(H l+1 ·W l+1 )’s output matrix Utilizing Masks right Perform an inverse expansion operation and output the matrix Restore the original convolution output tensor The mask-aware convolution calculation method in this back-propagation module includes: Utilizing Masks and right Rearrange the valid weight parameters in and skip the invalid weight parameters to get a (K′ l )×(C′ l ·R l ·S l ) Utilizing Masks The gradient of the output tensor of the lth convolutional layer in the neural network model Convert, rearrange its valid channels in order, skip invalid channels, and obtain (K′ l )×(H l+1 ·W l+1 ) Performing matrix multiplication Get a (C ′ l ·R l ·S l )×(H l+1 ·W l+1 )’s output matrix For the output matrix Based on the matrix to image algorithm, according to the mask Only valid channels are converted. Map to C l ×H l ×W l On the data layout; The neural network model is a picture classification model. The training data set It includes multiple pictures, the label is the category corresponding to the picture, and the output module is used to input the picture to be classified into the compression model to obtain the category of the picture to be classified.
3. A storage medium for storing a program for executing the acceleration method for neural network model compression as described in claim 1.
4. A client for use in the acceleration system for neural network model compression as described in claim 2.
Citation Information
Patent Citations
Mask-based depth neural network compression method
CN107689224A
Image classification method and device, data processing method and device
CN110188795A