A method and system for structured pruning
By automatically selecting the pruning layer to be pruned based on the importance sorting of weight parameters and activation frequency, the problem of mistakenly deleting important layers in the existing technology is solved, and the efficient compression and accuracy guarantee of deep neural networks on mobile devices is achieved.
Patent Information
- Application Number
- CN202210805865.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-08
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-07-08
AI Technical Summary
The prior art is very likely to accidentally delete important network layers when pruning deep neural networks, and unstructured pruning cannot accelerate on existing hardware, resulting in difficulty in deploying mobile devices.
By sorting based on the importance of weight parameters, gradients and activation frequency, the pruning layer is automatically selected and pre-pruning and fine-tuning are performed multiple times during the training process to ensure that the model accuracy is not lost.
It realizes that while compressing the model size, it reduces the probability of accidentally deleting important layers, and improves the operating efficiency and accuracy of the model on mobile devices.
Smart Images

Figure CN115222042B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural networks, specifically to the technical field of neural network model compression, and more specifically, to a method and system for structured pruning. Background Art
[0002] With the continuous development of artificial intelligence, the application scope of artificial neural networks (referred to as neural networks for short) is also constantly expanding. Neural networks can be applied to mobile devices to process images, texts, audios, etc. and perform corresponding prediction tasks. In the field of image recognition (such as image classification, object detection, and object tracking), high accuracy can be achieved by using deep neural networks (also referred to as deep learning models in some literature). However, due to the large number of various network layers (also referred to as processing layers in some literature, referring to complete layer structures such as convolutional layers and fully connected layers) in deep neural networks (such as hundreds of thousands or even millions), it has high requirements for the computing power and storage space of mobile devices. At present, there is often a large gap between the computing power and storage space of mobile devices and those of desktop computers and servers. Under the limited resources, there are huge obstacles to directly applying deep learning models, resulting in difficulty in directly deploying deep neural networks to mobile devices. Therefore, it is necessary to prune deep neural networks to reduce the redundant connections and parameters of the models, so as to compress the size of deep neural networks and accelerate the model inference speed.
[0003] Pruning methods are usually divided into structured pruning and unstructured pruning, among which:
[0004] Structured pruning is a method of pruning with network layers as the basic unit; when a network layer is pruned, the previous feature map (Feature Map, usually also referred to as a feature or feature vector) and the next feature map will change accordingly, but the structure of the model is not damaged and can still be accelerated by a GPU or other hardware. Therefore, this type of method is called structured pruning.
[0005] Unstructured pruning includes pruning of single weight parameters and / or single convolutional kernels. Since the pruned model is usually very sparse and destroys the original model structure, this type of method is called unstructured pruning. Although unstructured pruning can greatly reduce the number of model parameters and theoretical computational volume, the existing hardware architecture's computing method cannot accelerate it, so there is no improvement in the actual running speed, and specific hardware needs to be designed to possibly accelerate, and it is more difficult to transplant to mobile devices after pruning.
[0006] In the prior art, in most cases, the L1 norm is added to the scaling factor of the batch normalization (BN) layer during the training process for regularization sparsification. The importance of each channel is measured according to the magnitude of the BN scaling factor after sparsification, and then a threshold is manually selected to prune the channels where the BN scaling factor is less than the threshold to reduce the model size. Finally, the compressed model is fine-tuned to recover the accuracy of the pruned loss. However, this compression method is relatively cumbersome. This process requires sparse regular training of the model and manual selection of a threshold to determine which network layers with smaller modulus values to prune. However, it is difficult to determine the pruning ratio by manually selecting the threshold, and the manual selection of the threshold may accidentally delete some important network layers. For this reason, some technical solutions for automatically pruning network layers according to the magnitude of the weight parameters in the network layers have also emerged. However, the indicators referred to in these technical solutions are too one-sided and do not consider the impact of the sample input of the dataset during training on the output of each network layer. The probability of accidentally deleting some important network layers is still relatively high. Therefore, improvements need to be made to the prior art. Summary of the Invention
[0007] Therefore, the object of the present invention is to overcome the above-mentioned defects of the prior art and provide a method and system for structured pruning.
[0008] The object of the present invention is achieved by the following technical solutions:
[0009] According to a first aspect of the present invention, there is provided a method for structured pruning for pruning a deep neural network. The deep neural network includes a plurality of network layers. The method includes: S1. Setting a specified network layer in the deep neural network as a layer to be pruned, each layer to be pruned corresponding to a network layer, to obtain a deep neural network to be processed; S2. Training the deep neural network to be processed multiple times using an image dataset and performing multiple pre-pruning processes during the training process. Wherein, each pre-pruning process includes: determining the importance of all layers to be pruned on the image dataset, setting a plurality of layers to be pruned with relatively low importance ranking and that can meet the pruning amount of each pre-pruning after pre-pruning as being pre-pruned, and setting all elements contained in the output of the pre-pruned layer to be pruned to 0 during subsequent training. Wherein, the importance of each layer to be pruned on the image dataset is determined according to a predetermined calculation rule based on the weight parameters, gradients, and activation frequencies of the layer to be pruned. The activation frequency is the number of times the output size of the layer to be pruned exceeds a predetermined activation threshold during training; S3. When the number of pre-pruning processes reaches a predetermined number of pre-pruning times, fine-tuning the pre-pruned deep neural network to be processed using the image dataset, and pruning the network layer corresponding to the pre-pruned layer to be pruned to obtain a compressed deep neural network.
[0010] In some embodiments of the present invention, the importance of each pruning layer on the image dataset is determined according to the following calculation rules: Determine the saliency index based on the latest weight parameters and gradients of the network layer corresponding to the pruning layer; Calculate the weighted sum of the saliency index and the activation frequency of the pruning layer to obtain the importance of the pruning layer on the image dataset.
[0011] In some embodiments of the present invention, the activation frequency is determined in the following manner: Determine the output size corresponding to the output of each pruning layer when each sample in the image dataset is input into the depth neural network to be processed; Use the number of times that the output size of the pruning layer exceeds a predetermined activation threshold in the latest completed preset number of trainings as the activation frequency; wherein, the output size is obtained by taking the square root of the sum of the squares of the elements contained in the output, and the activation threshold is the average value of all output sizes in the latest completed preset number of trainings.
[0012] In some embodiments of the present invention, the saliency index is determined in the following manner: Multiply each corresponding element in the latest weight parameters and gradients in the network layer corresponding to the pruning layer and then sum to obtain the saliency index of the pruning layer.
[0013] In some embodiments of the present invention, the elements contained in the output of the pruning layer to be pre-pruned are set to 0 in the subsequent training in the following manner: Set a mask for each pruning layer respectively. The output of the pruning layer is the element-wise product of each element in the output of its corresponding network layer and its mask. Among them, the mask indicates whether the pruning layer is pre-pruned. A mask of 0 indicates pre-pruning, and a mask of 1 indicates not pre-pruning. All masks are initially set to 1.
[0014] In some embodiments of the present invention, step S3 includes: Fine-tuning the depth neural network after the last pre-pruning process multiple times using the image dataset; Remove the network layer corresponding to the pruning layer to be pre-pruned from the depth neural network to obtain a compressed depth neural network.
[0015] According to the second aspect of the present invention, a structured pruning system is provided, including: A model modeling interface for loading a depth neural network and receiving a user instruction to set a corresponding network layer in the depth neural network as a pruning layer to obtain a depth neural network to be processed; A pruning module for performing the method of structured pruning described in the first aspect to obtain a compressed depth neural network.
[0016] According to the third aspect of the present invention, a method for deploying a depth neural network includes: Obtain a compressed depth neural network, where the compressed depth neural network is obtained by using the method described in the first aspect or the system described in the second aspect; Deploy the compressed depth neural network on a mobile device for performing corresponding prediction tasks.
[0017] According to a fourth aspect of the present invention, an electronic device includes: one or more processors; and a memory, where the memory is used to store executable instructions; the one or more processors are configured to implement the steps of the method described in the first aspect or the third aspect by executing the executable instructions.
[0018] Compared with the prior art, the advantages of the present invention are as follows:
[0019] The present invention determines the importance of the layer to be pruned based on the weight parameters, gradients, and activation frequencies of the layer to be pruned, where the activation frequency is the number of times the output size of the layer to be pruned exceeds a predetermined activation threshold during training; on the one hand, the network layers that need to be pre-pruned can be automatically selected through importance ranking, and on the other hand, by considering the influence of the sample input of the image data set on the output of each layer to be pruned during training, the importance of the layer to be pruned on the data set can be evaluated more accurately, reducing the probability of erroneously deleting some important layers to be pruned; so as to better ensure the accuracy of the model while reducing the size of the model. Description of the Drawings
[0020] The following further describes the embodiments of the present invention with reference to the drawings, where:
[0021] Figure 1 It is a schematic flowchart of a method for structured pruning according to an embodiment of the present invention;
[0022] Figure 2 It is a schematic diagram of a coding module of a vision transformer model;
[0023] Figure 3 It is a schematic diagram of the principle of filling zeros according to a mask after pruning the coding module of the vision transformer model according to an embodiment of the present invention;
[0024] Figure 4 It is a schematic diagram of the structure of a VGG-16 model;
[0025] Figure 5 It is a schematic diagram of the structure of a compressed VGG-16 model obtained by performing structured pruning on the VGG-16 model according to an embodiment of the present invention. Detailed Embodiments
[0026] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below through specific embodiments with reference to the drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0027] As mentioned in the background art section, the prior art automatically deletes network layers based on the magnitude of weight parameters in the network layer. However, the metrics referred to in these technical solutions are too one-sided and do not consider the impact of the sample input of the dataset during training on the outputs of each network layer. The probability of erroneously deleting some important network layers is still relatively high. Therefore, the present application determines the importance of the layer to be pruned based on the weight parameters, gradients, and activation frequencies of the layer to be pruned, where the activation frequency is the number of times the output size of the layer to be pruned during training exceeds a predetermined activation threshold. On the one hand, the network layers to be pre-pruned can be automatically selected through importance ranking. On the other hand, by considering the impact of the sample input of the image dataset during training on the outputs of each layer to be pruned, the importance of the layer to be pruned on the dataset can be evaluated more accurately, reducing the probability of erroneously deleting some important layers to be pruned; so as to better guarantee the accuracy of the model while achieving compression of the model size.
[0028] According to an embodiment of the present invention, a method for structured pruning is provided for pruning a deep neural network, the deep neural network including a plurality of network layers, the method including: steps S1, S2, and S3. To better understand the present invention, the following will separately describe each step in detail with reference to specific embodiments.
[0029] Step S1: Obtain a deep neural network including a plurality of network layers, where a plurality of layers to be pruned are set, and each layer to be pruned includes one network layer.
[0030] According to an embodiment of the present invention, a user instruction is received to set a corresponding network layer in the deep neural network as a layer to be pruned. The layer to be pruned is provided with a mask, and the output of the layer to be pruned is the product of the output of the network layer it contains and the mask. The mask indicates whether the layer to be pruned is pre-pruned through 0 or 1, where 0 indicates being pre-pruned and 1 indicates not being pre-pruned, and all masks are initially set to 1. The deep neural network can be various existing or future deep neural networks (some literatures also refer to deep learning models or deep neural network models, and in some subsequent parts of this article, "model" is used for short). For example, existing transformers (such as Vision Transformer, Swin Transformer, etc.), VGG models (such as VGG-16 model), and the present invention does not impose any restrictions on this. The network layers included in the deep neural network are convolutional layers, attention mechanism layers, fully connected layers, and perceptron layers or combinations thereof. The user can specify the network layers involved in the pruning process; it should be understood that based on the type of network layers actually included in the deep neural network, different users' experiences, or different pruning requirements, the user can customize the network layers involved in the pruning process, that is: the user sets the specified multiple network layers as layers to be pruned; for example, sets at least a part of the convolutional layers and at least a part of the fully connected network layers as layers to be pruned.
[0031] Step S2: Use the image dataset to train the deep neural network to be processed multiple times and perform multiple pre-pruning processes during the training process. Each pre-pruning process includes: determining the importance of all layers to be pruned on the image dataset, setting multiple layers to be pruned with lower importance rankings and that can meet the pruning amount of each pre-pruning after being pre-pruned as pre-pruned, and setting all elements in the output of the pre-pruned layer to be 0 during subsequent training. The importance of each layer to be pruned on the image dataset is determined according to a predetermined calculation rule based on the weight parameters, gradients, and activation frequencies of the layer to be pruned, and the activation frequency is the number of times the output size of the layer to be pruned exceeds a predetermined activation threshold during training.
[0032] According to an embodiment of the present invention, step S2 includes performing a pre-pruning process on the layers to be pruned that are not yet pre-pruned currently each time the deep neural network to be processed is trained with the image dataset to reach a preset number of times. The pre-pruning process includes: S21. Determine the saliency index according to the latest weight parameters and gradients of the network layer corresponding to the layer to be pruned; S22. Calculate the weighted sum of the saliency index and activation frequency of the layer to be pruned to obtain the importance of the layer to be pruned on the image dataset; S23. Set multiple layers to be pruned with lower importance rankings on the image dataset and that can meet the pruning amount of each pre-pruning after being pre-pruned as pre-pruned, and set all elements in the output of the pre-pruned layer to be 0. The pruning amount of each pre-pruning is obtained by dividing the predetermined total pruning amount by the number of pre-pruning times. It should be understood that step S23 sets multiple layers to be pruned that were originally not pre-pruned and have lower importance rankings on the image dataset and can meet the pruning amount of each pre-pruning as pre-pruned. It should be understood that a lower importance ranking means a lower degree of importance; the importance ranking can be determined according to the numerical size of the importance.
[0033] To avoid continuously changing and organizing the structure of the model during multiple pruning processes, the present invention implements pre-pruning through a mask, setting the output of the layer to be pruned that needs to be pre-pruned to 0. According to an embodiment of the present invention, the output of the layer to be pruned is the product of the output of the network layer it contains and the set mask. This product is obtained by element-wise multiplication, that is, each element in the output of the network layer it contains is multiplied by the mask. Thus, if the mask of the layer to be pruned is 0, each element in the output of the layer to be pruned is 0, and the layer to be pruned is in a pre-pruned state; if the mask of the layer to be pruned is 1, the output of the layer to be pruned is equal to the output of the network layer it contains, and the layer to be pruned is in an un-pre-pruned state. The technical solution of this embodiment can at least achieve the following beneficial technical effects: By setting the mask, the present invention can avoid continuously changing and organizing the structure of the model during multiple pruning processes, reduce the operation difficulty of pruning the model, and improve the pruning efficiency.
[0034] For the output of the layer to be pruned, according to an embodiment of the present invention, the calculation method of the output of the layer to be pruned is as follows:
[0035] Y m = m·W·X
[0036] Wherein, m represents a mask, W represents the weight parameter of the processing layer corresponding to the layer to be pruned, X represents the input of the processing layer corresponding to the layer to be pruned, · represents element-wise multiplication calculation, and W·X represents the output of the processing layer corresponding to the pruning layer.
[0037] For the saliency index, any one of a variety of calculation methods can be used to determine it. According to an embodiment of the present invention, the saliency index is determined in the following manner: Multiply each corresponding element in the latest weight parameter and gradient in the network layer corresponding to the layer to be pruned and then sum them to obtain the saliency index of the layer to be pruned. Another example is that if you want to reduce the influence of the cancellation of negative and positive values when directly summing the products, according to an embodiment of the present invention, the saliency index is determined in the following manner: Take the absolute value of the product of each corresponding element in the latest weight parameter and gradient in the network layer corresponding to the layer to be pruned, and sum each absolute value to obtain the saliency index of the layer to be pruned. Still another example is that according to an embodiment of the present invention, the saliency index is determined in the following manner: Sum the product of the square of the element of the latest weight parameter in the network layer corresponding to the layer to be pruned and the square of each corresponding element in the gradient to obtain the saliency index of the layer to be pruned. In addition, the latest weight parameter and gradient can also be converted into matrices, and matrix multiplication can be used to calculate the saliency index; for example, according to an embodiment of the present invention, the saliency index is determined in the following manner: Convert the elements in the latest gradient of the network layer corresponding to the layer to be pruned into a row vector and convert the latest weight parameter into a column vector, and multiply the row vector by the column vector to obtain the saliency index of the layer to be pruned. It should be understood that the latest gradient corresponding to the network layer is the gradient used for the latest weight parameter update, and the latest weight parameter is the current weight parameter of the network layer, that is, the weight parameter obtained after the latest update. The technical solution of this embodiment can at least achieve the following beneficial technical effects: The saliency index of the present invention can represent the influence degree of the layer to be pruned on the model output distribution, so that the depth neural network to be processed can be better pruned based on the saliency index.
[0038] To facilitate the understanding that the saliency metric of the present invention can represent the influence degree of the layer to be pruned on the model output distribution, the applicant hereby gives the derivation process of the saliency metric. If the calculation difficulty is ignored, the saliency metric is calculated as: S[W|p(y|x)] = H[W|p(y|x)] - H[w|p(y|x)], where S is the saliency metric of the layer W to be pruned, p(y|x) is the class probability density distribution of the model output, x is the input sample, y is the sample true value, H[W|p(y|x)] is the information entropy of the layer W to be pruned, and H[w|p(y|x)] is the information entropy of the remaining network structure when the layer W to be pruned is removed. Activation frequency calculation: F[W] = ∑ x δ(A(x|W)>T), where x is the input sample, A(x|W) is the output size (or activation value) of the layer W to be pruned for the input sample x, T is the set threshold, δ is a 0-1 binary function, > is a conditional judgment, which is 1 if true and 0 otherwise. For more efficient pruning, it is necessary to derive and improve the calculation method of the saliency metric. The derivation process is as follows. For a perfectly trained deep neural network (accuracy 100%), the first term H[W|p(y|x)] in the formula S[W|p(y|x)] = H[W|p(y|x)] - H[w|p(y|x)] is 0. However, generally, it is very difficult for a deep neural network to reach an accuracy of 100%. Therefore, H[W|p(y|x)] cannot be omitted. In addition, it is very costly to test all the layers to be pruned one by one through the formula S[W|p(y|x)] = H[W|p(y|x)] - H[w|p(y|x)], especially for extremely deep convolutional neural networks. Therefore, a Taylor expansion approximation algorithm is proposed to obtain the saliency metric. By the Taylor polynomial, at w = 0, the Taylor expansion of H[W|p(y|x)] is obtained as:
[0039]
[0040] where R t (w = 0) is the first-order remainder, which can be ignored in the calculation process, is the gradient of the information entropy of the model output class probability density distribution with respect to the backpropagation of all the layers to be pruned. Therefore, according to the above formula combination, the saliency metric of the layer to be pruned is:
[0041]
[0042] where, Denote the latest gradient of the network layer corresponding to the layer to be pruned as \(g\), the latest weight parameter in the network layer corresponding to the layer to be pruned as \(w\), and \(\sum\) represents the sum of the products of each corresponding element in the latest weight parameter and gradient in the network layer corresponding to the layer to be pruned. The gradient and the weight parameter usually have multiple elements and correspond one by one (one element of the gradient controls the update of one element of the weight parameter). Therefore, the elements of the two can be multiplied and summed correspondingly. Thus, the saliency index of the layer to be pruned can measure the importance of the layer to be pruned to the probability density distribution of the model output class. Taking pruning a convolutional kernel as an example, if an input sample \(x\) is given, to determine the impact of deleting the layer to be pruned on the model output \(p(y|x)\), if the saliency index of a convolutional kernel is high, then this convolutional kernel has a high discriminability for the model output; if the saliency index is low, then this convolutional kernel has low or even no discriminability for the model output. Therefore, this convolutional kernel is redundant and can be safely deleted from the model.
[0043] For the activation frequency, any one of a variety of calculation methods can be used to determine it. According to an embodiment of the present invention, the activation frequency is determined as follows: Determine the output size corresponding to the output of each layer to be pruned when each sample in the image dataset is input into the deep neural network to be processed; Take the number of times that the output size of the layer to be pruned exceeds a predetermined activation threshold in the latest completed preset number of trainings as the activation frequency; wherein, the output size is obtained by taking the square root of the sum of the squares of the elements contained in the output, and the activation threshold is the average value of all output sizes in the latest completed preset number of trainings. Alternatively, change the output size to: the output size is the sum of the squares of the elements contained in the output. Or, change the activation threshold to: the activation threshold is the median, the first quartile or the third quartile, etc. of all output sizes in the latest completed preset number of trainings. The specific implementation manner can be determined by the implementer according to needs, and the present invention does not make any limitation thereto. For example, the activation frequency is calculated as: \(F[W]=\sum\) x \(\delta(A(x|W)>T)\), where \(x\) is the input sample, \(A(x|W)\) is the output size of the layer to be pruned \(W\) for the input sample \(x\), \(T\) is the set activation threshold, \(\delta\) is a 0-1 binary function, \(>\) is a conditional judgment, if true it is 1, otherwise it is 0. It can be seen from the above formula that when the feature extraction ability of a certain layer to be pruned is weak, its output size will be very small. Given an activation threshold \(T\), if it is less than this activation threshold, it is regarded that the feature extraction ability of this layer to be pruned is very weak, so it can be deleted; on the contrary, if it is greater than this activation threshold, it is regarded that the feature extraction ability of this layer to be pruned is very strong, so it can be retained. To avoid manual selection in the pruning process, the activation threshold is a set value, which can be determined by analyzing the average feature value of all layers to be pruned. For example: set it as the average value of all output sizes in the latest completed preset number of trainings. The 0-1 binary function \(\delta(x)\) is:
[0044]
[0045] In summary, by cumulatively activating the frequencies of all the layers to be pruned on the data set using the above method, the feature extraction ability of each layer to be pruned on the image data set can be obtained, and this is used to evaluate the importance of the layer to be pruned.
[0046] According to an embodiment of the present invention, the importance I of the layer to be pruned on the image data set is I = αS[W|p(y|x)] + βF[W], where α represents the weighting coefficient of the saliency index, and β represents the weighting coefficient of the activation frequency. α and β are two adjustable weighting coefficients. It can be seen from the formula that the importance of the layer to be pruned on the image data set is weighted by the saliency index and the activation frequency, and together they determine whether the layer to be pruned is pruned or retained. For different image data sets, the settings of α and β can be different, and the magnitudes of α and β can be determined by the implementer according to experience or the results of on-site debugging. The present invention does not impose any restrictions on this.
[0047] For the image data set, existing image data sets can be used. The prediction tasks can be image classification, object detection, tracking and recognition, etc.; for example, the Imagenet data set, the Cifar-10 data set, etc. The specific image data set can be set according to the needs of the implementer, and the present invention does not impose any restrictions on this.
[0048] For the loss function used in the training process, existing loss functions (such as the cross-entropy loss function, etc.) or loss functions formulated by the implementer according to needs can be used. The present invention does not impose any restrictions on this. For example, for the classification task, the cross-entropy is mainly used as the loss function: where (x i , y i ) is the sample, p(x i ) is the output of the model after Softmax processing, x i is the i-th input sample, y i is the true value (label) corresponding to the i-th input sample, and N represents the total number of samples for updating the weight parameters.
[0049] The total pruning amount can be set as the pruning rate (or the number of layers to be pruned), that is, the number of layers to be pruned that are pre-pruned divided by the number of all layers to be pruned specified. For example, the total pruning amount is: where n is the number of layers to be pruned that are pre-pruned in the deep neural network to be processed, and N is the number of all layers to be pruned in the deep neural network to be processed. In the actual dynamic iterative pruning process, only a certain proportion of the layers to be pruned need to be pre-pruned each time. For example, if the number of dynamic iterations is M times, then the pruning rate (corresponding to the pruning amount for each pre-pruning) for each iteration is R iis the actual pruning rate for the i-th iteration. According to an example of the present invention, it is assumed that the preset number set by the user is 5 times, and the number in the preset number refers to the number of times (Epoch) of training the deep neural network to be processed by fully utilizing the image dataset. The total pruning amount is 20%, and the number of pre-pruning times is 4. Then, after training every 5 Epochs, the network layers with the lowest 5% importance ranking are pre-pruned; after repeating 4 times, a total of 20% is pre-pruned.
[0050] Step S3: When the number of times of pre-pruning processing reaches the predetermined number of pre-pruning times, use the image dataset to fine-tune the pre-pruned deep neural network, and prune the network layer corresponding to the layer to be pruned that has been pre-pruned to obtain a compressed deep neural network.
[0051] According to an embodiment of the present invention, step S3 includes: S31. Use the image dataset to fine-tune the deep neural network after the last pre-pruning processing multiple times; S32. Remove the network layer corresponding to the layer to be pruned that has been pre-pruned from the deep neural network to obtain a compressed deep neural network.
[0052] Some deep neural networks may contain residual connections. If several network layers involved in the residual connection are removed, it will cause the dimensions of the two parts of data to be superimposed to be inconsistent (the dimension of one part of the data with several network layers removed becomes smaller, and the other part remains unchanged). Therefore, to eliminate the influence, according to an embodiment of the present invention, if the removed network layer has an impact on the data dimensions related to the residual connection, a zero-padding unit is added when removing the network layer. The zero-padding unit is used to pad zeros to the data corresponding to the residual connection in the compressed deep neural network according to the mask to eliminate the influence of the removed network layer on the data dimensions related to the residual connection.
[0053] According to an example of the present invention, during the compression process, the processing layer corresponding to the layer to be pruned with a mask of 1 and its parameters are retained, while the processing layer corresponding to the layer to be pruned with a mask of 0 is deleted; the actual operation method is to perform replacement according to the mask on the weight parameter variables of the model structure (such as nn.Parameter in pytorch). Since the residual connection requires the tensor dimensions to be consistent, if the removed network layer has an impact on the data dimensions related to the residual connection, zeros are supplemented at the output corresponding to the removed network layer through the zero-padding unit, so that problems such as dimension errors or dimension deviations will not occur during the residual calculation process.
[0054] For the convenience of understanding, the applicant gives the following examples of structural pruning according to the technical solution of the present invention on two existing deep neural networks.
[0055] Example 1: Structured Pruning of Vision Transformer Model Based on Importance (Salience Index and Activation Frequency)
[0056] This example uses a structured pruning scheme for deep learning models based on salience index to prune the Vision Transformer model. The whole process is as follows: First, construct the model structure of the model, initialize the weights and set the total pruning amount Suppose the pre-pruning process is performed M times, then the pruning amount for each pre-pruning Set the total number of training times to KM times and the predetermined number to K times. Then, perform a pre-pruning process every K times the model is trained with the image dataset. Among them, traverse the samples in the image dataset (it should be understood that it is set as the sample set used during training, that is, the training set), calculate the salience index values corresponding to each layer to be pruned, and the activation frequency of each pruning sample to obtain the importance of each layer to be pruned, and select the proportion or quantity with the lowest importance as R i The layers to be pruned with a mask value of 1 are set to be pre-pruned (the mask is changed from 1 to 0). Then, in K Epochs, retrain this model through stochastic gradient descent, and continue to calculate the importance of the pruned network on the image dataset during this period. Repeat the training and pre-pruning processes until the predetermined total pruning amount is reached.
[0057] Generally, the number of fine-tuning times (Epoch) after training can be set relatively large. When the model is pre-pruned, there can be a sufficiently large number of times to fine-tune the pruned model to recover the loss of accuracy. For example, first construct the structure of the Vision Transformer model (abbreviated as ViT model later). Refer to Figure 2 As shown, the ViT model includes multiple encoding modules (Blocks), and the encoding module is composed of a multi-head attention mechanism, a single-layer fully connected layer, a two-layer perceptron (MLP), and a residual connection stacked in sequence. Suppose each network layer corresponding to the multi-head attention mechanism, the linear fully connected layer, the first layer of the two-layer perceptron (MLP_FC1), and the second layer of the two-layer perceptron (MLP_FC2) in the ViT model is set as a layer to be pruned, add a mask to it, and judge the importance of each layer to be pruned by setting the salience index and activation frequency. During the model training process, the salience index of all layers to be pruned in the model will be updated with the training of the model, and the activation frequency will also be updated. Each pre-pruning uses the latest salience index and activation frequency to calculate the importance. Sort the importance of the layers to be pruned in descending order (from high to low), obtain the unit importance sorting results of all layers to be pruned, and find the layers to be pruned corresponding to the pruning amount (such as 5%) with the lowest importance ranking. Thus, according to the set pruning amount and the layers to be pruned, the ViT model can be processed for model training, pre-pruning, and compression.
[0058] The above process is executed in the operation process of dynamic iterative pruning. Among them, after training the model to reach the preset number of times, each time a certain amount or proportion of the layers to be pruned are pre-pruned; then the model is trained again, and the next pre-pruning process is carried out after reaching the preset number of times. This process is repeated multiple times until the last pre-pruning process is completed.
[0059] Since pruning will change the dimension of the residual output, directly adding the residuals will cause calculation errors. Therefore, it is necessary to set a zero-padding unit to pad the output of the pruned part of the network layer according to the mask to ensure the alignment of residual addition. The residual alignment method is referred to Figure 3 As shown, assume that the input of the encoding module of the vision transformer is 384-dimensional. After pruning, assume that the output dimension of the fully connected layer after pruning becomes 112-dimensional. Then, it is necessary to pad zeros according to the mask to output 384-dimensional data, and perform residual connection addition with the 384-dimensional data of the skip connection to output 384-dimensional data to avoid the problem of inconsistent dimensions at the residual addition.
[0060] The training environment or hyperparameters can be set according to the needs of the implementer. For example, the pruning of the ViT model is to build a model for training on 8 Nvidia RTX 3090 graphics cards. During training, the input size of the image is 224x224, and the batch size (Batch Size) of each graphics card is set to 8. To enhance the diversity of data and the generalization ability of the model, existing data augmentation methods such as random brightness change, random contrast change, random flipping, and random central cropping can also be used to enrich the samples in the image dataset. The AdamW optimizer is used, the momentum coefficient is set to 0.9, the weight decay (Weight Decay) is set to 0.0005, the initial learning rate is set to 0.0005, and annealing is performed according to the cosine annealing strategy (Cosine) of the learning rate. The total number of training times is set to 100 (pre-pruning is performed every 25 times of training), and the pruning rate can be set between 0 and 1 (for example, the pruning rate is 5% each time).
[0061] Example 2: Structured pruning of a convolutional network model based on importance
[0062] This example uses pruning of a convolutional network model. First, construct the convolutional network model structure, taking the VGG-16 model as an example. The basic unit of the VGG-16 model is a 3x3 convolutional layer, and the batch normalization layer (BN), activation layer ReLU, and pooling layer are stacked in sequence. The overall structure diagram is referred to Figure 4 .
[0063] This example modifies the model structures of the native torch.nn classes in the machine learning library Pytorch, such as the model structures of nn.Conv2d, nn.Linear, etc., and provides some interfaces (supporting components) that support structured pruning. Therefore, when building the VGG-16 model, these supporting components can be directly called in the form of interfaces to perform pruning and compression during model training. For example, nn.Conv2d of VGG-16 is a two-dimensional convolutional layer, which is replaced with a masked convolutional layer (MaskedConv), which is a convolutional layer that supports pre-pruning and masking operations, such as the calculation of saliency metrics, activation frequency calculation, importance calculation, and sorting, etc.
[0064] The training environment or hyperparameters, etc., can be set according to the needs of the implementer. Schematically, the pruning training is carried out on a Linux server with 4 NVIDIA GeForce RTX2080Ti graphics cards. The total number of training epochs is set to 60, the learning rate is initialized to 0.01, and then the learning rate is divided by 5 every 20 epochs. The data batch size is 128. The momentum coefficient is set to 0.9, and the weight decay is set to 0.0005.
[0065] Schematically, the original structure of VGG-16 is as Figure 4 shown, which is composed of multiple convolutional modules (each convolutional module contains one or more convolutional layers, and the convolutional layer is the smallest pruning unit), multiple pooling layers, and multiple fully connected layers stacked. Conv1-5 represents the numbers of 5 convolutional modules, and the number after the symbol "-" represents the serial number of the convolutional layer in the convolutional module. For example, Conv1-1 represents the convolutional layer 1 in convolutional module 1, and the rest are similar, which will not be elaborated here. Assume that through the method of the present invention, it is determined that the convolutional layer Conv5-1 and the convolutional layer Conv5-2 need to be pruned, then the final compressed model is as Figure 5 shown.
[0066] According to an embodiment of the present invention, a structured pruning system is provided, including: a model modeling interface for loading a deep neural network and receiving a user instruction to set a corresponding network layer in the deep neural network as a layer to be pruned, so as to obtain a deep neural network to be processed; a pruning module for performing the foregoing method of structured pruning to obtain a compressed deep neural network. According to an embodiment of the present invention, a mask is provided for the layer to be pruned, and the output of the layer to be pruned is set to the product of the output of the network layers it contains and the mask. Among them, the mask indicates whether the layer to be pruned has been pre-pruned. A mask of 0 indicates that it has been pre-pruned, and a mask of 1 indicates that it has not been pre-pruned. All masks are initially set to 1. In the model modeling interface, recognition components for a variety of common deep neural networks can be preset to recognize the types of network layers contained in the deep neural network and add masks to the specified network layer according to the received user instruction. For example: the model modeling interface can provide function support for replacing corresponding network layers (such as basic convolutional layers and basic linear fully connected layers), support mask operator calculation and mask acquisition for underlying feedforward calculations; this interface can be called during the model modeling stage to complete the modeling of the deep neural network to be processed (supporting the model modeling of the pruning version); and it supports the calculation of saliency metrics and the interface for activation frequency. The pruning module is used to provide a model training function and perform structured pruning (including processing such as training, importance calculation, pre-pruning, fine-tuning, and compression) according to the set pruning target (such as the total pruning amount). The technical solution of this embodiment can at least achieve the following beneficial technical effects: the model modeling interface of the present invention can load a deep neural network and customize the layer to be pruned according to the user instruction, so as to realize tooling by calling the interface, reduce costs and improve efficiency for compressing general deep neural networks.
[0067] According to an embodiment of the present invention, a method for deploying a deep neural network includes: obtaining a compressed deep neural network, where the compressed deep neural network is obtained by using the method of structured pruning or the structured pruning system described in the foregoing embodiment; deploying the compressed deep neural network on a mobile device for performing corresponding prediction tasks. The technical solution of this embodiment can at least achieve the following beneficial technical effects: the compressed deep neural network obtained by the technical solution of the present invention can better maintain the model accuracy, greatly reduce the number of parameters and the amount of calculation, and significantly improve the inference performance efficiency.
[0068] According to an embodiment of the present invention, an electronic device includes: one or more processors; and a memory, where the memory is used to store executable instructions; the one or more processors are configured to implement the steps of the method of structured pruning or the method of deploying a deep neural network by executing the executable instructions.
[0069] In summary, the present invention provides a structured pruning method and system, which solve the problems that the parameter redundancy of general deep learning models such as vision transformers and convolutional neural networks leads to high requirements for resources such as computing power, memory, and storage of mobile devices and slow running speed, and that the previous model pruning methods have a large loss of model accuracy and require a large amount of parameter tuning during the pruning process. Compared with the existing model compression methods, the present invention has the advantages of fast model accuracy recovery and better guarantee of model accuracy.
[0070] It should be noted that although the above steps are described in a specific order, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a changed order, as long as the required functions can be achieved.
[0071] The present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0072] The computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. The computer-readable storage medium may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing.
[0073] The various embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the disclosed embodiments.
Claims
1. A method of structured pruning for pruning a deep neural network, the deep neural network comprising a plurality of network layers, characterized in that, The method includes: S1. Set a specified network layer in the deep neural network as the layer to be pruned. Each layer to be pruned corresponds to a network layer, and a deep neural network to be processed is obtained. S2. Use the image dataset to train the deep neural network to be processed multiple times and perform multiple pre-pruning processes during the training process. Among them, each pre-pruning process includes: Determine the importance of all layers to be pruned on the image dataset. Set multiple layers to be pruned with relatively low importance rankings and that can meet the pruning amount of each pre-pruning after pre-pruning as the pre-pruned layers. And the elements contained in the output of the pre-pruned layer to be pruned are all set to 0 in subsequent training. Among them, the importance of each layer to be pruned on the image dataset is determined according to a predetermined calculation rule based on the weight parameters, gradients, and activation frequencies of the layer to be pruned. The activation frequency is the number of times the output size of the layer to be pruned exceeds a predetermined activation threshold during training. S3. When the number of pre-pruning processes reaches a predetermined number of pre-pruning times, use the image dataset to fine-tune the pre-pruned deep neural network to be processed, and perform pruning processing on the network layer corresponding to the pre-pruned layer to be pruned, and a compressed deep neural network is obtained.
2. The method according to claim 1, wherein The importance of each layer to be pruned on the image dataset is determined according to the following calculation rule: According to the latest weight parameters and gradients of the network layer corresponding to the layer to be pruned, determine the saliency index. Calculate the weighted sum of the saliency index and the activation frequency of the layer to be pruned to obtain the importance of the layer to be pruned on the image dataset.
3. The method according to claim 2, wherein The activation frequency is determined in the following manner: Determine the output size corresponding to the output of each layer to be pruned when each sample in the image dataset is input into the deep neural network to be processed. Take the number of times the output size of the layer to be pruned exceeds the predetermined activation threshold in the latest completed preset number of trainings as the activation frequency. Among them, the output size is obtained by taking the square root of the sum of the squares of the elements contained in the output, and the activation threshold is the average value of all output sizes in the latest completed preset number of trainings.
4. The method according to claim 2, wherein The saliency index is determined in the following manner: Multiply each corresponding element in the latest weight parameters and gradients in the network layer corresponding to the layer to be pruned and then sum to obtain the saliency index of the layer to be pruned.
5. The method according to any one of claims 1-4, characterized in that, The elements contained in the output of the pre-pruned layer to be pruned are all set to 0 in subsequent training in the following manner: Set a mask for each layer to be pruned. The output of the layer to be pruned is the element-wise product of the elements in the output of its corresponding network layer and its mask. Among them, the mask indicates whether the layer to be pruned is pre-pruned. The mask being 0 indicates pre-pruning, and the mask being 1 indicates not being pre-pruned. All masks are initially set to 1.
6. The method according to any one of claims 1-4, characterized in that Step S3 includes: Use the image dataset to fine-tune the deep neural network after the last pre-pruning process multiple times. Remove the network layer corresponding to the pre-pruned layer to be pruned from the deep neural network to obtain a compressed deep neural network.
7. A structured pruning system, characterized in that, It includes: A model modeling interface for loading the deep neural network and receiving a user instruction to set a corresponding network layer in the deep neural network as the layer to be pruned to obtain a deep neural network to be processed. A pruning module, which is used to execute the structured pruning method described in any one of claims 1-6, and obtain a compressed deep neural network.
8. A method for deploying a deep neural network, characterized in that, It includes: Obtain a compressed deep neural network, which is obtained by using the method described in any one of claims 1-6 or the system described in claim 7; Deploy the compressed deep neural network on a mobile device for performing corresponding prediction tasks.
9. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method described in any one of claims 1 to 6 or 8.
10. An electronic device, characterized in that, It includes: One or more processors; And A memory, where the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method described in any one of claims 1 to 6 or 8 by executing the executable instructions.
Citation Information
Patent Citations
Flexible deep learning network model compression method based on channel gradient pruning
CN112396179A
Global rank perception neural network model compression method based on filter feature map
CN114037844A