Adaptive Distillation Method for N:M Sparse Fine-tuning
The adaptive distillation method replaces the activation value with large differences in the student model, and combines the loss function of Kullback-Leibler divergence and average error, solving the problems of low efficiency and performance degradation of N:M sparse fine-tuning, achieving more efficient and accurate fine-tuning effects.
Patent Information
- Application Number
- CN202311122615.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-01
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-09-01
AI Technical Summary
The prior art is inefficient and has severe performance degradation in N:M sparse fine-tuning, especially in the case of high sparse rates and high-parameter efficient models.
An adaptive distillation method is proposed to generate a mask based on the differences in the feature maps of the student model and the teacher model, replace the activation values with larger differences in the student model, reduce the difficulty of learning the student model, and use the Kullback-Leibler divergence and mean error as the loss function to guide fine-tuning.
The rounds of N:M sparse fine-tuning are significantly reduced, and the fine-tuning efficiency and accuracy are improved. Especially in the case of high sparseness and high-parameter efficient models, the performance recovery is more significant.
Smart Images

Figure CN117172293B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to compression and acceleration of artificial neural networks, and in particular to an adaptive distillation method for N:M sparse fine-tuning. Background Art
[0002] The development of deep neural networks (DNNs) is inseparable from the ever-expanding memory footprint and computational requirements of DNNs. However, this often leads to challenges in training and deploying DNNs on resource-limited devices. In order to solve this problem, a large number of model compression methods have been proposed, including knowledge distillation [Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. In NeurIPS Workshops, 2014.], parameter quantization [Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, Joseph Naor, and Daniel Soudry. Accelerated sparse neural training: A provable and efficient method to find N:M transposable masks. In NeurIPS, 2021.], efficient model structure [Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.] and network sparseness [Yann LeCun, John Denker, and Sara Solla. Optimal brain damage. In NeurIPS, 1989.]. Network sparsity reduces model size by eliminating redundant parameters in DNN. Traditional network sparsity methods can be roughly divided into two categories: unstructured sparsity and structured sparsity. Unstructured sparsity can remove individual weights at any position in the DNN to achieve fine-grained compression, so robust performance can be maintained even at high sparsity rates. However, such fine-grained weight removal usually leads to irregular sparse matrices, which requires a lot of additional space to store the index matrix, resulting in limited acceleration on existing hardware. Unstructured sparsity can obtain a coarse-grained compressed network by eliminating entire weight blocks or convolution filters.Although unstructured sparsity has actual speedups, at similar sparsity rates, structured sparsity suffers greater performance degradation than unstructured sparsity [Alex Renda, Jonathan Frankle, and Michael Carbin. Comparing rewinding and fine-tuning in neural network pruning. In ICLR, 2020].
[0003] N:M fine-grained sparsity is a promising sparsity pattern that shows an excellent balance between acceleration effect and performance preservation. N:M fine-grained sparsity requires that at most N weights be retained for every M consecutive weights, making it possible to obtain actual acceleration on N:M sparse tensor cores [Olivier Giroux et al. Ronny Krashinsky. Nvidia ampere sparse tensorcore. https: / / developer.nvidia.com / blog / nvidia-ampere-architecture-in-depth / , 2020.]. NVIDIA proposed the Automatic Sparsity Process (ASP) [Asit Mishra, Jorge AlbericioLatorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius. Accelerating sparse deep neural networks. arXiv preprint arXiv:2104.08378, 2021.], which achieves 2:4 sparsity in three steps: (1) pre-training dense weights; (2) pruning based on the weight size of the pre-trained model; (3) fine-tuning the sparse weights using the training strategy in the pre-training phase. Since most pre-trained models are directly available on the Internet, the training cost of ASP mainly comes from the last step. Unfortunately, the fine-tuning phase is very heavy because it is equal to the number of training rounds of dense training. In addition, ASP will cause severe performance degradation in the following two cases: (1) pruning with high sparsity rate and (2) pruning parameter-efficient models. Therefore, it is crucial to improve the efficiency and performance of N:M sparse fine-tuning. One feasible way to alleviate the above problems is knowledge distillation (KD), which transfers knowledge from a larger teacher model (dense network) to a compact student model (sparse network). In fact, it has been documented that both logic-based knowledge distillation and feature map-based knowledge distillation can bring performance gains to unstructured sparse. Experiments have shown that logic-based knowledge distillation can bring performance gains to N:M sparse fine-tuning, but the effect of feature map-based knowledge distillation is limited.
[0004] Directly applying the knowledge distillation method to N:M sparse fine-tuning does not take into account the N:M sparse characteristics. Analysis shows that the number of parameters used by N:M sparse to generate each independent activation value is fixed compared to unstructured sparse, resulting in the inability to allocate the number of learnable parameters according to the difficulty of learning the activation value. Therefore, the effect of feature distillation is limited. In combination with the characteristics of N:M sparse, the present invention proposes an adaptive distillation method for N:M sparse fine-tuning. Summary of the invention
[0005] The purpose of the present invention is to provide an adaptive distillation method for N:M sparse fine-tuning, which makes N:M sparse fine-tuning faster and more accurate. Given a pre-trained model, this method can be used to obtain an N:M sparse model, and the number of fine-tuning rounds is much less than that of the traditional N:M workflow. Under the same number of fine-tuning rounds, this method can be used to obtain a more accurate N:M sparse model.
[0006] The present invention comprises the following steps:
[0007] 1) Prune the pre-trained model using the N:M pruning scheme based on weight size;
[0008] 2) Adaptive knowledge distillation is used to transfer the knowledge in the pre-trained model to the N:M sparse model. Adaptive knowledge distillation can automatically replace the activation values with large differences in the student feature map with the activation values in the teacher feature map according to the difference in activation values between the student model feature map and the teacher model feature map, thereby reducing the difficulty of student model learning and accelerating the convergence speed of student model fine-tuning.
[0009] In step 1), the specific steps of the weight-based N:M pruning scheme may be:
[0010] (1) Weight grouping: Divide the weights in the neural network into groups of M;
[0011] (2) Weight sparseness: Based on the absolute value of the weight in each group, retain the N weights with the largest absolute values.
[0012] In step 2), the specific steps of the adaptive knowledge distillation include:
[0013] (1) Input the same batch of images into the teacher model (dense model) and the student model (N:M sparse model) respectively, and obtain the logical prediction p of the teacher model and the student model respectively. t and p s , and the feature map F of the penultimate layer of the teacher model and the student model t and F s ;
[0014] (2) According to F t and Fs The difference is calculated to obtain the mask M;
[0015] (3) Use the mask to obtain the synthesized student feature map F m =M⊙F t +(1–M)⊙F s ;
[0016] (4) Calculate p t and p s The Kullback-Leibler divergence of t and F s The sum of these two items is used as the loss of the student model gradient back propagation to guide the fine-tuning of the student model.
[0017] Advantages and effects of the present invention:
[0018] The present invention can be applied to popular deep neural network models, including ResNet, ResNext, MobileNet-V2 and other models, to obtain corresponding N:M sparse models, and the number of fine-tuning rounds is less than that of the traditional N:M sparse process, and the accuracy of the fine-tuned model is higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 Schematic diagram of the method of the present invention.
[0020] Figure 2 Schematic diagram of fine-tuning convergence speed of the present invention. (a) is ResNet-50, and (b) is ResNext-50-32x4d. DETAILED DESCRIPTION
[0021] The following embodiments will further illustrate the present invention in conjunction with the accompanying drawings.
[0022] The motivation of the present invention is that the overhead of traditional N:M sparse fine-tuning is very high, because the number of rounds of fine-tuning is equal to the number of rounds of intensive training and when the sparsity rate is very high or the model itself is parameter-efficiently designed, the sparse model obtained by traditional N:M sparse fine-tuning suffers from serious loss of accuracy. Since knowledge distillation can transfer knowledge from models with strong capabilities to models with weak capabilities, and has been successfully applied to unstructured sparse fine-tuning, knowledge distillation can be introduced in the N:M sparse fine-tuning stage. However, experiments have found that traditional knowledge distillation methods cannot be generalized well to N:M sparse fine-tuning. To address this problem, adaptive knowledge distillation is proposed to reduce the difficulty of student model learning. The algorithm framework diagram of an embodiment of the present invention is shown in the figure. Figure 1As shown in the figure, the adaptive distillation method can determine which activation values to learn based on the difference between the feature maps of the student model and the teacher model to reduce the difficulty of feature distillation. The final loss function consists of the Kullback-Leibler divergence of the predicted distribution of the teacher model and the student model and the average error loss of the teacher model and the student model features.
[0023] TK Distillation Instructions:
[0024] Knowledge distillation can transfer knowledge from a strong teacher model to a weak student model, improving the accuracy of the student model. A simple knowledge distillation method allows the student model to imitate the logical output of the teacher model by matching the logical output of the teacher model and the student model. The distillation loss expression is as follows:
[0025]
[0026] Among them, p s and p t are the logistic outputs of the student model and the teacher model, σ(·) is the normalized exponential function, T is the distillation temperature, is the Kullback-Leibler divergence used to match the distribution of the logistic outputs of the student model and the teacher model.
[0027] Another commonly used knowledge distillation method is to perform knowledge distillation by migrating the features in the teacher model. The loss expression is as follows:
[0028]
[0029] in, and are the feature maps of the student model and the teacher model, respectively. B, N, H, and W represent the batch size, the number of channels, the height of the feature map, and the width of the feature map, respectively. The distance between two feature maps can be calculated, which can be a Euclidean distance function.
[0030] Adaptive Knowledge Distillation:
[0031] For N:M sparse models, ordinary feature distillation has limited effect. In order to make feature distillation better used in N:M sparse fine-tuning, the present invention adopts adaptive knowledge distillation. Unlike traditional feature distillation that allows the student model to directly learn the activation values in the feature map generated by the teacher model, the present invention first reduces the difficulty of student model learning by replacing the activation values with large differences in the student feature map and the teacher feature map with the activation values corresponding to the teacher feature map. The present invention uses a mask to complete the above operation, and the specific formula is as follows:
[0032] F m =M⊙Ft +(1-M)⊙F s (3)
[0033] Among them, F m is the student feature map after the replacement operation, M∈{0,1} B×N×H×W It is an adaptive mask. When the value of M is 1, it means that the activation value of the student feature map is replaced with the activation value of the corresponding position in the teacher feature map. When the value of M is 0, it means no replacement. The mask M is generated adaptively. It mainly refers to the difference ΔF between the teacher feature map and the student feature map. The calculation formula is as follows:
[0034] ΔF=|F t -F s |(4)
[0035] The present invention reduces the difficulty of feature distillation by replacing the activation values in the student feature map that are significantly different from those in the teacher feature map. Therefore, the generation formula of the adaptive mask M is as follows:
[0036]
[0037] Among them, b,n,h,w traverse B,N,H,W. k represents the number of activation value replacements, which is calculated by the following formula:
[0038]
[0039] in, is the rounding down operation, and Cos(·,·) is the cosine similarity calculation function. It can be seen that k and F t and F s When the distance between two feature maps is large, it means that there are many different activation values in the feature maps, so a larger k is needed to ensure stable training of the student model. On the contrary, when the distance in the feature maps decreases, a smaller k helps the student model extract knowledge from the teacher feature map.
[0040] Overall loss function:
[0041] Combining the knowledge distillation based on logical output in formula (1) and the adaptive feature distillation, the present invention finally uses the following loss function to fine-tune the student model:
[0042]
[0043] in, Calculate the mean square error of the two feature maps. α and β are hyperparameters that balance the two losses.
[0044] The performance comparison between the present invention (AdaDistil) and the traditional N:M fine-tuning method (ASP) on the CIFAR-10 dataset is shown in Table 1.
[0045] Table 1
[0046]
[0047] As can be seen from Table 1, in all N:M sparse patterns, the performance of AdaDistil proposed in the present invention is significantly better than ASP. For example, AdaDistil only needs one-third of the fine-tuning rounds of ASP to recover the performance degradation caused by N:M sparsity. And the method of the present invention achieves higher accuracy under the same fine-tuning rounds for all N:M sparse patterns. Specifically, the Top-1 accuracy of AdaDistil at 1:4 and 1:16 sparsity rates is 1.20% and 1.38% higher than ASP, respectively. These results well demonstrate the effect of the method proposed in the present invention on small data sets.
[0048] The performance comparison between the present invention (AdaDistil) and the traditional N:M fine-tuning method (ASP) on the ImageNet dataset is shown in Table 2.
[0049] Table 2
[0050]
[0051] As can be seen from Table 2, for the ResNet-50 network, the method of the present invention (AdaDistil) achieves accuracy comparable to that of the ASP method using only half the number of fine-tuning rounds in the 2:4 and 1:4 sparse modes. In addition, the method proposed in the present invention achieves higher accuracy using the same fine-tuning rounds at high sparsity rates (such as 1:4 and 1:16). For example, for ResNet-50, the Top-1 classification accuracy of AdaDistil in the 1:4 and 1:16 sparse modes is 0.29% and 0.36% higher than that of ASP, respectively. In order to verify the effectiveness of AdaDistil in different model architectures, the effect of the AdaDistil method is further evaluated on another large network (ResNext-50-32x4d) and two representative parameter-efficient networks (such as MobileNet-V2 and EfficientNet-B0). The results show that AdaDistil of the present invention outperforms the traditional ASP method on all network architectures. These experimental structures strongly demonstrate the effectiveness of the method of the present invention. Convergence speed is an important indicator for evaluating the efficiency of a method. Figure 2The curves of the accuracy of the present invention on ResNet-50 and ResNext-50-32x4d as the number of fine-tuning rounds are given. The results show that the convergence speed of the method proposed in the present invention is much better than the traditional N:M fine-tuning method.
Claims
1. An adaptive distillation method for N:M sparse fine-tuning, characterized by The following steps are involved: 1) Prune the pre-trained model using the N:M pruning scheme based on weight size; 2) Adaptive knowledge distillation is used to transfer the knowledge in the pre-trained model to the N:M sparse model. Adaptive knowledge distillation automatically replaces the activation values in the student feature map with the activation values in the teacher feature map according to the difference in activation values between the student model feature map and the teacher model feature map, so as to reduce the difficulty of student model learning and speed up the convergence speed of student model fine-tuning. The specific steps of the adaptive knowledge distillation include: (1) Input the same batch of images into the teacher model and the student model respectively, and obtain the logical prediction p of the teacher model and the student model respectively. t and p s , and the feature map F of the penultimate layer of the teacher model and the student model t and F s ; (2) According to F t and F s The difference is calculated to obtain the mask M; (3) Use the mask to obtain the synthesized student feature map. The specific formula is as follows: F m =M⊙F t +(1–M)⊙F s Among them, F m is the student feature map after the replacement operation, M∈{0,1} B×N×H×W It is an adaptive mask. When the value of M is 1, it means that the activation value of the student feature map is replaced with the activation value of the corresponding position in the teacher feature map. When the value of M is 0, it means no replacement. The mask M is adaptively generated, which is mainly based on the difference ΔF between the teacher feature map and the student feature map. The calculation formula is as follows: ΔF=|F t -F s | The difficulty of feature distillation is reduced by replacing the activation values in the student feature map that are significantly different from those in the teacher feature map. The generation formula of the adaptive mask M is as follows: Where b, n, h, w traverse B, N, H, W; k represents the number of activation value replacements, which is calculated by the following formula: in, is the rounding down operation, Cos(·,·) is the cosine similarity calculation function; k and F t and F s When the distance between two feature maps is large, it means that there are many activation values with large differences in the feature maps, so a larger k is needed to ensure stable training of the student model; on the contrary, when the distance in the feature maps decreases, a smaller k helps the student model extract knowledge from the teacher feature maps; Combining formulas The knowledge distillation based on logical output and the adaptive feature distillation in s and p t are the logistic outputs of the student model and the teacher model, σ(·) is the normalized exponential function, T is the distillation temperature, The Kullback-Leibler divergence is used to match the distribution of the student model and the teacher model logistic output; finally, the following loss function is used to fine-tune the student model: in, Calculate the mean square error of the two feature maps. α and β are hyperparameters that balance the two losses. (4) Calculate p t and p s The Kullback-Leibler divergence of t and F s The sum of these two items is used as the loss of the student model gradient back propagation to guide the fine-tuning of the student model.
2. The adaptive distillation method for N:M sparse fine-tuning according to claim 1, characterized in that In step 1), the weight-based N:M pruning scheme comprises the following specific steps: (1) Weight grouping: Divide the weights in the neural network into groups of M; (2) Weight sparseness: Based on the absolute value of the weight in each group, retain the N weights with the largest absolute values.