A knowledge distillation algorithm for neural network training
Through the knowledge distillation algorithm, feature information is extracted from complex models and guided simple model training, which solves the problem of deep learning network storage and computing resource consumption in computer vision tasks, and achieves performance improvement and resource conservation.
Patent Information
- Application Number
- CN202111329117.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-11-10
AI Technical Summary
Deep learning networks have problems with high storage space and computing resource consumption in computer vision tasks, and it is difficult to effectively deploy on hardware platforms. How to maintain performance while reducing the amount of model parameters.
Using the knowledge distillation algorithm, a simple student model training is guided by extracting useful feature information from complex teacher models, using attention mechanisms and activation information transfer modules, and designing feature conversion modules and loss functions to optimize the performance of the student model.
While keeping the model parameter quantity unchanged, the image classification performance of the simple model is significantly improved, the error rate is reduced, and the generalization ability of the model is improved.
Smart Images

Figure CN114169495B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural networks, and specifically relates to an algorithm for improving the performance of neural network models. Background Art
[0002] The performance of deep learning networks in computer vision tasks has reached an unprecedented height. However, at the same time, complex models bring high storage space and computational resource consumption, and it is difficult to be implemented on various hardware platforms. How to reduce the number of model parameters as much as possible while maintaining the performance of the model has become a problem to be solved, and the solution to this problem plays an important role in the deployment of deep learning networks. In the present invention, a method of multi-angle feature transfer is adopted to achieve the goal of improving the model performance while maintaining the model size. Summary of the Invention
[0003] The purpose of the present invention is to propose an algorithm for improving the performance of neural network models with relatively low complexity, so that the model can improve its own performance while maintaining the number of parameters.
[0004] In thermodynamics, "distillation" is a method of extracting the target liquid from a mixture. By analogy, useful feature information is extracted from a complex teacher model and transferred to a student model, and the student model learns useful knowledge to improve its own model performance. This method is called "knowledge distillation". The present invention provides an algorithm for using the activation information and attention mechanism information of a complex teacher model to improve the performance of a neural network model with relatively low complexity, which can be abbreviated as "a knowledge distillation algorithm for neural network training", and mainly involves a teacher model, a student model, and a feature conversion module; the feature transfer module mainly includes two parts: an attention mechanism transfer module (AT) and an activation information transfer module (AC). See Figure 1 as shown;
[0005] After feature extraction by the convolutional network, the input information X passes through convolution and the activation function σ to obtain the output information Y i :
[0006]
[0007] Use the sum of the squares of Y i to compress the information of C channels and place them on one channel for a compact spatial representation:
[0008]
[0009] Formula (2) is the processing of the attention mechanism transfer module (AT) for feature information. Through this method, the information of C channels can be compressed into one channel, so that a three-dimensional spatial vector can be transformed into a two-dimensional vector. Thus, the information of the attention mechanism transfer module is obtained.
[0010] For the transfer of activation information, a variant of Relu is designed as the activation function. As Figure 2 (b) shows, where m is the expected value of the negative response information:
[0011]
[0012] The output of the activation information transfer module (AC) is:
[0013]
[0014] The transfer of feature information is manifested as a loss function in the actual model training. By continuously reducing the value of the loss function, the student model's feature information can better imitate and learn the useful information of the teacher model. Therefore, a loss function for this network is designed.
[0015] The loss function L of the attention mechanism transfer module (AT) AT (T, S) is the mean square error of the attention mechanism information between the teacher model (T) and the student model (S):
[0016] L AT (T, S) = |F sum (Q(T)) - F sum (Q(S))| 2 , (5)
[0017] The loss function L of the activation information transfer module (AC) AC (T, S) is the mean square error of the activation information between the teacher and student models:
[0018] L AC (T, S) = |max(M, Q(T)) - max(M, Q(S))| 2 , (6)
[0019] The loss function L of the attention mechanism transfer module (AT) AT (T, S) and the loss function L of the activation information transfer module (AC) AC (T, S) have a ratio of 1:2, and the total loss of feature information transfer is:
[0020] L F = L AT + 2L AC , (7)
[0021] According to the above design, the algorithm provided by the present invention for improving the performance of a simple neural network model basically transfers the useful feature information of a complex teacher model to guide the student model to learn useful feature information through the transfer of activation information and attention mechanism information. The specific steps are as follows:
[0022] (1) First, determine the simple model to be trained, which is defined as the student model here; then select a model with higher complexity, specifically characterized by having more parameters than the student model and better performance on the target task than the student model, which is defined as the teacher model. The entire training process includes three parts: forward propagation, feature information transfer, and backpropagation.
[0023] (2) During the forward propagation process, as shown in formula (1), the input image information of the network undergoes mathematical multiplication and summation operations with the network weights to obtain the information of each feature layer and the final output information of the network. The forward propagation is carried out separately in the student model and the teacher model and does not affect each other.
[0024] (3) The transfer process of feature information mainly includes the transfer of two parts: attention mechanism information and activation information. The acquisition of attention mechanism information is as shown in formula (2), and the acquisition of activation information is as shown in formulas (3) and (4). The transfer of these two types of information is mainly manifested as the acquisition of the loss function, as shown in formulas (5) and (6). The loss function mainly includes two parts: attention mechanism information and activation information, and their weighted sum is as shown in formula (7).
[0025] (4) During the backpropagation process, mainly the information of the loss function is backpropagated. The main method is to use the chain rule to obtain the transformation value of the gradient information of each convolutional layer, and then add this transformation value to the original network weights to update the network parameters in this way. This process only updates the weight information of the student model, while the weight information of the teacher model remains unchanged.
[0026] (5) Repeat the steps (2), (3), and (4) according to the preset number of training iterations until the loss function value of the network no longer changes, and complete the entire network training process. Description of the Drawings
[0027] Figure 1 Overall architecture diagram of the method of the present invention.
[0028] Figure 2 Proposed activation function diagram. Detailed Embodiment
[0029] This section will further elaborate on how to use the feature information transfer method of the present invention to train a simple model.
[0030] Assume that the student model is MobileNet with a network parameter count of 4.23M, which has significantly fewer parameters compared to the corresponding teacher model Resnet50 (25.56M).
[0031] (1) As Figure 1 , the input to the network is an image. All convolutional layers of the student model MobileNet and the teacher model Resnet50 are evenly divided into four blocks, and the width and height of the output position features of each block are the same;
[0032] (2) Apply batch normalization and a 1×1 convolution operation to the output positions of each block of MobileNet to make the number of channels of the student model and the teacher model consistent, thus preparing for the following feature transfer;
[0033] (3) The input image can be regarded as a three-dimensional vector with RGB color channels, width W, and height H. This three-dimensional vector undergoes operations such as convolution, pooling, and activation function transformation with pre-set parameters in the student model MobileNet and the teacher model Resnet50 respectively, and the forward pass of the network is completed separately in the two branches of MobileNet and Resnet50;
[0034] (4) Obtain activation information and attention mechanism information at the output positions of each block of MobileNet and Resnet50. For activation information, use the sum of squares of corresponding positions of each feature channel to compress the information of all feature channels and place them on one channel for a compact spatial representation. For attention mechanism information, perform a non-linear transformation of the feature information, and design a variant of Relu as the activation function, where the threshold of Relu is the expected value of the negative response information. MobileNet and Resnet50 extract activation information and attention mechanism information at the corresponding feature output positions of the four blocks, and use the squared difference of the activation information and attention mechanism information at the corresponding positions as the final loss function value. This step completes the transfer process of the network's feature information;
[0035] (5) During the backpropagation process, use the chain rule to reverse-derive the value of the loss function to obtain the gradient update value of each layer of the network, and add it to the original network weights to update the network weights. This process is only performed on the student, while the teacher model keeps the original network weights unchanged.
[0036] (6) Repeat the three processes (3), (4), and (5) 200 times, and MobileNet converges.
[0037] During the model testing process, the teacher model Resnet50 does not participate. Only the forward propagation process of step (3) is performed on the student model. Therefore, the number of parameters of the network participating in the test remains the original number of parameters of MobileNet, which is 4.23M.
[0038] The Imagenet dataset has the advantages of a wide range of categories and a large amount of data, and can well detect the generalization ability of the model. Therefore, the present invention uses Imagenet as the test dataset. Using the proposed training method to test Imagenet, MobileNet can obtain an image classification performance of 28.21% (Top1 error rate) and 9.23% (Top5 error rate). Without using the proposed method for training, the image classification performance is 31.13% (Top1 error rate) and 11.24% (Top5 error rate). The Top1 error rate and the Top5 error rate are two metrics used to measure the image classification ability. The smaller their values, the better the performance of the model. The training method of the present invention can obtain a performance gain of 2.92% (Top1 error rate) and 2.01% (Top5 error rate) compared with the traditional training method, thereby proving the significant superiority of the model of the present invention.
Claims
1. A knowledge distillation method for training an image classification neural network, characterized in that, It involves a teacher model, a student model, and a feature transformation module; the feature transfer module includes an attention mechanism transfer module (AT) and an activation information transfer module (AC); After feature extraction by the convolutional network, the input image X passes through convolution and the activation function σ to obtain the output information Y i : Among them, are the convolution weights, and b i is the convolution bias. T represents the matrix transpose operation, and i represents the number of channels of the convolution; Use Y i to compress the information of C channels by the sum of squares and place them on one channel for a compact spatial representation: Formula (2) is the processing of the attention mechanism transfer module (AT) for feature information, which compresses the information of C channels into one channel, turning a three-dimensional spatial vector into a two-dimensional vector; thus, the information of the attention mechanism transfer module is obtained; For the transfer of activation information, a variant of Relu is designed as the activation function; where m is the expected value of the negative response information: What follows "|" is the judgment condition, that is, values less than 0 are filtered once, and then the expectations of these values greater than 0 are judged; The output of the activation information transfer module (AC) is: The transfer of feature information is manifested as a loss function in the actual model training. By continuously reducing the value of the loss function, the student model can better imitate and learn the useful information of the teacher model; for this purpose, a loss function for this network is designed: Attention mechanism transfer module (AT) loss function L AT (T, S) is the mean squared error of the attention mechanism information between the teacher model (T) and the student model (S): L AT (T, S) = |F sum (Q(T)) - F sum (Q(S))| 2 , (5) Among them, Q(T) and Q(S) are the attention mechanism information of the teacher model and the student model obtained from Formula 2 respectively; Activation Information Transfer Module (AC) Loss Function L AC (T, S) is the mean squared difference of the activation information between the teacher and student models: L AC (T,S) = |max(M, Q(T)) - max(M, Q(S))| 2 , (6) Wherein, M is a threshold value, which is the average value of Q(S) and Q(T); the ratio of the loss function L AT (T, S) of the attention mechanism transfer module (AT) to the loss function L AC (T, S) of the activation information transfer module (AC) is 1:2, then the total loss of feature information transfer is: L F = L AT + 2L AC , (7) Transfer the useful feature information of the complex teacher model to guide the student model to learn useful feature information through the transfer of activation information and attention mechanism information.
2. The knowledge distillation method for training an image classification neural network according to claim 1, wherein The specific steps are as follows: (1) First, determine the simple model to be trained, defined as the student model; then select a model with higher complexity, which is characterized by having more parameters than the student model and better performance on the target task than the student model, defined as the teacher model. The entire training process includes three parts: forward propagation, feature information transfer, and backward propagation; (2) During the forward propagation process, as shown in Formula (1), the input image information of the network undergoes mathematical multiplication and addition operations with the network weights to obtain the information of each feature layer and the final output information of the network; the forward propagation is carried out separately in the student model and the teacher model and does not affect each other; (3) During the feature information transfer process, it includes the transfer of two parts: attention mechanism information and activation information; The attention mechanism information is obtained as shown in Formula (2), and the activation information is obtained as shown in Formulas (3) and (4); the transfer of these two types of information is specifically manifested as the acquisition of the loss function, as shown in Formulas (5) and (6). The loss function includes two parts: attention mechanism information and activation information, and their weighted sum is as shown in Formula (7); (4) During the backward propagation process, the information of the loss function is backpropagated. The method is to use the chain rule to calculate the derivative to obtain the transformed value of the gradient information of each convolutional layer, and then add this transformed value to the original network weights to update the network parameters through this method; Only the weight information of the student model is updated in this process, while the weight information of the teacher model remains unchanged; (5) Repeat steps (2), (3), and (4) according to the preset number of training iterations until the value of the loss function of the network no longer changes, completing the entire network training process.
Citation Information
Patent Citations
Neural network weight initialization method based on transfer learning
CN111126599A
Knowledge distillation method, device and equipment based on multi-layer multi-attention migration
CN113326941A