An image classification method based on improved online knowledge distillation algorithm
By setting up the wrong question set module and reflection mechanism in the online knowledge distillation method, the feature learning of students' network is optimized, and the problem of insufficient supervision information in the existing online knowledge distillation method is solved, and the accuracy of image classification is improved.
Patent Information
- Application Number
- CN202210183421.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-11
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-02-11
AI Technical Summary
The existing online knowledge distillation methods lack stable and reasonable supervision information, which leads to inaccurate supervision information generated by the student network during preliminary training, affecting the performance of the distillation network, especially in image classification tasks, with low classification accuracy.
By setting up the wrong question set module to save the integrated output historical information of the student network, and using the reflection mechanism, the output distribution of the student network is far away from the wrong historical information and close to the correct historical information, thereby optimizing the feature learning of the student network.
By using the wrong question set module and reflection mechanism, students are forced to learn better features online, the accuracy of image classification is improved, and more accurate classification results are achieved.
Smart Images

Figure CN114549905B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image classification method based on an improved online knowledge distillation algorithm, and belongs to the technical field of model compression in deep learning. Background Art
[0002] Deep learning methods have made great progress in the field of computer vision. Researchers usually use large neural networks to improve recognition accuracy. However, large networks have many problems: the number of large network parameters is huge, which makes it impossible for some mobile devices or embedded devices to store such network models; the amount of calculation is huge, which makes it take a lot of time to perform an inference and cannot meet real-time requirements; and the huge energy consumption is the most critical problem on embedded devices, which must consider the need for battery life.
[0003] Therefore, model compression technology came into being. Today, model compression technology mainly includes model pruning, model quantization, knowledge distillation, lightweight network design and vector decomposition. Knowledge distillation is essentially a method based on transfer learning, which aims to distill the knowledge of the teacher network to the student network, and use the student network instead of the teacher network to perform reasoning during deployment, so as to achieve the purpose of model compression.
[0004] In recent years, the research on knowledge distillation methods has emerged in an endless stream. According to the distillation method, it can be divided into offline knowledge distillation and online knowledge distillation. Offline knowledge distillation is a two-stage distillation method of the teacher-student structure. The two-stage means that a large teacher network is trained first and then the knowledge of the teacher network is transferred to the student network. The training process of the teacher network and the student network is carried out separately, so a two-stage training process is required; while online knowledge distillation is a one-stage distillation method of the student-student structure, which means that multiple student networks learn from each other and are optimized simultaneously in one training stage. Online knowledge distillation is not like offline knowledge distillation, where the training of the teacher network and the student network is carried out separately. Therefore, from the perspective of implementation, online knowledge distillation is more convenient than offline knowledge distillation. From the perspective of training cost, online knowledge distillation saves the time and computing resource cost of training a teacher network alone. Therefore, the research on online knowledge distillation has certain significance.
[0005] However, the existing online knowledge distillation methods lack the stable and reasonable supervision information that the teacher network of offline knowledge distillation can provide. Therefore, the supervision information generated by several student networks of the online knowledge distillation method during the early training is not accurate enough, which will lead to the performance of the distilled network being damaged, resulting in the inability to complete computer vision tasks well, such as low classification accuracy when performing image classification. Summary of the invention
[0006] In order to improve the accuracy of image classification using a network obtained by an online knowledge distillation method, the present invention provides an image classification method based on an improved online knowledge distillation algorithm, the method comprising:
[0007] Step 1: Determine the student network set in the online knowledge distillation algorithm, and set the wrong question set to save the integrated output features of each student network during the training process. Train each student network through the training pictures in the public image classification dataset to update the network parameters and obtain the trained student network set. When updating the parameters of each student network, judge whether the integrated output features recorded in the wrong question set are correct by the real labels of the training pictures. If the recorded features are correct, update the network parameters so that the output features z of each network in the network set are i The distribution of z is close to the characteristic distribution recorded in the wrong question set. If the recorded characteristics are wrong, let z i The distribution is far away from the characteristic distribution recorded in the wrong question set;
[0008] Step 2: Apply different random transformations to the image to be classified to obtain the transformed image and input it into the trained student networks to obtain the output features of each student network. The output features of all student networks are averaged to obtain the integrated output features corresponding to the image to be classified, and the image to be classified is classified according to the integrated output features of the image to be classified.
[0009] Optionally, suppose the public image classification dataset contains N training images with a total of C categories, then the initialization wrong question set is The zero matrix of ; is a set of real numbers of dimension N×C;
[0010] In Step 1, each student network is trained to update the network parameters, including:
[0011] Preprocess the training images in the public image classification dataset;
[0012] Use the reflection mechanism to train each student network in the student network set, set the network training batch to 128, the learning rate to 0.1, the momentum to 0.9, the weight decay regularization coefficient to 0.0005, the total number of training rounds to 300, the learning rate decayed by 10 times at the 150th and 225th training rounds, the distillation temperature coefficient was set to 3.0, and the historical information retention rate γ of the wrong question set was 0.2;
[0013] When each student network propagates forward to update the network parameters, the output feature z of the i-th student network is obtained i , the output of all student networks is averaged to obtain the integrated output feature z e , in the forward direction, the integrated output is updated to the wrong question set, and the expression is:
[0014] NB←γNB+(1-γ)z e
[0015] When each student network back-propagates to update the network parameters, the real labels of the training images can be used to determine whether the features recorded in the wrong question set are correct. If the features recorded are correct, the output feature z of each student network in the student network set is i The distribution of z is close to the characteristic distribution recorded in the wrong question set. If the recorded characteristics are wrong, let z i The distribution of is far away from the characteristic distribution recorded in the wrong question set; that is, the loss function of each student network is set to:
[0016]
[0017] Among them, L HKD (NB, z i ) is the distillation loss function, L CE is the cross entropy loss function.
[0018] Optionally, when using the reflection mechanism to train each student network in the student network set, the initial reflection coefficient μ is set to μ 0 is 0.005, and the verification monitoring window size w is 10.
[0019] Optionally, after each student network completes a round of training and verification, the verification accuracy of each student network in the current round is recorded. When the verification accuracy of each student network exceeds the first threshold for w consecutive rounds, the reflection coefficient is set to a first predetermined value; when the verification accuracy of each student network exceeds the second threshold for ω consecutive rounds, the reflection coefficient is set to a second predetermined value.
[0020] Optionally, the first threshold is 75%, and the corresponding first predetermined value is 0.01, that is, μ=2μ 0 is 0.01.
[0021] Optionally, the second threshold is 93.5%, and the corresponding second predetermined value is 0.02, that is, μ=4μ 0 is 0.02.
[0022] Optionally, when training each student network by using training pictures in a public image classification data set to update network parameters, different preprocessing is used for different student networks to transform the same training pictures to obtain the input of each student network.
[0023] Optionally, the preprocessing includes random horizontal flipping, random cropping, padding the boundary part with zeros, adjusting the resolution to 32*32, and normalizing it.
[0024] The beneficial effects of the present invention are:
[0025] By using the wrong question set module to save the historical information of the network set integrated output, and using it as the supervision information for the student network optimization training, a reflection mechanism is proposed to make the output distribution of the student network move away from the wrong historical information and close to the correct historical information, forcing the student network to learn better features, thereby obtaining more accurate classification results when classifying images. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0027] Figure 1 Method structure diagram for online knowledge distillation for reflective collaborative learning.
[0028] Figure 2 This is a comparison chart of the accuracy results of the resnet32 network DML, ONE, OKDDip, KDCL and the method of the present invention on the CIFAR100 dataset. DETAILED DESCRIPTION
[0029] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0030] Embodiment 1:
[0031] This embodiment provides an image classification method based on an improved online knowledge distillation algorithm. Figure 1 , the method comprising:
[0032] Step 1: Determine the student network set in the online knowledge distillation algorithm, and set the wrong question set to save the integrated output features of each student network during the training process. Train each student network through the training pictures in the public image classification dataset to update the network parameters and obtain the trained student network set. When updating the parameters of each student network, judge whether the integrated output features recorded in the wrong question set are correct by the real labels of the training pictures. If the recorded features are correct, update the network parameters so that the output features z of each network in the network set are i The distribution of z is close to the characteristic distribution recorded in the wrong question set. If the recorded characteristics are wrong, let z i The distribution is far away from the characteristic distribution recorded in the wrong question set;
[0033] Step 2: Apply different random transformations to the image to be classified to obtain the transformed image and input it into the trained student networks to obtain the output features of each student network. The output features of all student networks are averaged to obtain the integrated output features corresponding to the image to be classified, and the image to be classified is classified according to the integrated output features of the image to be classified.
[0034] The student network set contains several student networks. The student network can be any neural network, such as ResNet, VGG, MobileNet, DenseNet, ShuffleNet and other public networks.
[0035] Embodiment 2
[0036] This embodiment provides an image classification method based on an improved online knowledge distillation algorithm. The method makes improvements on the existing online knowledge distillation algorithm for image classification tasks, and proposes a reflective collaborative learning online knowledge distillation algorithm based on a wrong question set mechanism, which is then applied to image classification tasks, thereby improving image classification accuracy.
[0037] The method comprises:
[0038] Step A.1: Set the model structure of the student network in the network set.
[0039] The step A.1 comprises:
[0040] (1) The model structure in the student network set should be selected according to the deployment scenario requirements. If you want speed, you should choose a lightweight network with small capacity, such as MobileNet, ShuffleNet, etc. If you want accuracy, you can choose a network with moderate or large capacity, such as DenseNet, VGG;
[0041] The structures of the neural networks in the student network set all adopt known public network structures. For example, the VGG (Simonyan K, Zisserman A. Very deep convolutional networks for large-scale image recognition [J]. arXiv preprint arXiv: 1409.1556, 2014.) network was proposed in 2014. It proved that the depth of the network can improve the performance of the network to a certain extent, and it uses continuously stacked small convolution kernels instead of a large convolution kernel to ensure that the nonlinearity of the network is increased under the condition of having the same receptive field. ResNet (He K, Zhang X, Ren S, et al. Deep residual learning for image recognition [C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770-778.) was proposed in 2015. It mainly refers to the VGG network structure and is modified on its basis. By adding residual units, the degradation problem of deep networks is alleviated. To a certain extent, the problem of gradient disappearance during training of deep networks is also solved by residual units. MobileNet (Howard AG, Zhu M, Chen B, et al. Mobilenets: Efficient convolutional neural networks for mobile vision applications [J]. arXiv preprint arXiv: 1704.04861, 2017.) is a lightweight deep network proposed in 2017. It mainly proposes to use deep separable convolution to reduce the number of parameters and calculations of standard convolution without reducing the recognition accuracy too much, thus alleviating the pressure of low computing power of mobile devices. And the ReLU6 activation function is used instead of ReLU to make the network more friendly to low-bit quantization.
[0042] (2) The number of networks in the set should be selected based on the graphics card storage space of the model training machine. If the graphics card storage space is sufficient, more networks can be selected for training, but the requirement of at least two networks participating in the training should be met.
[0043] Step B.1: If the dataset contains N training images with a total of C categories, then the set of wrong questions is initialized as The zero matrix of :
[0044] The step B.1 comprises:
[0045] (1) If the dataset contains N training images with a total of C categories, then the set of wrong questions to be initialized is The zero matrix of .
[0046] Step C.1. Define the reflection coefficient adjustment strategy.
[0047] The step C.1 comprises:
[0048] (1) Define the reflection coefficient adjustment strategy, and set the initial reflection coefficient μ = μ 0 is 0.005, and the verification monitoring window size w is 10. When the model completes a round of training and verification, the verification accuracy of the model in the current round is recorded. When the verification accuracy of the model exceeds 75% for w consecutive rounds, set μ = 2μ 0 As the model is further trained, when the verification accuracy of the model exceeds 93.5% for w consecutive rounds, μ = 4μ 0 is 0.02.
[0049] Step D.1, read the image data of the dataset, perform data enhancement on the image, adjust the resolution to 32*32 and normalize it.
[0050] The step D.1 comprises:
[0051] (1) Read the image data of the dataset, randomly flip the image horizontally, randomly crop it, fill the border with zeros, adjust the resolution to 32*32, and normalize it;
[0052] (2) The input of each network in the network set should use different random transformations. That is, if there are three models in the network set, multiple transformations are performed on the same input image to obtain the inputs of the three models respectively, and the three inputs are different. Finally, the resolution is adjusted to 32*32 so that its size can be input to the model.
[0053] Step E.1: Use the reflection mechanism to train the networks in the network set, set the batch size of network training to 128, and update the network ensemble output to the wrong question set according to the specified update strategy during network forward propagation.
[0054] The step E.1 comprises:
[0055] (1) Set the network training batch size to 128, the learning rate to 0.1, the momentum to 0.9, the weight decay regularization coefficient to 0.0005, the total number of training rounds to 300, the learning rate to decay by 10 times at the 150th and 225th training rounds, the distillation temperature coefficient to 3.0, and the historical information retention rate γ of the wrong question set to 0.2;
[0056] (2) When the network is propagated forward, the output feature z of the i-th network is obtained i , the output of all networks is averaged to get the integrated output feature z e , in the forward direction, the integrated output needs to be updated to the wrong question set, and the expression is:
[0057] NB←γNB+(1-γ)z e ;
[0058] (3) When back-propagating to update network parameters, the real labels of the training images can be used to determine whether the features recorded in the wrong question set are correct. If the features recorded are correct, the output feature z of each network in the network set is set to i The distribution of z is close to the characteristic distribution recorded in the wrong question set. If the recorded characteristics are wrong, let z i The distribution of is far away from the characteristic distribution recorded in the wrong question set, thus achieving the purpose of reflection. The loss function is expressed as:
[0059]
[0060] Among them, L HKD (NB, z i ) is the distillation loss function proposed by (Hinton G, Vinyals O, Dean J. Distilling the knowledge in a neural network [EB / OL]. https: / / arxiv.org / abs / 1503.02531), L CE is the cross entropy loss function.
[0061] Step F.1: If the verification accuracy of the network reaches a certain set condition, adjust the reflection coefficient during training.
[0062] The step F.1 comprises:
[0063] (1) If the integrated accuracy of the network set on the validation set meets the set conditions, the reflection coefficient μ is adjusted according to the set strategy, and a new round of training is carried out using the new reflection coefficient.
[0064] Step G.1: Complete the training of the network set and test and deploy the best performing model in the network set.
[0065] The step G.1 comprises:
[0066] (1) When deploying a model, if accuracy is a priority, all networks in the network set can be deployed, and the average value of the outputs of all networks in the network set can be taken as the final integrated result to obtain a more accurate recognition effect. If inference speed is a priority, the model with the highest verification accuracy in the network set can be selected for deployment, which can achieve both inference efficiency and accuracy.
[0067] Embodiment 3
[0068] This embodiment provides an image classification method based on an improved online knowledge distillation algorithm, and takes the application of ResNet110 network to classify images on the CIFAR100 dataset as an example. The CIFAR100 dataset is a public dataset in the field of image classification. The method includes:
[0069] A.1. Set the model structure of each network in the network set.
[0070] The step A.1 comprises:
[0071] (1) The selection of the model structure in the network set should be based on the deployment scenario requirements. If speed is required, a lightweight network with a small capacity should be selected. If accuracy is required, a network with a moderate or large capacity can be selected. In this embodiment, ResNet110 is selected as the student network in the student network set.
[0072] (2) The number of networks in the set should be selected based on the graphics card storage space of the model training machine. If the graphics card storage space is sufficient, more networks can be selected for training, but the requirement of at least two networks participating in the training should be met. In this embodiment, three networks are selected, that is, all three networks are ResNet110.
[0073] It should be noted that when selecting a specific type of student network, the student network can be the same network or different networks. For example, the three networks can be selected from VGG, ResNet and MobileNet respectively, or ResNet110 can be selected for all networks as in this embodiment.
[0074] B.1. If the dataset contains N training images with a total of C categories, then the set of wrong questions to be initialized is A zero matrix of N rows and C columns is initialized.
[0075] C.1. Define the reflection coefficient adjustment strategy.
[0076] The step C.1 comprises:
[0077] (1) Define the reflection coefficient adjustment strategy, and set the initial reflection coefficient μ = μ 0is 0.005, and the verification monitoring window size w is 10. When the model completes a round of training and verification, the verification accuracy of the model in the current round is recorded. When the verification accuracy of the model exceeds 75% for w consecutive rounds, set μ = 2μ 0 As the model is further trained, when the verification accuracy of the model exceeds 93.5% for w consecutive rounds, μ = 4μ 0 is 0.02.
[0078] D.1. Read the image data of the dataset, perform data augmentation on the image, adjust the resolution to 32*32 and normalize it.
[0079] The step D.1 comprises:
[0080] (1) Read the image data of the dataset and preprocess the image by randomly flipping it horizontally, randomly cropping it, filling the borders with zeros, adjusting the resolution to 32*32, and normalizing it;
[0081] (2) The input of each network in the network set should adopt different random preprocessing. That is, if there are three student networks in the network set, multiple transformations are performed on the same input image to obtain the input of each of the three models, and the three inputs are different. Finally, the resolution is adjusted to 32*32 so that its size can be input to the model.
[0082] E.1. Use the reflection mechanism to train the networks in the network set, set the batch size of network training to 128, and update the network ensemble output to the wrong question set according to the specified update strategy during network forward propagation.
[0083] The step E.1 comprises:
[0084] (1) Set the network training batch size to 128, the learning rate to 0.1, the momentum to 0.9, the weight decay regularization coefficient to 0.0005, the total number of training rounds to 300, the learning rate to decay by 10 times at the 150th and 225th training rounds, the distillation temperature coefficient to 3.0, and the historical information retention rate γ of the wrong question set to 0.2;
[0085] (2) When the network is propagated forward, the output feature z of the i-th network is obtained i , the output of all networks is averaged to get the integrated output feature z e , in the forward direction, the integrated output needs to be updated to the wrong question set, and the expression is:
[0086] NB←γNB+(1-γ)z e
[0087] (3) When back-propagating to update network parameters, the real labels of the training images can be used to determine whether the features recorded in the wrong question set are correct. If the features recorded are correct, the output feature z of each network in the network set is set to i The distribution of z is close to the characteristic distribution recorded in the wrong question set. If the recorded characteristics are wrong, let z i The distribution of is far away from the characteristic distribution recorded in the wrong question set, thus achieving the purpose of reflection. The loss function is expressed as:
[0088]
[0089] Among them, L HKD (NB, zi) is the distillation loss function proposed by (Hinton G, Vinyals O, Dean J. Distilling the knowledge in a neural network [EB / OL]. https: / / arxiv.org / abs / 1503.02531), L CE is the cross entropy loss function.
[0090] F.1. If the verification accuracy of the network reaches a certain set condition, adjust the reflection coefficient during training.
[0091] The step F.1 comprises:
[0092] (1) If the integrated accuracy of the network set on the validation set meets the set conditions, the reflection coefficient μ is adjusted according to the set strategy, and a new round of training is carried out using the new reflection coefficient.
[0093] G.1. Complete the training of the network set and test and deploy the best performing model in the network set.
[0094] The step G.1 comprises:
[0095] (1) When deploying a model, if accuracy is a priority, all networks in the network set can be deployed, and the average value of the outputs of all networks in the network set can be taken as the final integrated result to obtain a more accurate recognition effect. If inference speed is a priority, the model with the highest verification accuracy in the network set can be selected for deployment, which can achieve both inference efficiency and accuracy.
[0096] like Figure 2 As described above, this embodiment compares the image classification accuracy of the model obtained by compressing the image classification model using different knowledge distillation methods to classify all images of the CIFAR100 dataset, where:
[0097] Introduction to DML “Zhang Y,Xiang T,Hospedales TM,et al.Deep mutual learning[C] / / Proceedings of the IEEE Conference on Computer Vision andPattern Recognition.2018:4320-4328.”;
[0098] “Lan X,Zhu X,Gong S.Knowledge distillation by on-the-flynative ensemble[C] / / Proceedings of the 32nd International Conference onNeural Information Processing Systems.2018:7528-7538.”
[0099] “Chen D,Mei JP,Wang C,et al.Online Knowledge Distillation with Diverse Peers[C] / / Proceedings of the AAAI Conference onArtificial Intelligence.2020,34(04):3430-3437.”;
[0100] “Online Knowledge Distillationvia Collaborative Learning[C] / / Proceedings of the IEEE / CVF Conference onComputer Vision and Pattern Recognition.2020:11020-11029,” said KDCL.
[0101] The above four methods and the method of this application all use three network models for online knowledge distillation for image classification. Among the above four methods, the DML method adopts the idea of deep mutual learning, and each student network is optimized in turn in each training batch. When training one of the networks, the outputs of other networks are used as supervision information to guide the current network to be optimized for training. The ONE method uses a gate module to weighted average the outputs of multiple network branches to construct integrated supervision information, and then uses the integrated supervision information to optimize each network branch in turn. The KDDip method proposes a two-level distillation method. The first level of distillation is to use the attention mechanism to weight the outputs of several network branches to construct supervision information and supervise the optimization of several network branches. The second level of distillation is to use the average output of several network branches as supervision information to guide a student leader network to optimize. The KDCL method integrates the output results of all student networks, uses this result as supervision information, and guides each student network to optimize in turn. It can be seen that these four existing methods only focus on the current supervision information and ignore the supervision information that can be provided by historical records, so their classification accuracy cannot be further improved; Figure 2 The image classification accuracy of these four methods and the method of this application is given by Figure 2 It can be seen that the image classification accuracy obtained by the method of this application is 74.21%, which is higher than the four existing methods listed. However, it is very difficult to improve the accuracy of image classification by using networks distilled using different knowledge distillation methods. It can be seen that in previous studies, the highest value was 73.79%. This application uses the wrong question set module to save the historical information of the network set integrated output, and uses it as the supervision information for the student network optimization training. A reflection mechanism is proposed to make the output distribution of the student network away from the wrong historical information and close to the correct historical information, forcing the student network to learn better features, thereby obtaining more accurate classification results when classifying images, so that the image classification accuracy is improved to 74.21%.
[0102] Some steps in the embodiments of the present invention may be implemented using software, and the corresponding software program may be stored in a readable storage medium, such as a CD or a hard disk.
[0103] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An image classification method based on an improved online knowledge distillation algorithm, It is characterized in that The method comprises: Step 1: Determine the student network set in the online knowledge distillation algorithm, and set the wrong question set to save the integrated output features of each student network during the training process. Train each student network through the training pictures in the public image classification dataset to update the network parameters and obtain the trained student network set. When updating the parameters of each student network, judge whether the integrated output features recorded in the wrong question set are correct by the real labels of the training pictures. If the recorded features are correct, update the network parameters so that the output features z of each network in the network set are i The distribution of z is close to the characteristic distribution recorded in the wrong question set. If the recorded characteristics are wrong, let z i The distribution is far away from the characteristic distribution recorded in the wrong question set; Step 2: Apply different random transformations to the image to be classified to obtain the transformed image and input it into each trained student network to obtain the output features of each student network. The output features of all student networks are averaged to obtain the integrated output features corresponding to the image to be classified, and the image to be classified is classified according to the integrated output features of the image to be classified; Assume that the public image classification dataset contains N training images with a total of C categories, then the initialization set of wrong questions is The zero matrix of ; is a set of real numbers of dimension N×C; In Step 1, each student network is trained to update the network parameters, including: Preprocess the training images in the public image classification dataset; Use the reflection mechanism to train each student network in the student network set, set the network training batch to 128, the learning rate to 0.1, the momentum to 0.9, the weight decay regularization coefficient to 0.0005, the total number of training rounds to 300, the learning rate decayed by 10 times at the 150th and 225th training rounds, the distillation temperature coefficient was set to 3.0, and the historical information retention rate γ of the wrong question set was 0.2; When each student network propagates forward to update the network parameters, the output feature z of the i-th student network is obtained i , the output of all student networks is averaged to obtain the integrated output feature z e , in the forward direction, the integrated output is updated to the wrong question set, and the expression is: NB←γNB+(1-γ)z e When each student network back-propagates to update the network parameters, the real labels of the training images can be used to determine whether the features recorded in the wrong question set are correct. If the features recorded are correct, the output feature z of each student network in the student network set is i The distribution of z is close to the characteristic distribution recorded in the wrong question set. If the recorded characteristics are wrong, let z i The distribution of is far away from the characteristic distribution recorded in the wrong question set; that is, the loss function of each student network is set to: Among them, L HKD (NB, z i ) is the distillation loss function, L CE is the cross entropy loss function.
2. The method according to claim 1, It is characterized in that When the reflection mechanism is used to train each student network in the student network set, the initial reflection coefficient μ is set to μ 0 is 0.005, and the verification monitoring window size w is 10.
3. The method according to claim 2, It is characterized in that After each student network completes a round of training and verification, the verification accuracy of each student network in the current round is recorded. When the verification accuracy of each student network exceeds the first threshold for w consecutive rounds, the reflection coefficient is set to the first predetermined value; when the verification accuracy of each student network exceeds the second threshold for w consecutive rounds, the reflection coefficient is set to the second predetermined value.
4. The method according to claim 3, It is characterized in that The first threshold is 75%, and the corresponding first predetermined value is 0.01, that is, μ=2μ 0 is 0.
01.
5. The method according to claim 4, It is characterized in that The second threshold is 93.5%, and the corresponding second predetermined value is 0.02, that is, μ=4μ 0 is 0.
02.
6. The method according to claim 5, It is characterized in that When training each student network through the training pictures in the public image classification data set to update the network parameters, different preprocessing is used for different student networks to transform the same training pictures to obtain the input of each student network.
7. The method according to claim 6, It is characterized in that The preprocessing includes random horizontal flipping, random cropping, filling the boundary part with zeros, adjusting the resolution to 32*32, and performing normalization processing on it.
Citation Information
Patent Citations
Key point positioning model training method, positioning method, device and equipment
CN109740567A
Image recognition method and device, electronic equipment and storage medium
CN112001364A