A method for aligning class accuracy of different structural classification neural networks
Patent Information
- Application Number
- CN202310298351.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-24
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-03-24
AI Technical Summary
[0004]本发明的目的旨在针对两个不同的神经网络在分类任务下各类别精度难以对齐等问题,提供一种对齐不同结构分类神经网络类别精度的方法
[0028]第一,本发明指出工业界中部署新模型时存在各类别精度与原模型精度差异较大的问题,并设计有效的解决方法。
Smart Images

Figure CN116306872B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to computer vision, and more particularly to a method for aligning the accuracy of different structural classification neural network categories. Background Technology
[0002] In recent years, neural networks have achieved great success in the field of computer vision. Image classification, as a typical downstream task in computer vision, plays a significant role in many scenarios such as autonomous driving, medical image analysis, and behavior recognition. In these different tasks, model accuracy is one of the important indicators for measuring task completion. However, the academic community currently generally focuses on the average accuracy of the model, with little attention paid to the accuracy of each category individually, even though the accuracy of each category is more important in practical industrial deployments. Therefore, techniques for optimizing the accuracy of each category are urgently needed.
[0003] In industry, updating classification models typically involves fine-tuning the same model on a new dataset. This approach has the advantage of inheriting the weights of the previous model, preserving as much of its feature information as possible, which is naturally advantageous for maintaining accuracy across classes. However, in practice, model iteration also occurs alongside dataset updates. In model iteration, the new model parameters are usually randomly initialized, meaning they won't contain the feature information of the original model. While a naive training method might achieve comparable average classification accuracy to the original model, it struggles to match the individual accuracy across classes, making it unsuitable for real-world applications. Summary of the Invention
[0004] The purpose of this invention is to address the problem of difficulty in aligning the accuracy of different categories in classification tasks between two different neural networks, and to provide a method for aligning the category accuracy of classification neural networks with different structures. This method aligns the accuracy of different neural network models for the same dataset through a simple training approach.
[0005] The method for aligning the accuracy of different structural classification neural network categories according to the present invention includes the following steps:
[0006] 1) Pre-training: Pre-train the new structure model using an earlier version of the dataset to optimize the model weights;
[0007] 2) Knowledge distillation: Using neural network knowledge distillation techniques to perform knowledge transfer between new and old models;
[0008] 3) Fine-tuning the fully connected layer: Using the weight freezing technique, the shallow parameters of the neural network are frozen, and the categories with poor accuracy in the fully connected layer are fine-tuned.
[0009] In step 1), the pre-training involves sequentially feeding early version datasets into the new neural network model in version iteration order to train the new model. The original model is updated through iterative training on multiple version datasets. The early version datasets for the new model can use all versions of the datasets used in the iterative training of the original model, arranged in version order, as the training dataset for pre-training the new model. Alternatively, a subset of the datasets (e.g., an odd number of versions) can be used for pre-training to reduce training costs.
[0010] If the new model is pre-trained using all versions of the dataset, the sub-steps are as follows:
[0011] (1) Use the initial version of the dataset and pre-train the new model with a high learning rate.
[0012] (2) Use subsequent versions of the dataset and pre-train the new model with a low learning rate until iterates to the latest version of the dataset.
[0013] In step 2), the knowledge distillation involves using the output of the original model as the learning target of the new model and using a specific loss function to optimize the accuracy of the new model in each category.
[0014] The loss function used is as follows:
[0015]
[0016] In this loss function, α, β, and γ are hyperparameters used to balance the magnitude of the various losses.
[0017] In the formula The label-supervised loss aims to optimize the distance relationship between the output of the new model and the labels in the training set. As shown in the following equation, for multi-label image classification tasks, the BCE Loss can be used, where N is the number of samples, and y... i p is the category to which the i-th sample belongs. i This is the predicted value for the i-th sample:
[0018]
[0019] Knowledge distillation loss function Defined as the inter-class loss between the outputs of the original model and the new model, this knowledge distillation loss function can maintain the linear correlation between the output vectors of each teacher and student model; compared with the ordinary knowledge distillation loss function, its conditions are more relaxed; where B is the batch size, d is the distance metric, which can be the Pearson correlation coefficient, Y is the output of the model, t represents the original model, and s represents the new model.
[0020]
[0021] Knowledge distillation loss function Defined as the intra-class loss between the outputs of the original model and the new model, it aims to transfer the output distribution of the original model to the new model for different samples of the same class, where C is the number of classes;
[0022]
[0023] In step 3), the fine-tuning of the fully connected layer, for a new deep neural network model containing multiple layers, requires freezing all weights in the new model except for the last fully connected layer, i.e., preventing gradient updates of parameters in non-last layers, and updating only the parameters of the deepest fully connected layer of the new model during backpropagation. Its sub-steps are:
[0024] (1) For the pre-trained new model, freeze all weights except for the last fully connected layer, and fine-tune only the parameters of the deepest fully connected layer.
[0025] (2) Update the parameters of the deepest fully connected layer of the model using the latest version of the dataset to obtain a new model.
[0026] The method of this invention mainly includes a method for aligning the accuracy of different classification neural networks, which mainly includes three steps: pre-training, knowledge distillation, and fine-tuning of fully connected layers. It can effectively address the problem of large differences between the accuracy of each category and the original model when deploying new models in industry.
[0027] Compared with the prior art, the technical effects and advantages of the present invention are as follows:
[0028] First, this invention points out the problem that there are significant differences in accuracy between different categories and the original model when deploying new models in the industry, and designs an effective solution.
[0029] Second, this invention creatively reduces the accuracy differences between the new model and the original model across different categories by using three steps: iterative version data pre-training, knowledge distillation, and fine-tuning of the fully connected layer.
[0030] Third, the method proposed in this invention achieves good experimental results on multi-label classification datasets. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention.
[0032] Figure 2 This is a schematic diagram of the pre-training process of the new model in an embodiment of the present invention.
[0033] Figure 3 This is a schematic diagram of the knowledge distillation process according to an embodiment of the present invention.
[0034] Figure 4 This is a schematic diagram of the process of fine-tuning the fully connected layer in an embodiment of the present invention. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this invention clearer, the following embodiments will be used to further illustrate this invention in conjunction with the accompanying drawings.
[0036] like Figure 1 As shown, the embodiments of the present invention include the following steps:
[0037] 1) Pre-training: Pre-train the new structure model using an earlier version of the dataset to optimize the model weights;
[0038] 2) Knowledge distillation: Using neural network knowledge distillation techniques to perform knowledge transfer between new and old models;
[0039] 3) Fine-tuning the fully connected layer: Using the weight freezing technique, the shallow parameters of the neural network are frozen, and the categories with poor accuracy in the fully connected layer are fine-tuned.
[0040] 1) Pre-training
[0041] like Figure 2 As shown, in step 1), the pre-training involves sequentially feeding early version datasets into the new neural network model in version iteration order to train the new model. The original model is updated through iterative training on multiple version datasets. The early version datasets for the new model can use all versions of the datasets used in the iterative training of the original model, arranged in version order, as the training dataset for the new model to pre-train it. Alternatively, a subset of the datasets (e.g., using only odd-numbered or even-numbered versions) can be used for pre-training to reduce training costs.
[0042] If the new model is pre-trained using all versions of the dataset, the sub-steps are as follows:
[0043] (1) Use the initial version of the dataset and pre-train the new model with a high learning rate.
[0044] (2) Use subsequent versions of the dataset and pre-train the new model with a low learning rate until iterates to the latest version of the dataset.
[0045] 2) Knowledge distillation
[0046] like Figure 3 As shown, in step 2), the knowledge distillation is to use the output of the original model as the learning target of the new model and use a specific loss function to optimize the accuracy of the new model in each category.
[0047] The loss function used is as follows:
[0048]
[0049] In this loss function, α, β, and γ are hyperparameters used to balance the magnitude of the various losses.
[0050] In the formula The label-supervised loss aims to optimize the distance relationship between the output of the new model and the labels in the training set. As shown in the following equation, for multi-label image classification tasks, the BCE Loss can be used, where N is the number of samples, and y... i p is the category to which the i-th sample belongs. i This is the predicted value for the i-th sample:
[0051]
[0052] Knowledge distillation loss function Defined as the inter-class loss between the outputs of the original model and the new model, this knowledge distillation loss function can maintain the linear correlation between the output vectors of each teacher and student model; compared with the ordinary knowledge distillation loss function, its conditions are more relaxed; where B is the batch size, d is the distance metric, which can be the Pearson correlation coefficient, Y is the output of the model, t represents the original model, and s represents the new model.
[0053]
[0054] Knowledge distillation loss function Defined as the intra-class loss between the outputs of the original model and the new model, it aims to transfer the output distribution of the original model to the new model for different samples of the same class, where C is the number of classes;
[0055]
[0056] 3) Fine-tune the fully connected layer
[0057] like Figure 4 As shown, in step 3), the fine-tuning of the fully connected layer, for a new deep neural network model containing multiple layers, requires freezing all weights in the new model except for the last fully connected layer, that is, preventing gradient updates of parameters other than the last layer, and only updating the parameters of the deepest fully connected layer of the new model during backpropagation.
[0058] Its sub-steps are as follows:
[0059] (1) For the pre-trained new model, freeze all weights except for the last fully connected layer, and fine-tune only the parameters of the deepest fully connected layer.
[0060] (2) Update the parameters of the deepest fully connected layer of the model using the latest version of the dataset to obtain a new model.
[0061] The method of this invention mainly includes a method for aligning the accuracy of different classification neural networks, which mainly includes three steps: pre-training, knowledge distillation, and fine-tuning of fully connected layers. It can effectively address the problem of large differences between the accuracy of each category and the original model when deploying new models in industry.
[0062] The effects of the present invention will be further illustrated by simulation experiments below.
[0063] (1) Simulation conditions
[0064] This invention was developed on an NVIDIA A100 GPU environment, using the PyTorch deep learning framework. The primary language used in this invention is Python.
[0065] (2) Simulation content
[0066] Simulations were performed on a multi-label image classification dataset primarily used for scene recognition tasks. This dataset consists of 224x224 color images across 37 classes, totaling approximately 1.4 million images. Each class contains a maximum of 210,000 images and a minimum of 7,700 images. 1.3 million images are used as the training set, and 100,000 images as the test set. Each image can belong to multiple classes simultaneously, with a maximum of 8 classes and a minimum of not belonging to any of the 37 classes.
[0067] Table 1: Comparison of Simulation Results for Multi-Label Image Classification Datasets
[0068] BCE Loss 30.71% 6.23% Step 1 of the present invention 11.35% 3.48% Steps 1) and 2) of the present invention 8.87% 2.09% Steps 1), 2), and 3) of this invention 5.37% 1.49%
[0069] As shown in Table 1, the present invention was simulated on the multi-label image classification dataset mentioned above, where the original model was MobileNetV3-Large and the new model was MobileNetV3-Small. It can be seen that the three steps proposed in the present invention can effectively reduce the difference in accuracy of each category between the new model and the original model.
Claims
1. A method for aligning the accuracy of categories in classification neural networks with different structures, characterized in that... Includes the following steps: 1) Pre-training: The model with the new structure is pre-trained using an early version dataset to optimize the model weights. The early version dataset is a multi-label image classification dataset. The pre-training involves feeding the early version dataset into the new neural network model in the order of version iteration to train the new model. The original model version is updated through iterative training on multiple version datasets. The early version dataset of the new model uses all versions of the dataset used in the iterative training of the original model, i.e. the full dataset, and uses them as the training dataset for the new model in version order to pre-train the new model, or uses a portion of the dataset to pre-train the new model to reduce training costs. 2) Knowledge distillation: Using neural network knowledge distillation techniques to perform knowledge transfer between new and old models; Knowledge distillation loss Defined as the inter-class loss of the model output, this knowledge distillation loss maintains the linear correlation between the output vectors of each teacher and student model; where B is the batch size, d is the distance metric, which is the Pearson correlation coefficient, Y is the model output, t represents the old version of the online model, and s represents the new model; Knowledge distillation loss Defined as the intra-class loss of the model output, it aims to pass the ranking of the old model's outputs among different samples of the same class to the new model, where C is the number of classes; 3) Fine-tuning the fully connected layers: For the new model trained in step 2), the shallow parameters of the neural network are frozen using the weight freezing technique. That is, all weights in the new model except for the last fully connected layer are frozen to prevent gradient updates of the parameters of non-last fully connected layers. Using the latest version of the dataset, the parameters corresponding to the classes with poor accuracy in the deepest fully connected layer are fine-tuned and updated only to obtain the new model.
2. The method for aligning the accuracy of classification neural network categories with different structures as described in claim 1, characterized in that... In step 1), the new model is pre-trained using the entire dataset. The sub-steps are as follows: (1) Use the initial version of the dataset and pre-train the new model with a high learning rate; (2) Use subsequent versions of the dataset and pre-train the new model with a low learning rate until iterates to the latest version of the dataset.
3. The method for aligning the accuracy of classification neural network categories with different structures as described in claim 1, characterized in that... In step 1), the partial dataset may be either an odd-numbered version of the dataset or an even-numbered version of the dataset.
4. The method for aligning the accuracy of classification neural network categories with different structures as described in claim 1, characterized in that... In step 2), the knowledge distillation involves using the output of the original model as the learning target of the new model and using a loss function to optimize the accuracy of the new model in each category. The loss function used is as follows: The loss These are hyperparameters used to balance the magnitude of various losses; In the formula For supervised loss, which aims to optimize the distance relationship between the output of the new model and the labels on the training set, BCELoss is used, where the loss... It is the total number of samples. It is the first The category to which each sample belongs. It is the first Predicted values for each sample: 。
Citation Information
Patent Citations
Model distillation method and system and text retrieval method
CN114328834A