An image classification model training method and classification method based on knowledge transfer
Through the self-supervised enhancement and category enhancement methods combined with knowledge transmission, the catastrophic forgetting problem in incremental learning is solved, the semantic relationship utilization of the model between the old and new data sets is realized, and the generalization and migration performance of the model is improved.
Patent Information
- Application Number
- CN202211126235.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-09-16
AI Technical Summary
The existing technology has catastrophic forgetting problems in the incremental learning process, and the semantic relationships and key features between new and old data are not fully utilized, resulting in a degradation of model performance.
The image data set is processed by self-supervised enhancement and category enhancement methods, and the model parameters are updated in combination with cross-entropy loss, knowledge distillation loss and knowledge transmission loss. The semantic relationship between the new and old data sets is obtained by using optimal knowledge transmission, and the feature extraction network is initialized.
It alleviates the catastrophic forgetting problem in the incremental learning process, improves the generalization and transferability of the model, maintains the ability to classify old data, and at the same time improves the ability to classify new data.
Smart Images

Figure CN115471700B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, in particular to the field of image classification in the field of computer vision, and more particularly to an image classification model training method and a classification method based on knowledge transfer. Background Art
[0002] In the prior art, deep learning is widely used in the field of image processing, and deep convolutional neural networks in deep learning have achieved remarkable success in retinal blood vessel segmentation, retinal disease classification, and fundus disease image abnormality detection. However, these successes depend to a large extent on a large amount of labeled data. If there is not enough labeled data to train the network, the network's processing capabilities in retinal blood vessel segmentation, retinal disease classification, and fundus disease image abnormality detection may deteriorate. In some fields, such as the field of medical imaging, it takes a lot of time, manpower, and money to collect and label data, so it is difficult to build an effective data set in a short time. The traditional deep learning method is a batch learning method that requires all the data to train the model. If there is no large amount of labeled data to train the model, the model may not perform well. In addition, if new labeled data is received later, the entire model needs to be retrained, which will increase the cost of training and may also cause a waste of resources.
[0003] Incremental learning provides a new idea to solve the above problems. Updating the model according to the new data that keeps coming can effectively alleviate the current dilemma. Incremental learning updates the trained model and uses the new data to improve the generalization ability of the model, while retaining the existing knowledge as much as possible. In the incremental process, the old data is not available. With incremental learning, there is no need to be limited by the difficulty of building a data set in a short time, nor is there a need to retrain the entire model. However, the new data that keeps coming may be accompanied by changes in different collection devices or collection users, which may cause changes in the feature space and data distribution of the new data, which may cause data difference problems. Therefore, if the new data is directly used to fine-tune the trained model, the classification accuracy of the old data may drop sharply, causing a catastrophic forgetting problem.
[0004] In view of the above problems, Chinese Patent Application CN113066025A proposes an image dehazing method based on incremental learning, feature, and attention transfer. However, this method does not solve the problem of catastrophic forgetting in the increment well. Chinese Patent Application CN106022368A proposes an incremental trajectory anomaly detection method based on incremental kernel principal component analysis, which uses kernel principal component analysis to achieve incremental anomaly detection. Chinese Patent Application CN112990280A proposes a method that adopts class increment in image classification problems. US Patent Application US2020302230A1 proposes a method that applies class incremental learning to the field of object detection.
[0005] Although many methods for alleviating catastrophic forgetting have been proposed in the prior art, the problem of catastrophic forgetting still challenges the incremental learning process based on data increment. And the currently widely adopted solution of jointly using the distillation loss and the cross-entropy loss as the loss function to update the model parameters to alleviate catastrophic forgetting in the incremental learning process has the following disadvantages: 1. It does not explore which features extracted by the deep learning model in the incremental process are key features; 2. It does not fully utilize the semantic relationship between the new and old data in the incremental process. This makes the model after incremental learning still face catastrophic forgetting. Summary of the Invention
[0006] Therefore, the purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a method for training an image classification model based on knowledge transfer and an image classification method.
[0007] The purpose of the present invention is achieved by the following technical solutions:
[0008] According to the first aspect of the present invention, there is provided a method for training an image classification model based on knowledge transfer, which is used for incrementally training a pre-trained image classification model. Among them, the pre-trained image classification model includes a feature extraction network and a classifier. The method includes incrementally training the pre-trained image classification model with a new image data set in the following manner: S1. Perform enhancement processing on the current new image data set; S2. Initialize the image classification model with the parameters of the image classification model after the previous training, and train it with the enhanced current new image data set until convergence, where the model parameters are updated using the cross-entropy loss, the knowledge distillation loss, and the knowledge transfer loss during the training process.
[0009] Preferably, the step S1 includes: performing self-supervised enhancement processing on the current new image data set, or performing class enhancement processing on the current new image data set, or first performing self-supervised enhancement processing on the current new image data set and then performing class enhancement processing.
[0010] Preferably, the self-supervised enhancement process flips the samples in the current new image dataset at rotation angles of 90°, 180°, and 270°; the class enhancement process means randomly sampling two samples of different classes from the current new image dataset and generating a new sample for class expansion in the following way to be added to the current new image dataset:
[0011]
[0012] Among them, represents the new sample generated based on samples and ; represents the sample in class in the dataset, represents the sample in class β in the dataset, and μ represents the interpolation coefficient.
[0013] Preferably, step S2 includes: S21. Using the optimal knowledge transfer method to transfer the feature extraction network parameters of the image classification model after the previous training to the feature extraction network of the image classification model to obtain the initial image classification model for the current training; S22. Training the initial image classification model for the current training with the current new image dataset until convergence, and updating the parameters of the initial image classification model for the current training using cross-entropy loss, knowledge distillation loss, and knowledge transfer loss.
[0014] In some embodiments of the present invention, in step S22, the total loss function is calculated according to the following formula:
[0015]
[0016] Among them, represents the total loss function, represents the cross-entropy loss function, represents the knowledge distillation loss function, represents the knowledge transfer loss function, and represent hyperparameters, x represents the sample, and y represents the label of sample x.
[0017] In some embodiments of the present invention, the knowledge distillation loss is calculated using the following formula:
[0018]
[0019] Among them, represents the number of samples in the image dataset of the previous training after enhancement processing, represents the sample, represents the softmax function, Represents the classifier of the image classification model after the previous training, Represents the feature extraction network of the image classification model after the previous training, Represents the classifier of the initial image classification model for the current training, Represents the feature extraction network of the image classification model for the current training.
[0020] In some embodiments of the present invention, the knowledge transfer loss is calculated using the following formula:
[0021]
[0022] Wherein, Represents the number of samples of the new image dataset for the current training after augmentation processing, Represents a sample, Represents the softmax function, Represents the classifier of the image classification model after the previous training, Represents the feature extraction network of the image classification model after the previous training, Represents the classifier of the initial image classification model for the current training, Represents the feature extraction network of the initial image classification model for the current training.
[0023] According to the second aspect of the present invention, there is provided an image classification method, characterized in that the method includes: T1. Obtain an image to be processed; T2. Process the image to be processed using the image classification model trained by the method of the first aspect of the present invention to obtain a classification result.
[0024] Compared with the prior art, the advantages of the present invention are: (1) Adopting self-supervised learning data augmentation and class augmentation methods to make the model learning in the incremental training process more generalizable and transferable; (2) Utilizing optimal knowledge transfer to obtain the semantic relationship between the new and old datasets, realizing the transfer of the model feature space, and alleviating the catastrophic forgetting problem in the incremental training process. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The following further describes the embodiments of the present invention with reference to the accompanying drawings, wherein:
[0026] Figure 1 Is a schematic flowchart of an image classification model training method according to an embodiment of the present invention;
[0027] Figure 2 Is a schematic diagram of the comparison experiment results for 2 incremental trainings according to an embodiment of the present invention;
[0028] Figure 3 Is a schematic diagram of the comparison experiment results for 5 incremental trainings according to an embodiment of the present invention;
[0029] Figure 4 Schematic diagram of the comparative experiment results of 10 incremental trainings according to the embodiments of the present invention;
[0030] Figure 5 Schematic diagram of the confusion matrix for incremental training according to the DER method;
[0031] Figure 6 Schematic diagram of the confusion matrix for incremental training according to the MUC method;
[0032] Figure 7 Schematic diagram of the confusion matrix for incremental training according to the embodiments of the present invention. Detailed implementation manners
[0033] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0034] As mentioned in the background art, when the prior art overcomes the catastrophic forgetting problem, there are still two deficiencies: on the one hand, it does not explore which of the features extracted by the deep learning model during the incremental process are key features; on the other hand, it does not make full use of the semantic relationship between the new and old data during the incremental process. Aiming at the defects of the prior art, the present invention proposes a training method for an image classification model based on knowledge transfer, which performs incremental training on the image classification model pre-trained on the basic labeled data set using new image data. Among them, during the training process, a self-supervised enhancement method and a class enhancement method are used to enhance the new image data set, so that the model learns more generalizable and transferable representations during the incremental learning process, and then the optimal knowledge transfer is used to obtain the semantic relationship between the new and old data sets, realizing the transfer of the model feature space, thereby alleviating the catastrophic forgetting problem during the incremental learning process. Generally speaking, as Figure 1 shown, a training method for an image classification model based on knowledge transfer of the present invention includes performing incremental training on the pre-trained image classification model using a new image data set in the following manner: enhancing the current new image data set (including self-supervised enhancement and class enhancement) to achieve data representation enhancement, initializing the image classification model with the parameters of the image classification model after the previous training (preferably using the optimal knowledge transfer method for initialization), and training it with the enhanced current new image data set until convergence, where the model parameters are updated using cross-entropy loss, knowledge distillation loss, and knowledge transfer loss during the training process.
[0035] To better understand the present invention, the present invention will be described separately from five aspects: new image dataset division, data augmentation, key knowledge transfer, training process, and verification process, in combination with specific embodiments below.
[0036] I. New Image Dataset Division
[0037] The incremental learning method is different from traditional deep learning methods. It does not require training the model with all the data, but updates the model according to the continuously arriving new data. To better train the model, the new image dataset is divided into multiple sub-datasets for incremental training. One sub-dataset is used for one incremental training. By dividing the new image dataset, the change of categories in the incremental training process can be controlled, so as to calculate the sample centers between different categories and the cost between samples subsequently.
[0038] According to an embodiment of the present invention, the new image dataset is divided in the following way: the new image dataset is equally divided into 10 sub-datasets, and one sub-dataset is used for one incremental training. Among them, all sub-datasets contain the same image categories, and the proportion of image categories in each sub-dataset is the same. According to an example of the present invention, if the new image dataset includes image samples of four categories A, B, C, and D, where there are 4000 image samples of category A, 3000 image samples of category B, 2000 image samples of category C, and 1000 image samples of category D, the total 10000 images of the four categories are equally divided into 10 sub-datasets. Each sub-dataset includes 1000 images, and the proportion of the four categories A, B, C, and D in each sub-dataset is 4∶3∶2∶1.
[0039] II. Data Augmentation
[0040] In incremental learning, as the number of incremental training times increases, the model will forget old knowledge, which may lead to a decline in model performance. From the perspective of spectral analysis, the spectral components with large eigenvalues in the incremental process are not easily forgotten. Therefore, high-representation data augmentation based on spectral analysis can be studied to expand the spectral components, so as to obtain more diverse and transferable representation information in the incremental process. Specifically, the similarity of the feature space before and after the new dataset is calculated to quantify the sensitivity of different directions in the model's deep feature space. First, use the old dataset to train the pre-trained image classification model (the image classification model generally includes a feature extraction network and a classifier, which are known in the field of image classification and will not be elaborated in the present invention) to obtain the old feature extraction network , and then use the new dataset to update the old feature extraction network to obtain the new feature extraction network It should be noted that the feature extraction network can adopt a Convext network, a ResNet network, or other neural networks. Then, the old feature extraction network is used again and the new feature extraction network to process the old dataset to obtain the depth features mapped by the old feature extraction network and the depth features mapped by the new feature extraction network Then, the depth features mapped by the old feature extraction network and the depth features mapped by the new feature extraction network are decomposed into different directions. The decomposition method is as follows:
[0041]
[0042] Among them, represents the number of dataset samples, represents the i-th sample, represents the depth features of the i-th sample, represents the transpose of represents the dimension of the feature space, represents the eigenvalue corresponding to the j-th dimension of the feature space, represents the eigenvector corresponding to the j-th dimension of the feature space, represents the transpose of Through decomposition, two sets of eigenvectors can be obtained. Among them, one set of vectors is used to represent the original representation information , and the other set of vectors is used to represent the new representation information Finally, according to the two sets of eigenvectors obtained by decomposition, the forgetfulness and transferability of each direction in the model depth feature space are studied. The angular relationship ψ is used to explore the distance between the spaces corresponding to the old feature extraction network and the new feature extraction network
[0043]
[0044] Among them, represents the -th eigenvector with the -th largest eigenvalue in the feature space corresponding to the old feature extraction network , represents the -th eigenvector with the -th largest eigenvalue in the feature space corresponding to the new feature extraction network . And It should be noted that the spatial distance between feature extraction networks can be understood as the distance between two two-dimensional planes in a three-dimensional coordinate system.
[0045] During the incremental learning process, the retention of old knowledge is reflected at the representation level. The shape of the representation feature distribution, that is, the covariance, should not change much. If the direction of a feature vector only changes slightly after updating the feature extraction network, it is reflected as a very small change angle. Therefore, the size of the change angle of the feature vector after incremental training can reflect which are the key features and which are the features with stronger transferability and less likely to be forgotten. When the angle change is small, it indicates that the offset between the front and back feature spaces is small, and this feature vector is not easily forgotten and has stronger transferability; when the angle change is large, it means that the offset between the front and back feature spaces is large, and this feature vector is more easily forgotten. Through the above research, it is found that after updating the feature extraction network, the larger the eigenvalue of the feature vector, the smaller the angle change, which also means that the larger the eigenvalue of the feature vector, the less likely it is to be forgotten and the more transferable it is after incremental training. At the same time, it indicates that the feature vector with a larger eigenvalue is a key vector. It can be seen that learning more transferable features during the incremental learning process can alleviate the catastrophic forgetting problem.
[0046] As can be seen from the foregoing spectral analysis research, the feature vector with a larger eigenvalue is a key vector. In order to learn more key vectors to alleviate the catastrophic forgetting problem during the incremental learning process, the present invention enhances the number of feature vectors with larger eigenvalues that the model sees by enhancing the data during the incremental process. Further, it can enable the model to learn more transferable features to alleviate catastrophic forgetting. According to an embodiment of the present invention, the present invention proposes a self-supervised enhancement method to perform self-enhancement processing on the data to improve the diversity of data features, and proposes a category enhancement method to perform category enhancement processing on the data to improve the diversity of data categories.
[0047] According to an embodiment of the present invention, the present invention performs self-supervised enhancement processing on the data in a rotational manner. Specifically, the samples in the dataset are rotated by a preset angle. According to an embodiment of the present invention, the preset angle can be 90°, 180°, and 270°. Based on such rotational processing, if the original dataset includes K types of samples, the K types of samples in the original dataset after self-supervised enhancement processing will be expanded to 4K types of samples, and a new label will be assigned to the rotated samples according to the rotation angle. The self-supervised enhancement method relaxes certain invariance constraints when simultaneously learning the original task and the self-supervised task compared with the currently widely used 4-way self-supervised task, which is beneficial to learning richer features during the incremental learning process.
[0048] According to an embodiment of the present invention, the present invention performs category enhancement processing on data in the form of sample combination. Specifically, the present invention randomly samples two samples of different categories from the data set and generates a new sample for adding to the data set according to the following method:
[0049]
[0050] Among them, represents the new sample generated based on samples and . represents the sample in category α in the data set, represents the sample in category β in the data set, and μ represents the interpolation coefficient. According to an embodiment of the present invention, μ is sampled from the interval [0.3, 0.7]. Based on such category expansion processing, if the original data set includes K categories of samples, the K categories of samples in the original data set after category enhancement processing will be expanded to M + K categories of samples, where M = K(K - 1) / 2. Since the data set after category enhancement processing includes more categories, it is possible to see more categories during the incremental learning process to learn transferable and diverse features.
[0051] III. Key knowledge transfer
[0052] As can be seen from the previous embodiments, by performing enhancement processing on the data set, the model can learn more generalizable and transferable features during the incremental learning process. Only by transmitting these features during the incremental learning process can the model's catastrophic forgetting be alleviated. According to an embodiment of the present invention, during the incremental training process, the present invention uses key knowledge transfer to transfer the feature extraction network parameters of the image classification model after the previous training to the feature extraction network of the image classification model to obtain the initial image classification model for the current training. Through knowledge transfer, the feature extraction network can be better initialized, so that the initialized feature extraction network retains the knowledge learned before and avoids catastrophic forgetting. To better understand the present invention, the basic principle of key knowledge transfer is described below.
[0053] During the incremental process, there is a mapping relationship, i.e., a semantic relationship, between the old and new datasets. And since the old and new models between the old and new datasets are related, the similarity between the old and new datasets can help achieve incremental training. As the number of incremental training times increases, the number of old datasets becomes larger and larger, and the probability of the old and new datasets being related to each other also becomes greater, thus promoting the transfer of key features. Driven by the semantic relationship, the present invention connects the old and new datasets through model reuse. Assuming that the semantic relationship between the old and new datasets has been extracted, semantic mapping can transfer a feature extraction network from the original data to the target data, that is, the semantic mapping takes the original feature extraction network as the input and generates a feature extraction network that matches the target data.
[0054] Semantic mapping can capture the correlation between the old and new datasets and can transfer the original feature extraction network into a target feature extraction network . Therefore, semantic mapping can be used to weight the transformed predictions. Semantic mapping encodes the sample-level semantic association between the old data with dimension α and the new data with dimension β. The more associations there are between the old and new data, the larger the value of the corresponding semantic mapping.
[0055] Define , where is a d-dimensional simple form, is a d-dimensional positive real number, represents the normalized marginal probability of the importance of each sample in the old dataset with dimension α, represents the normalized marginal probability of the importance of each sample in the new dataset with dimension β, , are set to a uniform distribution and have no information prior. Introduce a cost matrix to describe the sample changes in the old and new datasets and guide the transfer. Its elements indicate the cost required to link the old dataset to the new dataset. Considering semantic mapping as the coupling of two distributions can connect the samples between the tasks with the lowest transportation cost and optimize it by minimizing:
[0056] (1)
[0057] where T1 represents that the sum of all elements in the matrix is 1, represents the transpose matrix of T1, T represents semantic mapping, , which shows how to align the old dataset with the new dataset. At this time, the probability mass of a sample will migrate to a similar sample at a relatively low cost, and the correct alignment mapping relationship between the old dataset and the new dataset will also be output. By applying the semantic mapping T, the present invention can convert the trained feature extraction network from the previous task into the feature extraction network for the current task. It should be noted that the present invention uses the Sinkhorn algorithm to solve the optimal transport problem.
[0058] From the above, it can be seen that the cost matrix C characterizes the relationship between the old and new datasets. If you want to solve the cost matrix C, you need to first solve the sample centers of each category in the dataset, and then use the sample centers to calculate the costs between different categories in the dataset and construct the cost matrix C. The sample centers are solved in the following way:
[0059] (2)
[0060] Among them, represents the sample center of category n, represents the number of samples in the dataset, represents the label of the i-th sample, represents the i-th sample, represents the feature extraction network.
[0061] If the samples between two different categories in the same dataset are related, their corresponding sample centers will also be close to each other. The pairwise Euclidean distance is used to measure the cost between samples:
[0062] (3)
[0063] Among them, represents the cost between the samples of category n and category m, represents the sample center of category n, represents the sample center of category m. The larger the distance, the greater the difference between the samples of the two categories in the dataset, and the more difficult it is to reuse the specific coefficients of the previously well-trained model for samples with greater differences between different categories.
[0064] In the above solution process, after solving formula (2), the sample centers of different categories can be obtained. Then, based on formula (3), the cost matrix C can be obtained, and based on formula (1), the semantic mapping relationship T based on optimal knowledge transfer can be obtained. Since the key knowledge transfer technology is a technology known to those skilled in the art, the present invention will not elaborate too much. In the following embodiments, the semantic mapping relationship T is used to represent the key knowledge transfer.
[0065] In incremental training, when faced with new tasks, the model needs to quickly adapt to the feature space of the new data, that is, to improve the classification ability of the new data set while reducing the catastrophic forgetting of the old data set. In order to better enable the model to quickly adapt to the feature space of the new data in the incremental learning process, the present invention uses optimal knowledge transfer to solve the feature space adaptation problem through semantic mapping between the new and old data sets. Under the guidance of Building a new feature extraction network , that is, using the optimal transmission method based on the old feature extraction network Parameters to initialize the new feature extraction network , based on this feature extraction network The old feature extraction network is well utilized and the semantic relationship between samples is preserved. At the same time, the calibration between the new and old samples is maintained because the relationship between samples in the new and old datasets is captured through semantic mapping. Therefore, even without training, the new feature extraction network obtained after optimal knowledge transfer has the ability to accurately classify the new dataset while maintaining the ability to classify the old dataset.
[0066] 4. Training Process
[0067] According to one embodiment of the present invention, the present invention uses multiple sub-datasets divided by a new image data set to perform multiple incremental training on a pre-trained image classification model. In order to more clearly illustrate the relevant technical features of the present invention, the present invention takes a complete incremental training as an example for introduction, wherein each incremental training includes steps S1 and S2. Each step is described in detail below.
[0068] In step S1, the sub-dataset of the current training is enhanced. The sub-dataset of the current training is first subjected to self-supervised enhancement, and then subjected to category enhancement. Specifically, the process of enhancement is described by taking the image samples of category a and category b in the sub-dataset of the current training as an example. First, self-supervised enhancement is performed, that is, the image samples of category a and the image samples of category b are rotated by 90°, 180° and 270°. After the self-supervised enhancement, the image categories are expanded from the original 2 categories to 8 categories. Then, the image samples after the self-supervised enhancement are subjected to category enhancement, that is, the expanded 8 categories of image samples are subjected to category enhancement according to the processing method described in the aforementioned embodiment. After the category enhancement, the image categories are expanded from 8 categories to 36 categories. It should be noted that the sub-dataset can be enhanced only by the self-supervised enhancement method, or the sub-dataset can be enhanced only by the category enhancement method.
[0069] In step S2, the parameters of the image classification model after the previous training are used to initialize the image classification model, and the enhanced current training sub-dataset is used to train it until convergence. During the training process, the cross-entropy loss, knowledge distillation loss, and knowledge transfer loss are used to update the model parameters.
[0070] According to an embodiment of the present invention, step S2 includes steps S21 and S22.
[0071] In step S21, the optimal knowledge transfer method is used to transfer the feature extraction network parameters of the image classification model after the previous training to the feature extraction network of the image classification model to obtain the initial image classification model for the current training. That is, the semantic mapping relationship T between the previous training sub-dataset and the current training sub-dataset is calculated according to the optimal knowledge transfer method, and under the guidance of the semantic mapping relationship the feature extraction network of the image classification model after the previous training is used to construct a new feature extraction network , , and the parameters of the new feature extraction network are used to initialize the feature extraction network of the image classification model for the current training , obtaining the initial image classification model for the current training. Using the semantic mapping relationship T to guide the feature extraction network of the image classification model after the previous training to construct a new feature extraction network , that is, using the semantic mapping relationship T as a constraint condition, taking the feature extraction network parameters as input, substituting them into the semantic mapping relationship T to output the feature extraction network parameters.
[0072] In step S22, the current training initial image classification model is trained until convergence using the current training sub-dataset, and the cross-entropy loss, knowledge distillation loss, and knowledge transfer loss are used to update the parameters of the current training initial image classification model.
[0073] According to an embodiment of the present invention, the following method is used to update the parameters of the current training initial image classification model in step S22:
[0074]
[0075] represents the total loss function, represents the cross-entropy loss function, represents the knowledge distillation loss function, represents the knowledge transfer loss function, and denotes hyperparameters, \(x\) denotes a sample, and \(y\) denotes the label of sample \(x\). According to an example of the present invention, the hyperparameters and have a value of 10.
[0076] According to an embodiment of the present invention, the knowledge distillation loss is calculated in the following manner:
[0077]
[0078] wherein, denotes the number of samples in the sub-dataset of the previous training after augmentation processing, denotes a sample, denotes the softmax function, denotes the classifier of the image classification model after the previous training, denotes the feature extraction network of the image classification model after the previous training, denotes the classifier of the initial image classification model of the current training, denotes the feature extraction network of the image classification model of the current training. It should be noted that the image classification model after the previous training will not be retrained according to the sub-dataset of the current training during the current incremental training.
[0079] According to an embodiment of the present invention, the knowledge transfer loss is calculated in the following manner:
[0080]
[0081] wherein, denotes the number of samples in the sub-dataset of the current training after augmentation processing, denotes a sample, denotes the softmax function, denotes the classifier of the image classification model after the previous training, denotes the feature extraction network of the image classification model after the previous training, denotes the classifier of the initial image classification model of the current training, denotes the feature extraction network of the initial image classification model of the current training.
[0082] It should be noted that the cross-entropy loss is a common loss function in the field of deep learning, so it will not be described in detail in the present invention.
[0083] V. Verification process
[0084] To verify the effectiveness of the present invention, six different fundus disease image datasets are used in the comparative experiment of the present invention to compare the differences between the present invention and the prior art. The first dataset includes 12,238 high-quality clinical fundus disease image samples and their corresponding annotation labels, and there are image samples of four types of fundus diseases in the dataset, namely age-related macular degeneration (AMD), diabetic retinopathy (DR), glaucoma, and normal retina image samples. The second dataset includes 13,812 fundus disease image samples with uneven quality levels and their corresponding annotation labels, and there are image samples of two types of fundus diseases in the dataset, namely age-related macular degeneration (AMD) and diabetic retinopathy (DR) image samples. The third dataset includes 1,748 retinal color image samples, and all the image samples in the dataset are diabetic retinopathy (DR) image samples. The fourth dataset includes 15 fundus disease image samples each from normal patients, glaucoma patients, and diabetic (DR) patients. The fifth dataset includes 650 glaucoma retina image samples. The sixth dataset includes 40 digital retina image samples and their corresponding annotation labels.
[0085] The incremental learning experiment is carried out using the first dataset to simulate different data ratios. Specifically, the first dataset is divided into a training set and a test set, and the selection ratio of the training set to the test set is 4:1. And the test set is equally divided into 2, 5, 10 equal parts respectively, and 2 incremental trainings, 5 incremental trainings, and 10 incremental trainings are carried out respectively according to the content of the embodiments of the present invention. After each incremental training is completed, the test set is used to test the image classification model obtained after the current training to obtain the model accuracy of the image classification model after the current training. The accuracy results of the image classification models after 2, 5, and 10 incremental trainings are as Figures 2 - 4As shown, it can be seen from the figure that the accuracy rates of the image classification models obtained after 2 incremental trainings, 5 incremental trainings, and 10 incremental trainings respectively according to the embodiments of the present invention are all better than those of the image classification models obtained after incremental training according to the prior art. Oracle represents the test accuracy obtained by training using all data. Among them, the prior art includes End-to-End Incremental Learning (EEIL), Learning a Unified Classifier Incrementally via Rebalancing (LUCIR), Learning without Forgetting (LwF), Memory Aware Synapses (MAS), Learning without Memorizing (LwM), and Multi-Classifier (MUC).
[0086] Furthermore, in order to simulate the practical problems of data quality differences and incomplete data categories when the dataset used in the current incremental training is different from the dataset used in the previous incremental training in terms of data acquisition devices or data acquisition personnel. First, use part of the data in the first dataset to complete an incremental training, then put the data in the other five datasets in subsequent incremental trainings. Finally, by obtaining the confusion matrix of each incremental training process, compare the differences between the present invention and the prior art, and verify the effectiveness of the present invention in reducing misclassification situations. The verification results are as Figures 5 - 7 shown. The confusion matrix is an analysis table that summarizes the prediction results of a classification model in machine learning. It summarizes the records in the dataset according to two criteria: the true category and the category predicted by the classification model in matrix form. The rows of the matrix represent the true values, and the columns of the matrix represent the predicted values. The darker the color on the diagonal from the upper left to the lower right in the matrix pair, the better the effect of the model. Therefore, it can be seen from the figure that when there are problems of data quality differences and incomplete data categories, the accuracy of the image classification model after incremental training according to the present invention is significantly better than that of the image classification model obtained by training with the multi-classifier-based incremental learning method, and there is little difference from the accuracy of the image classification model obtained by training with the Dynamically Expandable Representation for Class Incremental Learning (DER) method.
[0087] As can be seen from the above, based on spectral analysis, it is crucial to study the features with certain attributes during the incremental process. Then, self-supervised learning data augmentation and class augmentation methods are adopted to enable the incremental model to learn more generalizable and transferable representations. Further, to alleviate the catastrophic forgetting problem during the incremental process, the optimal knowledge transfer is utilized to obtain the semantic relationship between the new and old datasets, realizing the transfer of the model feature space and alleviating the catastrophic forgetting problem during the incremental training process.
[0088] It should be noted that although the above steps are described in a specific order, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order, as long as the required functions can be achieved.
[0089] The present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0090] The computer-readable storage medium can be a tangible device that retains and stores instructions for use by an instruction execution device. The computer-readable storage medium may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing.
[0091] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the technical field to understand the disclosed embodiments.
Claims
1. A method for training an image classification model based on knowledge transfer, which is used to perform incremental training on a pre-trained image classification model, wherein, The pre-trained image classification model includes a feature extraction network and a classifier. It is characterized in that the method includes incrementally training the pre-trained image classification model with a new image dataset in the following manner: S1. Perform enhancement processing on the current new image dataset; S2. Initialize the image classification model with the parameters of the image classification model after the previous training, and train it with the enhanced current new image dataset until convergence. During the training process, cross-entropy loss, knowledge distillation loss, and knowledge transfer loss are used to update the model parameters, where: The knowledge distillation loss is calculated using the following formula: Among them, represents the number of samples of the image dataset of the previous training after enhancement processing, represents a sample, represents the softmax function, represents the classifier of the image classification model after the previous training, represents the feature extraction network of the image classification model after the previous training, represents the classifier of the initial image classification model of the current training, represents the feature extraction network of the image classification model of the current training; The knowledge transfer loss is calculated using the following formula: Among them, represents the number of samples of the new image dataset for the current training after enhancement processing, represents a sample, represents the softmax function, represents the classifier of the image classification model after the previous training, represents the feature extraction network of the image classification model after the previous training, represents the classifier of the initial image classification model for the current training, represents the feature extraction network of the initial image classification model for the current training.
2. The method according to claim 1, characterized in that, The step S1 includes: Performing self-supervised enhancement processing on the current new image dataset, or performing class enhancement processing on the current new image dataset, or first performing self-supervised enhancement processing on the current new image dataset and then performing class enhancement processing.
3. The method according to claim 2, characterized in that: The self-supervised enhancement processing is to flip the samples in the current new image dataset at rotation angles of 90°, 180°, and 270°; The class enhancement processing refers to randomly sampling two samples of different classes from the current new image dataset and generating a new sample for class augmentation in the following manner to be added to the current new image dataset: Among them, represents a new sample generated based on the sample and represents a sample in class α in the dataset, represents a sample in class β in the dataset, and μ represents the interpolation coefficient. 4. The method according to claim 1, wherein The step S2 includes: S21. Use the optimal knowledge transfer method to transfer the feature extraction network parameters of the image classification model after the previous training to the feature extraction network of the image classification model to obtain the initial image classification model for the current training; S22. Train the initial image classification model for the current training with the current new image dataset until convergence, and use cross-entropy loss, knowledge distillation loss, and knowledge transfer loss to update the parameters of the initial image classification model for the current training.
5. The method according to claim 4, characterized in that, In step S22, the total loss function is calculated according to the following formula: Among them, represents the total loss function, represents the cross-entropy loss function, represents the knowledge distillation loss function, represents the knowledge transfer loss function, and represents hyperparameters, x represents a sample, and y represents the label of sample x.
6. An image classification method, characterized in that, The method includes: T1. Obtain the image to be processed; T2. Use the image classification model trained by the method according to any one of claims 1-5 to process the image to be processed to obtain a classification result.
7. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method according to any one of claims 1-6.
8. An electronic device, characterized in that, Including: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the electronic device to implement the steps of the method according to any one of claims 1-6.
Citation Information
Patent Citations
Incremental track anomaly detection method based on incremental kernel principle component analysis
CN106022368A
Image big data-oriented class increment classification method, system and device and medium
CN112990280A
Image defogging method based on incremental learning and feature and attention transfer
CN113066025A
Method of incremental learning for object detection
US20200302230A1
Distributed model cooperative training method and system
CN114626550A