A Robust Image Classification Method and System Based on Label Embedding
By learning label embedding and using it as the training target of image classification neural network, the problem of insufficient robustness caused by label noise in image classification is solved, and higher classification accuracy and robustness are achieved, especially effective image classification in the label noise environment.
Patent Information
- Application Number
- CN202210087643.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-01-25
AI Technical Summary
The existing image classification method is not robust enough when facing label noise, resulting in a degradation of classification performance. In the existing method, the single-hot encoded label form fails to effectively utilize similarity information between categories.
A robust image classification method based on label embedding is adopted. By learning ideal label embedding and using it as the goal of image classification neural network training, the label embedding vector in the target comprehensive loss function is used to reduce the impact of label noise, and the parameters are updated by gradient descent method to enhance the robustness of the image classifier.
It effectively alleviates the negative impact of noise labels in single-hot encoding form on image classification, improves the accuracy and robustness of image classification, especially in the presence of label noise, which can better mine similarity information between categories.
Smart Images

Figure CN114445662B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pattern recognition, and particularly to a robust image classification method and system based on label embedding. Background Art
[0002] In an album management system, since there are a large number of image types and quantities, it is very difficult to find pictures of a target person. Currently, an image classifier is used to classify pictures of different people to quickly find pictures of the target person.
[0003] Ideally, for the image classification task, image training samples with correct labels are given to train an image classification neural network to obtain an image classifier. However, in reality, the image training samples obtained often contain label noise, that is, the labels of many image training samples are incorrect, and this incorrect supervision information will seriously damage the classification performance of a robust image classifier. In this case, researching robust image classification methods has great research significance and practical value.
[0004] More and more scholars at home and abroad have paid attention to the problem of robust image classification and have done a lot of research in this area. The main methods in recent papers are as follows: Nagarajan Natarajan et al. designed an unbiased estimator for the loss function in 2013 to mitigate the negative impact of label noise on image classification; Giorgio Patrini et al. decomposed the loss function into two parts in 2016: the first part is independent of label noise, and only the second part considers the influence of label noise. After the decomposition, by adjusting the traditional stochastic gradient descent method, the purpose of training a robust image classifier is achieved; Daiki Tanaka et al. proposed a method in 2018, specifically: after each round of training, use the results estimated by the model to modify the labels of all image samples; then use the modified labels for the next round of training, thus forming an iterative process to finally obtain a robust image classifier; Fabrice Muhlenbach et al. proposed using the neighborhood relationship between image samples to screen out image samples that may carry noise in 2004; Lu Jiang et al. proposed MentorNet in 2018, which is to pre-train an auxiliary network specifically for screening noise image samples. After screening by the auxiliary network, the remaining clean image samples are used to update the parameters of the image classification neural network.
[0005] Although the above methods can all achieve good results, they still have a common drawback. That is, during the process of optimizing the classifier parameters, the actually used labels all adopt the one-hot encoding form. That is to say, only one element in all positions of each vector is 1, and the other positions are 0. Such a discrete-value label representation form does not contain the similarity information between different category labels and may not be the most suitable for robust image classification. Summary of the Invention
[0006] The object of the present invention is to provide a robust image classification method and system based on label embedding to achieve the purpose of quickly and accurately classifying the portrait pictures stored in the album management system.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A robust image classification method based on label embedding, comprising:
[0009] Obtain a target picture to be classified; the target picture is a picture containing a portrait area;
[0010] Classify the target picture to be classified according to a robust image classifier to obtain a classification result of the target picture;
[0011] Wherein, the robust image classifier is trained based on a target comprehensive loss function; the target comprehensive loss function includes three sub-functions, namely an original label image classification loss sub-function, a label learning loss sub-function, and a new label image classification loss sub-function; the original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the sample picture and the original label classification result of the portrait area; the label learning loss sub-function is used to represent the loss value when the original label classification result of the portrait area learns the new label classification result of the portrait area; the new label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the sample picture and the new label classification result of the portrait area.
[0012] Optionally, it further includes a training method for the robust image classifier; the training method includes:
[0013] Obtain training data; the training data includes a plurality of data pairs; the data pair includes a sample picture and the original label classification result of the portrait area corresponding to the sample picture; the original label classification result of the portrait area corresponding to the sample picture contains noise information;
[0014] Determine the target comprehensive loss function;
[0015] Determine the final parameters of the image classification neural network based on the training data, the target comprehensive loss function, and the image classification neural network;
[0016] Determine a robust image classifier based on the final parameters of the image classification neural network.
[0017] Optionally, the obtaining of the training data specifically includes:
[0018] Obtain the original data; the original data includes a plurality of original data pairs; the original data pair includes an original sample picture and the original label classification result of the portrait area corresponding to the original sample picture; the original label classification result of the portrait area corresponding to the original sample picture contains noise information;
[0019] Perform normalization processing on the original data to obtain the training data.
[0020] Optionally, the target comprehensive loss function includes a first target comprehensive loss function and a second target comprehensive loss function; the determining of the target comprehensive loss function specifically includes:
[0021] Determine a first target sample set and a second target sample set; the data pairs in the first target sample set and the second target sample set are both part of the training data, and the data pairs in the first target sample set are different from the data pairs in the second target sample set;
[0022] Determine a first original label image classification loss sub-function of the first image classification neural network and a second original label image classification loss sub-function of the second image classification neural network; both the first original label image classification loss sub-function and the second original label image classification loss sub-function are cross-entropy loss functions; the first original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the second calibrated sample picture in the second target sample set and the original label classification result of the portrait area corresponding to the second calibrated sample picture; the second original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the first calibrated sample picture in the first target sample set and the original label classification result of the portrait area corresponding to the first calibrated sample picture; the first calibrated sample picture is any sample picture in the first target sample set; the second calibrated sample picture is any sample picture in the second target sample set;
[0023] Determine the first label learning loss sub-function of the first image classification neural network and the second label learning loss sub-function of the second image classification neural network; the first label learning loss sub-function is used to represent the loss value between the first learning result and the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture; the second label learning loss sub-function is used to represent the loss value between the second learning result and the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture; the first learning result is the learning result obtained by fitting and learning the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture with the original label classification result of the portrait region corresponding to the second calibrated sample picture; the second learning result is the learning result obtained by fitting and learning the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture with the original label classification result of the portrait region corresponding to the first calibrated sample picture; wherein, the first learning result is the new label classification result of the portrait region corresponding to the second calibrated sample picture, and the second learning result is the new label classification result of the portrait region corresponding to the first calibrated sample picture;
[0024] Determine the first new label image classification loss sub-function of the first image classification neural network and the second new label image classification loss sub-function of the second image classification neural network; the first new label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture in the second target sample set and the first learning result; the second new label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture in the first target sample set and the first result;
[0025] The first target comprehensive loss function includes three sub-functions, namely the first original label image classification loss sub-function, the first label learning loss sub-function, and the first new label image classification loss sub-function;
[0026] The second target comprehensive loss function includes three sub-functions, namely the second original label image classification loss sub-function, the second label learning loss sub-function, and the second new label image classification loss sub-function.
[0027] Optionally, based on the training data, the target comprehensive loss function, and the image classification neural network, determine the final parameters of the image classification neural network, specifically including:
[0028] Determine the final parameters of the first image classification neural network and the final parameters of the second image classification neural network based on the first target sample set, the second target sample set, the first target comprehensive loss function, the second target comprehensive loss function, the first image classification neural network, and the second image classification neural network;
[0029] Wherein, the robust image classifier is determined according to the final parameters of the first image classification neural network or determined according to the final parameters of the second image classification neural network.
[0030] Optionally, the first image classification neural network and the second image classification neural network are both the same image classification neural network; the image classification neural network is a network obtained by adding an embedding layer on the basis of a convolutional neural network; the embedding layer is used to convert the original label classification result of the portrait area into a new label classification result of the portrait area.
[0031] A robust image classification system based on label embedding, comprising:
[0032] A data acquisition module, configured to acquire a target picture to be classified; the target picture is a picture containing a portrait area;
[0033] A classification result determination module, configured to classify the target picture to be classified according to a robust image classifier to obtain a classification result of the target picture;
[0034] Wherein, the robust image classifier is trained based on a target comprehensive loss function; the target comprehensive loss function includes three sub-functions, namely an original label image classification loss sub-function, a label learning loss sub-function, and a new label image classification loss sub-function; the original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the sample picture and the original label classification result of the portrait area; the label learning loss sub-function is used to represent the loss value when the original label classification result of the portrait area learns the new label classification result of the portrait area; the new label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the sample picture and the new label classification result of the portrait area.
[0035] Optionally, it further includes a training module; the training module includes:
[0036] A training data acquisition unit, configured to acquire training data; the training data includes a plurality of data pairs; the data pair includes a sample picture and the original label classification result of the portrait area corresponding to the sample picture; the original label classification result of the portrait area corresponding to the sample picture contains noise information;
[0037] A target comprehensive loss function determination unit for determining a target comprehensive loss function;
[0038] A final parameter determination unit of the image classification neural network for determining final parameters of the image classification neural network based on the training data, the target comprehensive loss function, and the image classification neural network;
[0039] A robust image classifier determination unit for determining a robust image classifier based on the final parameters of the image classification neural network.
[0040] Optionally, the training data acquisition unit specifically includes:
[0041] An original data subunit for acquiring original data; the original data includes a plurality of original data pairs; each original data pair includes an original sample picture and an original label classification result of the portrait area corresponding to the original sample picture; the original label classification result of the portrait area corresponding to the original sample picture contains noise information;
[0042] A normalization processing subunit for performing normalization processing on the original data to obtain training data.
[0043] Optionally, the target comprehensive loss function includes a first target comprehensive loss function and a second target comprehensive loss function; the target comprehensive loss function determination unit specifically includes:
[0044] A target sample set subunit for determining a first target sample set and a second target sample set; the data pairs in the first target sample set and the second target sample set are both part of the training data, and the data pairs in the first target sample set are different from the data pairs in the second target sample set;
[0045] The original label image classification loss sub-function determination subunit is used to determine the first original label image classification loss sub-function of the first image classification neural network and the second original label image classification loss sub-function of the second image classification neural network; both the first original label image classification loss sub-function and the second original label image classification loss sub-function are cross-entropy loss functions; the first original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture in the second target sample set and the original label classification result of the portrait region corresponding to the second calibrated sample picture; the second original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture in the first target sample set and the original label classification result of the portrait region corresponding to the first calibrated sample picture; the first calibrated sample picture is any sample picture in the first target sample set; the second calibrated sample picture is any sample picture in the second target sample set;
[0046] The label learning loss sub-function determination subunit is used to determine the first label learning loss sub-function of the first image classification neural network and the second label learning loss sub-function of the second image classification neural network; the first label learning loss sub-function is used to represent the loss value between the first learning result and the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture; the second label learning loss sub-function is used to represent the loss value between the second learning result and the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture; the first learning result is the learning result obtained after the original label classification result of the portrait region corresponding to the second calibrated sample picture fits and learns the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture; the second learning result is the learning result obtained after the original label classification result of the portrait region corresponding to the first calibrated sample picture fits and learns the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture; wherein, the first learning result is the new label classification result of the portrait region corresponding to the second calibrated sample picture, and the second learning result is the new label classification result of the portrait region corresponding to the first calibrated sample picture;
[0047] A new label image classification loss sub - function determination subunit is used to determine the first new label image classification loss sub - function of the first image classification neural network and the second new label image classification loss sub - function of the second image classification neural network; the first new label image classification loss sub - function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture in the second target sample set and the first learning result; the second new label image classification loss sub - function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture in the first target sample set and the first result.
[0048] The first target comprehensive loss function includes three sub - functions, namely the first original label image classification loss sub - function, the first label learning loss sub - function, and the first new label image classification loss sub - function.
[0049] The second target comprehensive loss function includes three sub - functions, namely the second original label image classification loss sub - function, the second label learning loss sub - function, and the second new label image classification loss sub - function.
[0050] According to the specific embodiments provided by the present invention, the following technical effects are disclosed:
[0051] Based on the problems mentioned in the background art, the present invention provides a robust image classification method and system based on label embedding. On the basis of the prior art, the present invention seeks a more reasonable label form to improve the effect of label - robust image classification. To achieve this goal, the present invention learns an ideal label embedding and then adds this learned label embedding as a new learning target to the target comprehensive loss function for training the robust image classifier. Different from the one - hot encoding label form of either 1 or 0, each element in the label embedding vector is a continuous value within the range of 0 to 1. During the process of label embedding learning, the similarity information represented by continuous values between different categories can be gradually mined. Finally, the robust image classifier can greatly alleviate the negative impact of the noisy labels in the one - hot encoding form on image classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0053] Figure 1 It is a schematic flowchart of the robust image classification method based on label embedding of the present invention.
[0054] Figure 2 Schematic diagram of the training process of the robust image classifier based on label embedding of the present invention;
[0055] Figure 3 Result curve graph of the present invention under the condition of 50% noise rate on the Cifar10 dataset;
[0056] Figure 4 Result curve graph of the present invention under the condition of 60% noise rate on the Cifar10 dataset;
[0057] Figure 5 Result curve graph of the present invention under the condition of 70% noise rate on the Cifar10 dataset;
[0058] Figure 6 Schematic diagram of the structure of the robust image classification system based on label embedding of the present invention. Detailed implementation manners
[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0060] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0061] The object of the present invention is to provide a robust image classification method and system based on label embedding; the present invention learns an ideal label embedding as a new training target for the image classification neural network, adds a cross-entropy loss that minimizes the distance between the output of the image classification neural network and the label embedding in the target loss function, then updates the parameters of the image classification neural network using the gradient descent method, and then classifies the image test data, thereby training a robust image classifier, and using this classifier to classify pictures.
[0062] An example is that the robust image classifier in the present application can classify pictures, thereby labeling pictures of different categories (for example, picture classification based on different users) for easy viewing and searching by users. In addition, the classification labels of these pictures can also be provided to the album management system for classification management, saving the management time of users, improving the efficiency of album management, and enhancing the user experience.
[0063] For example, the robust image classifier in the present application can classify and tag pictures of different users. For instance, when the robust image classifier in the present application is applied to a smart terminal, it can classify pictures in the album of the smart terminal by person, that is, pictures including user A can be labeled with the tag of user A; pictures including user B can be labeled with the tag of user B, realizing the classification of the album based on different users.
[0064] Embodiment 1
[0065] As Figure 1 shown, a robust image classification method based on label embedding provided in this embodiment includes:
[0066] Step 101: Obtain a target picture to be classified; the target picture is a picture containing a portrait area.
[0067] Step 102: Classify the target picture to be classified according to the robust image classifier to obtain the classification result of the target picture.
[0068] Among them, the robust image classifier is trained based on a target comprehensive loss function; the target comprehensive loss function includes three sub-functions, namely, an original label image classification loss sub-function, a label learning loss sub-function, and a new label image classification loss sub-function; the original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the sample picture and the original label classification result of the portrait area; the label learning loss sub-function is used to represent the loss value when the original label classification result of the portrait area learns the new label classification result of the portrait area; the new label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the sample picture and the new label classification result of the portrait area.
[0069] As a preferred implementation manner, the robust image classification method provided in this embodiment further includes a training method for the robust image classifier; the training method includes:
[0070] Obtain training data; the training data includes a plurality of data pairs; the data pair includes a sample picture and the original label classification result of the portrait area corresponding to the sample picture; the original label classification result of the portrait area corresponding to the sample picture contains noise information.
[0071] Determine the target comprehensive loss function.
[0072] Based on the training data, the target comprehensive loss function, and the image classification neural network, determine the final parameters of the image classification neural network.
[0073] Based on the final parameters of the image classification neural network, determine the robust image classifier.
[0074] Among them, the obtaining of the training data specifically includes:
[0075] Obtain the original data; the original data includes a plurality of original data pairs; the original data pair includes an original sample picture and the original label classification result of the portrait area corresponding to the original sample picture; the original label classification result of the portrait area corresponding to the original sample picture contains noise information; perform normalization processing on the original data to obtain the training data.
[0076] The target comprehensive loss function includes a first target comprehensive loss function and a second target comprehensive loss function; the determination of the target comprehensive loss function specifically includes:
[0077] Determine a first target sample set and a second target sample set; the data pairs in the first target sample set and the second target sample set are both part of the training data, and the data pairs in the first target sample set are different from the data pairs in the second target sample set.
[0078] Determine a first original label image classification loss sub-function of the first image classification neural network and a second original label image classification loss sub-function of the second image classification neural network; both the first original label image classification loss sub-function and the second original label image classification loss sub-function are cross-entropy loss functions; the first original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the second calibrated sample picture in the second target sample set and the original label classification result of the portrait area corresponding to the second calibrated sample picture; the second original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the first calibrated sample picture in the first target sample set and the original label classification result of the portrait area corresponding to the first calibrated sample picture; the first calibrated sample picture is any sample picture in the first target sample set; the second calibrated sample picture is any sample picture in the second target sample set.
[0079] Determine the first label learning loss sub-function of the first image classification neural network and the second label learning loss sub-function of the second image classification neural network; the first label learning loss sub-function is used to represent the loss value between the first learning result and the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture; the second label learning loss sub-function is used to represent the loss value between the second learning result and the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture; the first learning result is the learning result obtained by fitting and learning the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture with the original label classification result of the portrait region corresponding to the second calibrated sample picture; the second learning result is the learning result obtained by fitting and learning the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture with the original label classification result of the portrait region corresponding to the first calibrated sample picture; wherein, the first learning result is the new label classification result of the portrait region corresponding to the second calibrated sample picture, and the second learning result is the new label classification result of the portrait region corresponding to the first calibrated sample picture.
[0080] Determine the first new label image classification loss sub-function of the first image classification neural network and the second new label image classification loss sub-function of the second image classification neural network; the first new label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture in the second target sample set and the first learning result; the second new label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture in the first target sample set and the first result.
[0081] The first target comprehensive loss function includes three sub-functions, namely the first original label image classification loss sub-function, the first label learning loss sub-function, and the first new label image classification loss sub-function.
[0082] The second target comprehensive loss function includes three sub-functions, namely the second original label image classification loss sub-function, the second label learning loss sub-function, and the second new label image classification loss sub-function.
[0083] Based on the training data, the target comprehensive loss function, and the image classification neural network, determine the final parameters of the image classification neural network, specifically including:
[0084] Based on the first target sample set, the second target sample set, the first target comprehensive loss function, the second target comprehensive loss function, the first image classification neural network, and the second image classification neural network, determine the final parameters of the first image classification neural network and the final parameters of the second image classification neural network.
[0085] Wherein, the robust image classifier is determined according to the final parameters of the first image classification neural network or according to the final parameters of the second image classification neural network.
[0086] The first image classification neural network and the second image classification neural network are both the same image classification neural network; the image classification neural network is a network obtained by adding an embedding layer on the basis of a convolutional neural network; the embedding layer is used to convert the original label classification result of the portrait area into a new label classification result of the portrait area.
[0087] Embodiment 2
[0088] This embodiment provides a training method for a robust image classifier based on label embedding. This training method learns an ideal label embedding as a new learning target for the image classification neural network, so as to enhance the robustness of the image classification neural network through the semantic information contained in the label embedding to complete robust image classification.
[0089] Combined with Figure 2 , the training method for the robust image classifier based on label embedding provided in this embodiment includes the following steps.
[0090] The first step is to normalize the image data and initialize the parameters of the image classification neural network; specifically: perform normalization processing on the image training samples and initialize the parameters of the two image classification neural networks. Among them, each dimensional feature value of the normalized samples is between 0 and 1.
[0091] The second step is to select a set of clean image samples; specifically: the two image classification neural networks perform forward propagation on all the image training samples, calculate the loss function values of all the image training samples, and sort them from small to large according to the loss function values to obtain their respective selected sets of clean image samples.
[0092] In the third step, learn label embeddings and use the label embeddings to improve the robustness of the image classification neural network. Specifically: First, the two image classification neural networks pass the sets of clean image samples they each selected to the other, and then perform forward propagation again to obtain their respective outputs. Then, the two image classification neural networks pass the outputs obtained at this time to the other as the targets for label embedding learning. Next, the two image classification neural networks calculate the cross-entropy loss for their respective label embeddings and outputs and add it to the objective loss function. Finally, update the parameters of the image classification neural network by the gradient descent method.
[0093] This embodiment can mine the similarity information between different categories through the learning of label embeddings, thereby enhancing the robustness of the image classification neural network to label noise. This embodiment uses the gradient descent method to solve the model to obtain the optimal solution, and thus can classify the image samples in the test set.
[0094] Embodiment Three
[0095] During the training process of the robust image classifier, this embodiment needs to train two image classification neural networks (neural networks) f and g simultaneously, and denote the parameters of the two image classification neural networks as ω f and ω g . Let denote the image training set with label noise containing n samples, where, and represent the d-dimensional feature vector and the one-hot encoded noisy label vector of the i-th sample in the training set (c represents the number of categories), respectively. Specifically, it includes the following:
[0096] First, select a mini-batch subset from the complete image training set, and then pass this subset to the two image classification neural networks f and g respectively, so as to obtain the outputs of the two image classification neural networks for this mini-batch subset.
[0097] Then, use the cross-entropy loss function to calculate the loss function values of each image sample in this mini-batch subset, and then select a part of the image samples with the smallest loss function values. Let the sets of possibly clean image samples selected by the two image classification neural networks (i.e., the part of the image samples with the smallest loss function values above) be D f and D g .
[0098] Next, the two image classification neural networks pass the subsets of the clean image samples they each selected to the other, and then perform forward propagation again to obtain the outputs f(D g ) and g(D f) Thus, the image classification losses of the two image classification neural networks can be constructed:
[0099]
[0100] and
[0101]
[0102] Next, it is introduced how to complete the learning of label embeddings and how to use label embeddings to enhance the robustness of the image classification neural network. Specifically, in the above steps, the two image classification neural networks will each pass their outputs f(D g ) and g(D f ) to the other party. After the two image classification neural networks obtain the outputs of the other party, they use these as the learning objectives for their own label embeddings and the original labels. Therefore, it is necessary to minimize the differences between the label embeddings e f and e g and the outputs g(D f ) and f(D g ) of the other party to learn the ideal label embeddings. The specific loss function is:
[0103]
[0104] and
[0105]
[0106] Now, use the learned label embeddings as a better supervision information than the one-hot encoded label vectors to guide the training of the two image classification neural networks. Specifically, minimize the distances between the outputs f(D g ) and g(D f ) of the two image classification neural networks and their corresponding label embeddings e f and e g . The loss function is:
[0107]
[0108] and
[0109]
[0110] In summary, combine the three sub-objective functions proposed above to obtain the final objective loss function of the two image classification neural networks.
[0111] Finally, use the gradient descent method to update the network parameters until convergence. Classify the test data according to the obtained model, so as to finally complete the robust image classification.
[0112] In this embodiment, the Cifar10 dataset is used to prove the performance of the proposed method. The dataset contains 50,000 RGB images as the training set and 10,000 RGB images as the test set. The image size is 32×32, and there are a total of 10 categories. Since the original dataset is clean, in this embodiment of the present invention, 50%, 60%, and 70% of the samples are randomly selected from the original training set and symmetric label noise is added, that is, the original label will be randomly flipped to any other category according to a given probability. Figure 3 , Figure 4 and Figure 5 respectively show the test set accuracies of the Cifar10 dataset in the above three cases. As Figure 3-5 can be seen, in the three cases of noise rates of 50%, 60%, and 70%, the method provided in this embodiment of the present invention can also achieve a very high accuracy on this dataset. Therefore, the robust image classification method proposed in this embodiment of the present invention has high practical value.
[0113] Embodiment 4
[0114] As Figure 6 shown, this embodiment provides a robust image classification system based on label embedding, including:
[0115] A data acquisition module 601, configured to acquire a target picture to be classified; the target picture is a picture containing a portrait area.
[0116] A classification result determination module 602, configured to classify the target picture to be classified according to a robust image classifier to obtain a classification result of the target picture.
[0117] Wherein, the robust image classifier is trained based on a target comprehensive loss function; the target comprehensive loss function includes three sub-functions, namely an original label image classification loss sub-function, a label learning loss sub-function, and a new label image classification loss sub-function; the original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the sample picture and the original label classification result of the portrait area; the label learning loss sub-function is used to represent the loss value when the original label classification result of the portrait area learns the new label classification result of the portrait area; the new label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the sample picture and the new label classification result of the portrait area.
[0118] Wherein, the robust image classification system based on label embedding provided in this embodiment further includes a training module; the training module includes:
[0119] A training data acquisition unit for acquiring training data; the training data includes a plurality of data pairs; each data pair includes a sample image and an original label classification result of the portrait area corresponding to the sample image; the original label classification result of the portrait area corresponding to the sample image contains noise information.
[0120] A target comprehensive loss function determination unit for determining a target comprehensive loss function.
[0121] A final parameter determination unit of the image classification neural network for determining the final parameters of the image classification neural network based on the training data, the target comprehensive loss function, and the image classification neural network.
[0122] A robust image classifier determination unit for determining a robust image classifier based on the final parameters of the image classification neural network.
[0123] The training data acquisition unit specifically includes:
[0124] An original data subunit for acquiring original data; the original data includes a plurality of original data pairs; each original data pair includes an original sample image and an original label classification result of the portrait area corresponding to the original sample image; the original label classification result of the portrait area corresponding to the original sample image contains noise information; a normalization processing subunit for performing normalization processing on the original data to obtain training data.
[0125] The target comprehensive loss function includes a first target comprehensive loss function and a second target comprehensive loss function; the target comprehensive loss function determination unit specifically includes:
[0126] A target sample set subunit for determining a first target sample set and a second target sample set; the data pairs in the first target sample set and the second target sample set are both part of the training data, and the data pairs in the first target sample set are different from the data pairs in the second target sample set.
[0127] The original label image classification loss sub-function determination subunit is used to determine the first original label image classification loss sub-function of the first image classification neural network and the second original label image classification loss sub-function of the second image classification neural network; both the first original label image classification loss sub-function and the second original label image classification loss sub-function are cross-entropy loss functions; the first original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture in the second target sample set and the original label classification result of the portrait region corresponding to the second calibrated sample picture; the second original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture in the first target sample set and the original label classification result of the portrait region corresponding to the first calibrated sample picture; the first calibrated sample picture is any sample picture in the first target sample set; the second calibrated sample picture is any sample picture in the second target sample set.
[0128] The label learning loss sub-function determination subunit is used to determine the first label learning loss sub-function of the first image classification neural network and the second label learning loss sub-function of the second image classification neural network; the first label learning loss sub-function is used to represent the loss value between the first learning result and the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture; the second label learning loss sub-function is used to represent the loss value between the second learning result and the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture; the first learning result is the learning result obtained after the original label classification result of the portrait region corresponding to the second calibrated sample picture fits and learns the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture; the second learning result is the learning result obtained after the original label classification result of the portrait region corresponding to the first calibrated sample picture fits and learns the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture; wherein, the first learning result is the new label classification result of the portrait region corresponding to the second calibrated sample picture, and the second learning result is the new label classification result of the portrait region corresponding to the first calibrated sample picture.
[0129] A new label image classification loss sub-function determination subunit is configured to determine a first new label image classification loss sub-function of the first image classification neural network and a second new label image classification loss sub-function of the second image classification neural network; the first new label image classification loss sub-function is used to represent a loss value between a predicted classification result corresponding to a portrait region feature map of a second calibrated sample picture in the second target sample set and the first learning result; the second new label image classification loss sub-function is used to represent a loss value between a predicted classification result corresponding to a portrait region feature map of a first calibrated sample picture in the first target sample set and the first result.
[0130] The first target comprehensive loss function includes three sub-functions, namely a first original label image classification loss sub-function, a first label learning loss sub-function, and a first new label image classification loss sub-function.
[0131] The second target comprehensive loss function includes three sub-functions, namely a second original label image classification loss sub-function, a second label learning loss sub-function, and a second new label image classification loss sub-function.
[0132] Compared with the prior art, the significant advantages of the present invention are as follows: (1) Utilize the semantic information contained in label embedding to alleviate the negative impact of label noise on image classification; (2) The learning of label embedding and the training of the parameters of the image classification neural network can promote each other, and finally an accurate and stable classification effect can be obtained.
[0133] Different from the existing robust image classification methods, the present invention uses a learnable label embedding as the target to be fitted during the training of the image classification neural network, so as to utilize the semantic information contained in the label embedding to help the image classification neural network overcome the negative impact of label noise, and finally train a well-performing image classification neural network in a dataset with label noise. Experiments prove that the method proposed by the present invention can train an image classification neural network that is robust to label noise. This method can be applied to the image classification task with label noise and achieve satisfactory classification results.
[0134] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0135] In this article, specific examples are used to elaborate on the principles and implementation modes of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation modes and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.
Claims
1. A robust image classification method based on label embedding, characterized in that Including: Obtain a target image to be classified; The target image is an image containing a portrait area; Classify the target image to be classified according to a robust image classifier to obtain a classification result of the target image; Wherein, the robust image classifier is trained based on a target comprehensive loss function; The robust image classification method based on label embedding further includes a training method for the robust image classifier; the training method includes: Obtain training data; the training data includes a plurality of data pairs; each data pair includes a sample image and an original label classification result of the portrait area corresponding to the sample image; the original label classification result of the portrait area corresponding to the sample image contains noise information; Determine a target comprehensive loss function; Based on the training data, the target comprehensive loss function, and an image classification neural network, determine the final parameters of the image classification neural network; Based on the final parameters of the image classification neural network, determine a robust image classifier; The target comprehensive loss function includes a first target comprehensive loss function and a second target comprehensive loss function; determining the target comprehensive loss function specifically includes: Determine a first target sample set and a second target sample set; the data pairs in the first target sample set and the second target sample set are both part of the training data, and the data pairs in the first target sample set are different from the data pairs in the second target sample set; Determine a first original label image classification loss sub-function of a first image classification neural network and a second original label image classification loss sub-function of a second image classification neural network; both the first original label image classification loss sub-function and the second original label image classification loss sub-function are cross-entropy loss functions; the first original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the second calibrated sample image in the second target sample set and the original label classification result of the portrait area corresponding to the second calibrated sample image; the second original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait area feature map of the first calibrated sample image in the first target sample set and the original label classification result of the portrait area corresponding to the first calibrated sample image; the first calibrated sample image is any sample image in the first target sample set; the second calibrated sample image is any sample image in the second target sample set; Determine a first label learning loss sub-function of the first image classification neural network and a second label learning loss sub-function of the second image classification neural network; the first label learning loss sub-function is used to represent the loss value between the first learning result and the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture; the second label learning loss sub-function is used to represent the loss value between the second learning result and the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture; the first learning result is the learning result obtained by fitting and learning the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture with the original label classification result of the portrait region corresponding to the second calibrated sample picture; the second learning result is the learning result obtained by fitting and learning the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture with the original label classification result of the portrait region corresponding to the first calibrated sample picture; wherein, the first learning result is the new label classification result of the portrait region corresponding to the second calibrated sample picture, and the second learning result is the new label classification result of the portrait region corresponding to the first calibrated sample picture; Determine a first new label image classification loss sub-function of the first image classification neural network and a second new label image classification loss sub-function of the second image classification neural network; the first new label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture in the second target sample set and the first learning result; the second new label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture in the first target sample set and the first result; The first target comprehensive loss function includes three sub-functions, namely a first original label image classification loss sub-function, a first label learning loss sub-function, and a first new label image classification loss sub-function; The second target comprehensive loss function includes three sub-functions, namely a second original label image classification loss sub-function, a second label learning loss sub-function, and a second new label image classification loss sub-function.
2. The robust image classification method based on label embedding according to claim 1, wherein, The obtaining of the training data specifically includes: Obtain original data; the original data includes a plurality of original data pairs; the original data pair includes an original sample picture and the original label classification result of the portrait region corresponding to the original sample picture; the original label classification result of the portrait region corresponding to the original sample picture contains noise information; Perform normalization processing on the original data to obtain training data.
3. A robust image classification method based on label embedding according to claim 1, wherein Based on the training data, the target comprehensive loss function, and the image classification neural network, determine the final parameters of the image classification neural network, specifically including: Determine the final parameters of the first image classification neural network and the final parameters of the second image classification neural network based on the first target sample set, the second target sample set, the first target comprehensive loss function, the second target comprehensive loss function, the first image classification neural network, and the second image classification neural network; Among them, the robust image classifier is determined according to the final parameters of the first image classification neural network or according to the final parameters of the second image classification neural network.
4. A robust image classification method based on label embedding according to claim 1 or 3, characterized in that The first image classification neural network and the second image classification neural network are both the same image classification neural network; the image classification neural network is a network obtained by adding an embedding layer on the basis of a convolutional neural network; the embedding layer is used to convert the original label classification result of the portrait area into a new label classification result of the portrait area.
5. A robust image classification system based on label embedding, characterized in that, It includes: A data acquisition module, which is used to acquire the target picture to be classified; The target picture is a picture containing a portrait area; A classification result determination module, which is used to classify the target picture to be classified according to the robust image classifier to obtain the classification result of the target picture; Among them, the robust image classifier is trained based on the target comprehensive loss function; The robust image classification system based on label embedding further includes a training module; the training module includes: A training data acquisition unit, which is used to acquire training data; the training data includes multiple data pairs; the data pair includes a sample picture and the original label classification result of the portrait area corresponding to the sample picture; the original label classification result of the portrait area corresponding to the sample picture contains noise information; A target comprehensive loss function determination unit, which is used to determine the target comprehensive loss function; An image classification neural network final parameter determination unit, which is used to determine the final parameters of the image classification neural network based on the training data, the target comprehensive loss function, and the image classification neural network; A robust image classifier determination unit, which is used to determine the robust image classifier based on the final parameters of the image classification neural network; The target comprehensive loss function includes a first target comprehensive loss function and a second target comprehensive loss function; the target comprehensive loss function determination unit specifically includes: A target sample set subunit, which is used to determine a first target sample set and a second target sample set; the data pairs in the first target sample set and the second target sample set are both part of the training data, and the data pairs in the first target sample set are different from the data pairs in the second target sample set; The original label image classification loss sub-function determination subunit is used to determine the first original label image classification loss sub-function of the first image classification neural network and the second original label image classification loss sub-function of the second image classification neural network; both the first original label image classification loss sub-function and the second original label image classification loss sub-function are cross-entropy loss functions; the first original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture in the second target sample set and the original label classification result of the portrait region corresponding to the second calibrated sample picture; the second original label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture in the first target sample set and the original label classification result of the portrait region corresponding to the first calibrated sample picture; the first calibrated sample picture is any sample picture in the first target sample set; the second calibrated sample picture is any sample picture in the second target sample set; The label learning loss sub-function determination subunit is used to determine the first label learning loss sub-function of the first image classification neural network and the second label learning loss sub-function of the second image classification neural network; the first label learning loss sub-function is used to represent the loss value between the first learning result and the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture; the second label learning loss sub-function is used to represent the loss value between the second learning result and the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture; the first learning result is the learning result obtained after the original label classification result of the portrait region corresponding to the second calibrated sample picture fits and learns the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture; the second learning result is the learning result obtained after the original label classification result of the portrait region corresponding to the first calibrated sample picture fits and learns the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture; wherein, the first learning result is the new label classification result of the portrait region corresponding to the second calibrated sample picture, and the second learning result is the new label classification result of the portrait region corresponding to the first calibrated sample picture; The new label image classification loss sub-function determination subunit is used to determine the first new label image classification loss sub-function of the first image classification neural network and the second new label image classification loss sub-function of the second image classification neural network; the first new label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the second calibrated sample picture in the second target sample set and the first learning result; the second new label image classification loss sub-function is used to represent the loss value between the predicted classification result corresponding to the portrait region feature map of the first calibrated sample picture in the first target sample set and the first result; The first target comprehensive loss function includes three sub-functions, namely the first original label image classification loss sub-function, the first label learning loss sub-function, and the first new label image classification loss sub-function; The second target comprehensive loss function includes three sub-functions, namely the second original label image classification loss sub-function, the second label learning loss sub-function, and the second new label image classification loss sub-function.
6. A robust image classification system based on label embedding according to claim 5, characterized in that, The training data acquisition unit specifically includes: The original data sub-unit is used to acquire original data; the original data includes a plurality of original data pairs; the original data pair includes an original sample picture and the original label classification result of the portrait area corresponding to the original sample picture; the original label classification result of the portrait area corresponding to the original sample picture contains noise information; The normalization processing sub-unit is used to perform normalization processing on the original data to obtain training data.
Citation Information
Patent Citations
Image classification method based on semi-supervision
CN112200245A
Image classification method and device, electronic equipment and storage medium
CN113033689A