Robust image classification method for suppressing noise tag interference

Through dynamic threshold and weighting strategies, and memory compensation is implemented later in the training, the problem of degraded generalization performance of deep neural networks on noise-labeled datasets is solved, achieving higher robustness and generalization capabilities.

CN120014322APending Publication Date: 2025-05-16XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510011520.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-04
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When processing data sets with noise labels, the prior art tends to overfit the noise labels, resulting in a degradation of the generalization performance of the network and it is difficult to effectively utilize the potential value of clean and difficult samples.

Method used

A weighted memory compensation algorithm (WMCDH) based on dynamic thresholds is proposed. By dynamically adjusting the threshold and sample weights, clean labels, difficult labels and noise labels are separated, and a memory compensation mechanism is implemented later in the training stage to strengthen the network's learning of high-trust clean samples.

Benefits of technology

Effectively suppress noise label interference, improve the robustness and generalization capabilities of the model, make full use of the information of clean and difficult samples, and improve model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014322A_ABST
    Figure CN120014322A_ABST
Patent Text Reader

Abstract

The invention discloses a robust image classification method for suppressing noise label interference, and belongs to the field of computer vision and artificial intelligence. The method comprises the following steps: S1, inputting a data set with a noise tag into a backbone network to carry out preheating training of a plurality of epochs, and recording the loss of all samples in the preheating training; s2, calculating the average loss of each sample according to the loss recorded in S1; s3, calculating a dynamic threshold value, and dividing the sample into three areas with clean labels, noise labels and difficult labels in combination with the distribution of the average loss obtained in the S2; s4, assigning different weights to the three areas with the clean labels, the labels with the noise and the difficult labels obtained through division in the step S3, and calculating composite weighting loss; s5, designing a memory compensation mechanism, strengthening the learning of the network on a high-credibility clean sample in the later stage of training, compensating knowledge forgetting caused by noise, inhibiting the interference of a noise label sample, and finally outputting a trained network model; and S6, dynamically updating the loss record and the average loss of the sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and artificial intelligence, and in particular to a robust image classification method for suppressing noise label interference by mining clean difficult samples. Background Art

[0002] Deep Neural Networks (DNNs) have achieved great success in dealing with various tasks. However, the success of DNNs depends largely on large-scale data sets with high-quality annotations. In many practical application scenarios, it is very difficult to obtain large-scale and accurately annotated data sets. Data collected by crowdsourcing, automatic annotation, web crawlers and other technologies will inevitably contain noisy labels. Since deep neural networks have strong learning capabilities, previous studies have shown that DNNs are prone to overfitting these erroneous labels on data sets with noisy labels, resulting in a decrease in the generalization performance of the network. Since label noise is widely present in many high-end and precise practical application scenarios, such as transportation, remote sensing, medical care and agriculture. Therefore, it is of far-reaching significance to propose a learning method that is robust to noisy labels for the problem of label noise for practical scene applications. The present invention belongs to the field of computer vision and efficient and reliable artificial intelligence, and is applied to image intelligent processing tasks in visual task scenarios such as general images, medical images, traffic images and remote sensing images.

[0003] In noisy label learning, in order to suppress the interference of noisy labels and improve the robust performance of the network, the main strategies used include designing robust loss functions, calculating noise transfer matrices, label correction, early stopping and sample selection. In particular, compared with other methods, sample selection methods have shown very promising results in dealing with noisy label problems. The core idea of ​​sample selection methods is to screen samples that may have clean labels according to the predicted probability or loss value of the samples during training, and train the network based on these clean label samples. Malach et al. select samples based on the prediction difference between two classifiers. Jiang et al. use clean samples to build a pre-trained mentor network to guide the training of the student network. Han et al. train two networks simultaneously and select samples with loss values ​​below a threshold to train the other network. Yu et al. use the same threshold to select data for another model based on small loss samples with inconsistent predictions. Bai et al. use momentum to refine the classifier alternately to extract samples. Han et al. use the forgetting phenomenon of the model to identify and filter out samples that may have noisy labels. Wei et al. selected samples with smaller losses by calculating the joint loss and simultaneously updated the parameters of the two networks. Li et al. used two networks to select samples from each other in combination with semi-supervised learning techniques. These algorithms generally use small loss samples as clean samples to ensure that the screened samples have high credibility. However, in order to ensure that the selected samples are clean enough, they tend to ignore some clean and difficult samples with larger losses. Although these CHSs have larger losses, they usually contain more representative feature information and may be close to the classification boundary or belong to complex categories. Therefore, excluding them from training is likely to lead to the loss of valuable information and may aggravate the class imbalance problem. This limitation restricts the generalization performance of the network and its adaptability to challenging tasks.

[0004] Obviously, effectively mining clean and difficult samples in the LNL task is of great significance to improving the robustness of learning. Exploring how to make full use of the potential value of such samples will provide new solutions for improving model performance. In order to learn more clean and difficult samples during training, Zhu et al. proposed an HSA-NRL difficult sample perception algorithm, but the model paid too much attention to the processing of difficult samples in the early stage and easily neglected the learning of simple clean samples, resulting in overfitting or unstable learning. Yuan et al. proposed a Late Stopping algorithm, which monitors the behavior of the model by observing the changes in model confidence over time. Although Late Stopping has noticed the importance of difficult samples, it directly merges all difficult samples into the clean sample set, which easily leads to too many noisy difficult samples being retained, thus affecting the learning performance of the model. Cordeiro et al. proposed a PropMix algorithm to filter difficult samples through a confidence-based screening mechanism. Although this method can separate samples that may contain noise, it regards these samples as unlabeled data and updates the labels, but fails to directly select actually valuable samples from the difficult label sample area, which may not fully utilize the potential information of CHSs. Existing methods have certain limitations in solving the CHSs problem, especially in accurately identifying and effectively utilizing clean and difficult samples. Therefore, further exploring effective mining and utilization strategies for CHSs is an important direction to improve the robustness of the LNL model.

[0005] In the field of noisy label learning (LNL), many existing heuristic methods are based on the memory effect of deep neural networks. Zhang et al. showed that deep neural networks have strong learning ability and can gradually remember samples of the entire data set. Arpit et al. further found that deep neural networks first learn training samples with clean labels and then learn training samples with noisy labels. These studies reveal the reason why the generalization performance of deep neural networks decreases in noisy label learning, that is, as the training process deepens, the network gradually overfits the noisy label samples, weakening the retention of the correct memory learned earlier. Since samples with noisy labels are usually learned in the later stage of training, if a memory compensation strategy is adopted at this stage, it can effectively reduce the forgetting of correct memory and thus improve the robustness of the model. The core idea of ​​memory compensation is to re-strengthen the network's memory by using highly credible clean samples in the later stage of training so that it can recover from the interference of noisy samples. This strategy provides a new idea for further improving the network's ability to handle noisy labeled data, and further proves the important application value of memory effect in LNL.

[0006] The sample selection method has achieved good performance in training with noisy labels. In order to reduce the interference of noisy label samples, the sample selection algorithm often restricts the network to learn samples with large loss in clean label samples. Figure 1 As shown in (a), the clean samples with large losses in the blue dotted box are called clean hard samples (CHSs). These clean hard samples are often close to the decision boundary of the classifier. CHSs and some noise label samples often have the same loss, which causes them to overlap and cannot be distinguished. In the learning task with noisy labels, the identification of CHSs is quite challenging. Since these samples are fuzzy in the overall classification, they are easily misclassified. In order to keep the training samples clean, most of the existing noisy label learning methods try to eliminate the noisy label samples with large losses as much as possible, which inevitably leads to the inclusion of many clean hard samples. However, these CHSs are more representative of the category in terms of feature information, and it is necessary to consider their positive impact on learning, because this is crucial to achieve the best generalization performance. Therefore, when designing the LNL algorithm, how to effectively identify and reasonably utilize these CHSs becomes one of the key issues to improve the robustness and performance of the network. Summary of the invention

[0007] The purpose of the present invention is to propose a robust image classification method that suppresses noise label interference to solve the problems mentioned in the background technology.

[0008] To solve the above problems, the present invention specifically adopts the following technical solutions:

[0009] A robust image classification method for suppressing noise label interference comprises the following steps:

[0010] S1. Input the dataset with noise labels into the backbone network for several epochs of warm-up training, and record the loss of all samples in the warm-up training.

[0011] S2, calculate the average loss of each sample based on the loss recorded in S1;

[0012] S3, calculate the dynamic threshold, combine the distribution of the average loss obtained in S2, and divide the samples into three areas: those with clean labels, those with noisy labels, and those with difficult labels;

[0013] S4, assign different weights to the three regions with clean labels, noisy labels and difficult labels obtained in S3, and calculate the composite weighted loss;

[0014] S5. Design a memory compensation mechanism to strengthen the network’s learning of highly reliable clean samples in the later stages of training, compensate for knowledge forgetting caused by noise, suppress the interference of noisy label samples, and finally output the trained network model;

[0015] S6. Dynamically update the loss record and average loss of the sample.

[0016] Preferably, S2 specifically includes the following contents:

[0017] Define the historical loss record of each sample i as R i ={L i,1 ,L i,2 ,…,L i,t}, then the average loss is defined as:

[0018]

[0019] At the end of each epoch, the latest loss is recorded in the historical loss:

[0020] R i ←R i ∪{L i,t+1}

[0021] Among them, t represents the epoch number of the current training.

[0022] Preferably, the dynamic threshold in S3 includes a cleaning threshold T clean and noise threshold T noise , and its setting rules are as follows:

[0023]

[0024]

[0025] in, Indicates sorting the average loss of samples; η is a hyperparameter; γ represents the noise rate; N represents the total number of samples; T represents the total training cycle;

[0026] As the training progresses, Increase, the noise threshold T noise Gradually decrease to reduce the influence of noise samples.

[0027] Preferably, the division rule of combining the dynamic threshold and the average loss distribution in S3 to divide the samples into three regions of clean labels, noisy labels and difficult labels is:

[0028] Clean label area:

[0029] Region with noisy labels:

[0030] Difficulty labeling area:

[0031] Preferably, the calculation formula of the composite weighted loss in S4 is as follows:

[0032]

[0033] Among them, L base (p i ,y i ) represents the cross entropy loss of the i-th sample; wi represents the weight of the sample, which is dynamically adjusted based on its loss size;

[0034] The weight calculation rules for samples in different regions are as follows:

[0035] Clean area sample weight: w i =w clean , when the sample loss is less than the cleaning threshold T clean When , give it the maximum weight w clean =1;

[0036] Noise area sample weight: w i =w noise , when the sample loss is greater than the noise threshold T noise When , give it the minimum weight w clean =0;

[0037] Overlap Area Domain sample weight: w i ∈[w noise ,w clean ], for samples between the clean threshold and the noise threshold, their weights are calculated based on the sample and T clean The distance between them decreases linearly according to the following formula:

[0038]

[0039] Among them, L i is the loss of the ith sample.

[0040] Preferably, the memory compensation mechanism in S5 specifically includes the following contents:

[0041] When T clean =T noise , it enters the memory compensation stage; in the memory compensation stage, only highly reliable clean samples are selected for training:

[0042]

[0043] Preferably, at the end of each epoch, the latest loss is recorded into the historical loss:

[0044] R i ←R i ∪{L i,t+1}

[0045] Recalculate the average loss

[0046]

[0047] Compared with the prior art, the present invention provides a robust image classification method for suppressing noise label interference, which has the following beneficial effects:

[0048] (1) The present invention proposes a robust image classification method that suppresses noise label interference, and further proposes a weighted memory compensation algorithm based on dynamic threshold (WMCDH), which optimizes model performance by dynamically adjusting the threshold, updating the weights of samples according to different noise rates and training epochs, and implementing a memory compensation strategy.

[0049] (2) The present invention dynamically assigns different weights to difficult samples according to the distance between the difficult samples and the boundary of the clean sample set, thereby more effectively utilizing the clean labeled samples in the difficult sample set. This strategy is more advantageous than simply selecting clean samples.

[0050] (3) Based on the memory effect of deep neural networks, the present invention proposes a memory compensation mechanism, which further suppresses the influence of noisy labels in difficult samples and improves the generalization ability of the model. According to current understanding, this is the first time that the concept of memory compensation has been introduced in the field of noisy label learning, filling this research gap. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0052] Figure 1 is the average loss distribution of clean label samples and noisy label samples on the CIFAR-10 dataset containing 50% symmetric noise;

[0053] Figure 2 It is a block diagram of the WMCDH network structure mentioned in the embodiment of the present invention;

[0054] Figure 3 This is a schematic diagram of the dynamic sample selection mentioned in the embodiment of the present invention. DETAILED DESCRIPTION

[0055] The technical solution and its effects of the present invention are further described below in conjunction with specific implementation modes / examples. The following implementation modes / examples are only used to illustrate the content of the present invention, and the present invention is not limited to the following implementation modes or examples.

[0056] First, the problem solved by the present invention is described:

[0057] Dataset Contains K categories and N samples, is a sample in the d-dimensional data space, y i ∈y={1,…,K} is the true label corresponding to xi, and (xi,yi) satisfies the independent and identical step condition. The classification task is to obtain a mapping function f(•;Θ) from x→y such that the empirical risk R of parameter Θ D (f) Minimization.

[0058]

[0059] Where L represents the loss function.

[0060] As the data set continues to grow, various noises will inevitably be introduced into the data set during the data collection process, which will convert the originally clean data set D into a data set with noisy labels. In Represents a label that may be noisy. Performing standard classification training on the , we will get a small batch B at a certain moment t The formula is:

[0061]

[0062] Where b is B t The number of samples included.

[0063] In B t Up along R D (f) The direction that minimizes Θ t The parameters are updated:

[0064]

[0065] Where η is the learning rate.

[0066] From formula (3), we can see that when training with dataset D~, the network is no longer noise-tolerant and is prone to overfitting to noise labels, resulting in reduced generalization performance. Therefore, it is of great significance to minimize the impact of noise labels on training.

[0067] In order to solve the above problems, the model can learn as many useful clean samples as possible during training, especially clean difficult samples. The present invention proposes a weighted memory compensation algorithm (WMCDH) based on dynamic threshold. Specifically, WMCDH dynamically divides the training samples into clean label samples, difficult label samples and noise label samples (such as Figure 1 (b)). WMCDH focuses on selecting clean labeled samples with small losses in the early stage and retaining as many CHSs as possible in the difficult sample set. For difficult samples, WMCDH dynamically assigns different weights according to their distance from the boundary of the clean labeled sample set. That is, the distance from the clean threshold T clean The closer the difficult sample is, the larger its weight is, and vice versa. As the training progresses, WMCDH slides to reduce the noise threshold T noise , to avoid fitting too many noise label samples. However, when learning clean and difficult samples, it is inevitable to introduce certain noise label samples. According to the memory effect of deep neural networks, overfitting these noise label samples may cause the correct information learned previously to be forgotten. To solve this problem, we introduced a "memory compensation" mechanism. That is, in the later stage of training, let the noise threshold T noise Reduced to less than the clean threshold T clean A small number of high-confidence clean samples are used to compensate the network’s memory, thereby preventing the network from forgetting useful memories and further suppressing the interference of noise label samples.

[0068] The following is an explanation of a robust image classification method for suppressing noise label interference proposed by the present invention in conjunction with relevant drawings and specific examples. The specific contents are as follows.

[0069] Embodiment 1:

[0070] See also Figure 2 , the network structure diagram of WMCDH is as follows Figure 2As shown. The present invention first inputs the data set with noisy labels into the backbone network for preheating training for several epochs, and records the losses of all samples in the preheating training. The average loss of each sample in all epochs is calculated based on the recorded losses. According to the average loss distribution of the samples, the samples are divided into three regions with clean labels, noisy labels and difficult labels using a dynamic threshold. The difficult label region is the overlapping area of ​​the clean label samples and the noisy label samples. Therefore, the difficult label region contains some clean label samples with larger losses and some noisy label samples with smaller losses. In each training epoch, for different regions, we assign different weights to samples and obtain a composite loss function. After an epoch ends, the loss of each sample is added to the historical record of the loss as a basis for calculating the average loss distribution in the next epoch. In particular, in each epoch, the present invention dynamically changes the threshold T2. When T clean <T noise When T , we use the distance weight method to dynamically modify the weight for samples in difficult label areas. clean >T noise When training the network, only some clean label samples are selected to compensate for the impact of noisy labels on network performance and further improve the robust performance of the network. The specific details are as follows:

[0071] (1) Dynamic Threshold

[0072] Threshold T clean and T noise The dynamic adjustment of is the core of this method, which aims to gradually increase the participation ratio of clean samples and reduce the interference of noise samples. clean and T noise It can more accurately classify clean samples and noise samples. The present invention records the historical losses of each sample in the previous rounds of warm-up training and calculates their mean. This can smooth out single loss anomalies caused by training fluctuations, thereby more stably judging sample quality.

[0073] Define the historical loss record of each sample i as R i ={L i,1 ,L i,2 ,…,L i,t}, where t is the number of epochs in the current training. Its mean is defined as:

[0074]

[0075] At the end of each epoch, the present invention records the latest loss into the historical loss:

[0076] R i ←R i ∪{Li,t+1}(5)

[0077] The average loss strategy can dynamically adjust T clean and T noise , which makes sample classification more stable and helps to make better use of training data. The threshold setting rules are as follows:

[0078] Cleaning threshold T clean :

[0079]

[0080] in It means sorting the mean loss of samples, η is a hyperparameter; γ represents the noise rate; N represents the total number of samples.

[0081] Noise threshold T noise :

[0082]

[0083] Where T represents the total training cycle.

[0084] As the training progresses, Increase, the noise threshold T noise Gradually decrease to reduce the influence of noise samples.

[0085] According to the threshold T clean and T noise , the sample can be divided into the following three areas:

[0086] Cleaning samples:

[0087] Noise sample:

[0088] Overlapping (difficult) samples:

[0089] This dynamic threshold adjustment method can gradually optimize the sample classification strategy, allowing clean samples to participate in training to a greater extent and improve model performance.

[0090] The steps of the WMCDH algorithm are as follows:

[0091]

[0092]

[0093] (2) Weighted loss

[0094] In order to effectively utilize the clean labeled samples in the clean sample area and the overlapping sample area, and suppress the interference of the noisy labeled samples in the noisy sample area, the present invention proposes a loss-based weighting strategy. The weighted loss calculation formula is as follows:

[0095]

[0096] Among them, L base (p i ,y i ) is the cross entropy loss of the ith sample, wi is the weight of the sample, which is dynamically adjusted based on its loss size.

[0097] The weight calculation rules for samples in different regions are as follows:

[0098] Clean area sample weight (w i =w clean ): When the sample loss is less than the cleaning threshold T clean When , give it the maximum weight w clean =1.

[0099] Noise area sample weight (w i =w noise ): When the sample loss is greater than the noise threshold T noise When , give it the minimum weight w clean =0.

[0100] Overlap Area Domain sample weight (w i ∈[w noise ,w clean ]): For samples between the clean threshold and the noise threshold, their weights are calculated based on the sample and T clean The distance between them decreases linearly according to the following formula:

[0101]

[0102] Where L i is the loss of the ith sample.

[0103] The distribution of clean label samples and noisy label samples in the overlapping area is uneven. Figure 3 As shown, the closer to the threshold T clean The higher the proportion of clean labels, the higher the probability of noisy labels. This method dynamically allocates the weights of samples in overlapping areas, which not only enhances the learning of clean label samples in overlapping sample areas, but also reduces the impact of noisy label samples.

[0104] (3) Memory compensation

[0105] Although the weighting strategy reduces the impact of noise label samples in the overlapping sample area, it is still inevitably affected to a certain extent. According to the memory effect of deep neural networks, overfitting these noise label samples will cause the network to gradually forget the correct knowledge learned earlier. Therefore, in order to reduce the negative impact of noise label samples, we proposed a memory compensation mechanism: in the later stage of training, the network's learning of a small number of highly reliable clean samples is re-strengthened to compensate for the knowledge forgotten due to noise.

[0106] When T clean =T noise When , it enters the memory compensation stage. In the memory compensation stage, only highly reliable clean samples are selected for training:

[0107]

[0108] In the memory compensation stage, continue to reduce the threshold T noise , so that T clean >T noise ,like Figure 3 (d) This adjustment allows the model to pay more attention to high-confidence clean samples with smaller losses at the end, while suppressing the fitting of noise samples with larger losses. This is because memory compensation makes full use of the memory effect of the deep learning model, strengthens the model's learning ability for clean samples, and thus improves the training effect.

[0109] (4) Dataset

[0110] The effectiveness of the proposed method is verified on four artificial noise datasets: MNIST, F-MNIST, CIFAR-10 and CIFAR-100, and three real noise datasets: Food-101, Clothing1M and WebVision. These datasets are widely used in the test of noisy labeled image classification and can be regarded as benchmark datasets.

[0111] MNIST. This dataset contains 70,000 images of handwritten digits, each of which is 28×28 pixels in size and is grayscale. The images are divided into 10 categories, each containing about 7,000 images. 6,000 images are used for training and 1,000 images are used for testing. The entire dataset is divided into a training set of 60,000 images and a test set of 10,000 images.

[0112] F-MNIST. This dataset is a popular benchmark dataset for image classification. The dataset contains 70,000 grayscale images of clothing items, each of which is 28×28 pixels in size. The data structure of F-MNIST is similar to MNIST, dividing the images into 10 categories, each containing about 7,000 images. 60,000 images are used for training and 10,000 images are used for testing. It is widely used in research in the fields of neural networks and computer vision.

[0113] CIFAR. This dataset contains two datasets, CIFAR-10 and CIFAR-100. Both datasets contain 60,000 long color images, each of which is 32×32 pixels in size. CIFAR-10 divides images into 10 categories, each of which contains 6,000 images, of which 5,000 are used for training and 1,000 are used for testing. CIFAR-100 divides images into 100 categories, each of which contains 600 images, of which 500 are used for training and 100 are used for testing. Both datasets contain a training set of 50,000 images and a test set of 10,000 images.

[0114] Food-101. This dataset contains 101 different types of food categories, totaling about 100,000 images, each of which is 112×112 pixels in size. These images are collected from the Internet and cover common types of food, such as fruits, vegetables, staple foods, desserts, etc. For each category, 250 manually reviewed clean test images and 750 training images with real-world label noise are provided. There are about 1,000 high-quality images for each category, 750 of which are used for training and 250 for testing. The entire dataset is divided into a training set of 75,750 images and a test set of 25,250 images.

[0115] Clothing1M. This dataset is a large-scale natural scene dataset with noisy labels. The images come from online shopping websites. Since the sample labels are generated by keywords, label noise is easily introduced during the collection process. The dataset is divided into 14 categories, containing 1 million noisy data, and 50,000, 14,000, and 10,000 clean data for training, verification, and testing respectively.

[0116] WebVision. This dataset is a large-scale dataset for training image classification models. It uses 2.4 million images crawled from websites using 1,000 concepts in ImageNet ILSVRC12, with each category containing approximately 2,400 to 3,600 images. WebVision has a similar category setting to ImageNet, covering 1,000 categories and sharing the same category labels as the ImageNet dataset. Since the images are automatically crawled from the Internet, the dataset contains a large number of noisy labels, which is suitable for research tasks such as learning with noisy labels.

[0117] (5) Evaluation comparison

[0118] The present invention compares the test accuracy of the proposed method with several state-of-the-art methods. Compared with many state-of-the-art methods, the proposed method achieves better results on both artificially synthesized noise datasets and real noise datasets. On datasets containing artificially synthesized noise, the present invention adds four types of artificially synthesized noise, namely symmetric noise, pairflip noise, triangular noise and instance-dependent noise, to each dataset. Some experimental results are shown in Tables 1 and 2.

[0119] Table 1 Test accuracy on the cifar10 dataset containing four types of artificial noise

[0120]

[0121] Table 2 Test accuracy on the Food-101 dataset containing real noise

[0122]

[0123] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and improved concepts of the present invention within the technical scope disclosed by the present invention, and they should be covered by the protection scope of the present invention.

Claims

1. A robust image classification method for suppressing noise label interference, characterized in that: The following steps are involved: S1. Input the dataset with noise labels into the backbone network for several epochs of warm-up training, and record the loss of all samples in the warm-up training. S2, calculate the average loss of each sample based on the loss recorded in S1; S3, calculate the dynamic threshold, combine the distribution of the average loss obtained in S2, and divide the samples into three areas: those with clean labels, those with noisy labels, and those with difficult labels; S4, assign different weights to the three regions with clean labels, noisy labels and difficult labels obtained in S3, and calculate the composite weighted loss; S5. Design a memory compensation mechanism to strengthen the network’s learning of highly reliable clean samples in the later stages of training, compensate for knowledge forgetting caused by noise, suppress the interference of noisy label samples, and finally output the trained network model; S6. Dynamically update the loss records and average loss of samples.

2. The method according to claim 1, characterized in that The S2 specifically includes the following contents: Define the historical loss record of each sample i as R i ={L i,1 ,L i,2 ,…,L i,t }, then the average loss is defined as: At the end of each epoch, the latest loss is recorded in the historical loss: R i ←R i ∪{L i,t+1 } Among them, t represents the epoch number of the current training.

3. The method according to claim 2, characterized in that The dynamic threshold in S3 includes a cleaning threshold T clean and noise threshold T noise , and its setting rules are as follows: in, Indicates sorting the average loss of samples; η is a hyperparameter; γ represents the noise rate; N represents the total number of samples; T represents the total training cycle; As the training progresses, Increase, the noise threshold T noise Gradually decrease to reduce the influence of noise samples.

4. The method according to claim 3, characterized in that The rule described in S3 for combining the dynamic threshold with the distribution of the average loss to divide the samples into three regions: those with clean labels, those with noisy labels, and those with difficult labels is: Clean label area: Region with noisy labels: Difficulty labeling area:

5. The method according to claim 4, characterized in that The calculation formula of the composite weighted loss described in S4 is as follows: Among them, L base (p i ,y i ) represents the cross entropy loss of the i-th sample; wi represents the weight of the sample, which is dynamically adjusted based on its loss size; The weight calculation rules for samples in different regions are as follows: Clean area sample weight: w i =w clean , when the sample loss is less than the cleaning threshold T clean When , give it the maximum weight w clean =1; Noise area sample weight: w i =w noise , when the sample loss is greater than the noise threshold T noise When , give it the minimum weight w clean =0; Overlap Area Domain sample weight: w i ∈[w noise ,w clean ], for samples between the clean threshold and the noise threshold, their weights are calculated based on the sample and T clean The distance between them decreases linearly according to the following formula: Among them, L i is the loss of the ith sample.

6. The method according to claim 5, characterized in that The memory compensation mechanism described in S5 specifically includes the following contents: When T clean =T noise , it enters the memory compensation stage; in the memory compensation stage, only highly reliable clean samples are selected for training:

7. The method according to claim 6, characterized in that At the end of each epoch, the latest loss is recorded in the historical loss: R i ←R i ∪{L i,t+1 } Recalculate the average loss: