Sample selection method based on model prediction confidence coefficient change trend

By analyzing the confidence difference time series of deep learning models in training iterations and judging the confidence change trend, the problem of difficulty in balancing accuracy and recall in the current technology during sample selection is solved, and more accurate sample selection is achieved.

CN120045940APending Publication Date: 2025-05-27HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510124267.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art has difficulty in balancing accuracy and recall when selecting samples, resulting in the correctly marked samples that may be missed or the wrongly marked samples are selected.

Method used

The sample selection method based on the model predicts the confidence change trend, and the confidence change trend is judged by collecting and analyzing the confidence difference time series of the model in different training iterations, and the Mann-Kendall trend test method is used to accurately distinguish the correct labels and wrong labels in high-loss data.

Benefits of technology

It effectively alleviates the problem of traditional small loss strategies ignoring the correct labels in high-loss samples, and improves the accuracy and recall rate of sample selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045940A_ABST
    Figure CN120045940A_ABST
Patent Text Reader

Abstract

The invention discloses a sample selection method based on a model prediction confidence coefficient change trend. The method comprises the following steps: 1) obtaining all to-be-selected sample data; 2) forming a k classification data set containing noise labels for the acquired sample data; 3) for each sample in the training set, collecting the following confidence difference: 4) collecting the confidence difference in different training iterations to obtain the following confidence difference time sequence: 5) judging the confidence change trend of the confidence difference time sequence, and if the confidence difference time sequence has the rising trend, judging that the confidence difference time sequence does not have the rising trend, and if the confidence difference time sequence does not have the rising trend, judging that the confidence difference time sequence does not have the rising trend; and if not, the sample is regarded as a potential correct labeled sample. By using the method provided by the invention, the correct label can be identified from the high-loss sample, so that the problem that the correct label in the high-loss sample is ignored during sample selection of a traditional small-loss strategy is relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to deep learning technology, and in particular to a sample selection method based on the change trend of model prediction confidence. Background Art

[0002] The impact of noisy labels in training data on deep learning models cannot be ignored. Modern deep neural networks can often easily fit the wrong labels in the dataset, learn the wrong mapping relationship between sample features and categories, and affect the generalization ability of the model. To alleviate the adverse effects of noisy labels on model training, researchers have developed a large number of sample selection methods, aiming to select correctly labeled samples from the training set, regard the remaining mislabeled samples as unlabeled samples, and finally use semi-supervised learning technology for model training.

[0003] Common sample selection methods usually regard samples with smaller losses as correctly labeled samples, that is, adopt the small-loss trick. However, some correctly labeled samples have inherent learning difficulties, resulting in a high-loss phenomenon similar to that of mislabeled samples in the early stage of training. Therefore, it is usually difficult to achieve a balance between accuracy and recall by simply setting a threshold based on the sample loss size, that is: a lower threshold may lead to the omission of a large number of correctly labeled samples that are difficult to learn, while a higher threshold may lead to the selection of a large number of mislabeled samples. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a sample selection method based on the change trend of model prediction confidence for the defects in the prior art.

[0005] The technical solution adopted by the present invention to solve its technical problems is: a sample selection method based on the change trend of model prediction confidence, including the following steps:

[0006] 1) Obtain all the sample data to be selected;

[0007] 2) For the obtained sample data, form a k-classification dataset containing noisy labels, denoted as where represents the feature of the i-th sample in the dataset, is the observed label of the i-th sample;

[0008] is the sample feature space, and the label space is a discrete label set; denote y i as the true label of the sample x i ;

[0009] 3) For each sample in the noisy training set Collect the following confidence gaps:

[0010]

[0011] wherein represents the confidence difference between the model prediction on its labeled label i for the sample \(x\) in the \(t\)-th training iteration and the confidence on the class \(c\),

[0012] 4) After collecting these confidence gaps in different training iterations, the following confidence difference time series is obtained:

[0013]

[0014] 5) Judge the confidence change trend of the confidence difference time series. If the confidence gap time series has an upward trend, that is, the confidence gap between the labeled label and other label classes tends to increase with the training iteration, then regard this sample as a potentially correctly labeled sample.

[0015] According to the above scheme, in step 5), the Mann-Kendall trend test method is used to judge the confidence change trend of the confidence difference time series.

[0016] According to the above scheme, in step 5), the Mann-Kendall trend test method is used to judge the confidence change trend of the confidence difference time series, specifically as follows:

[0017] Use \(MKTest(\cdot)\) to represent the Mann-Kendall test process, which takes the confidence difference time series \(D\) c i (\(t\)) as the input and outputs the standardized test statistic \(Z\). The alternative hypothesis is that the confidence gap time series has an upward trend;

[0018] The sample selection criterion is:

[0019]

[0020] wherein, \(\alpha\) is the selected significance level, and \(Z\) 1-α is the \(100\times(1 - \alpha)\) percentile of the standard normal distribution. If a sample satisfies the sample selection criterion in the \(t\)-th iteration, it indicates that all confidence gaps have an upward trend, and it is regarded as a correctly labeled sample. \(C\) t is the set of correctly labeled samples selected in the \(t\)-th iteration.

[0021] According to the above scheme, the setting range of \(t\) is \([30, 100]\).

[0022] A sample selection method based on the changing trend of model prediction confidence, comprising the following steps:

[0023] 1) Obtain all the sample data to be selected, and use the small-loss sample selection strategy to select the correct label;

[0024] 2) For all the sample data to be selected, select the correct label based on the sample selection method of the changing trend of model prediction confidence;

[0025] 2.1) Form a k-classification data set containing noisy labels, denoted as where represents the feature of the i-th sample in the data set, is the observed label of the i-th sample;

[0026] is the sample feature space, and the label space is the discrete label set; denote y i as the true label of the sample x i ;

[0027] 2.2) For each sample in the noisy training set Collect the following confidence gaps:

[0028]

[0029] In the formula, represents the confidence difference between the model prediction on its labeled label i and the confidence on the category c for the sample x in the t-th training iteration,

[0030] 2.3) After collecting these confidence gaps in different training iterations, obtain the following confidence difference time series:

[0031]

[0032] 2.4) Judge the changing trend of the confidence of the confidence difference time series. If the confidence gap time series has an upward trend, that is, the confidence gap between the labeled label and other label categories tends to increase with the training iteration, then regard this sample as a potentially correctly labeled sample.

[0033] 3) Use the union of the samples selected in steps 1) and 2) as the new training sample set, and continue training to obtain the correctly labeled sample set.

[0034] According to the above scheme, in step 1), the small-loss sample selection strategy includes the sample selection method based on model prediction and the sample selection method based on training dynamics.

[0035] According to the above solution, in step 3), the loss function used in training is as follows:

[0036]

[0037] where is the cross-entropy loss function;

[0038]

[0039] represents the expectation of the sample in the data distribution D , is the indicator function.

[0040] The beneficial effects produced by the present invention are as follows:

[0041] The present invention proposes a sample selection method based on confidence trend tracking, which can accurately distinguish the correct labels and wrong labels in high-loss data by monitoring the change trend of the model prediction confidence. Even if the loss of correctly labeled data is large, the method proposed by the present invention can identify the correct labels from high-loss samples, thus alleviating the problem that the traditional small-loss strategy ignores the correct labels in high-loss samples during sample selection. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0043] Figure 1 is the flowchart of the method of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0045] As Figure 1 shown, a sample selection method based on the change trend of model prediction confidence includes the following steps:

[0046] 1) Obtain all the sample data to be selected;

[0047] 2) For the obtained sample data, form a k-classification data set containing noisy labels, denoted as where represents the feature of the i-th sample in the data set, is the observed label of the i-th sample;

[0048] is the sample feature space, and the label space is a set of discrete labels; denote y i as the true label of sample x i .

[0049] 3) For each sample in the noisy training set collect the following confidence gaps:

[0050]

[0051] where represents the confidence difference between the model prediction on its labeled label i for sample x and the confidence on class c,

[0052] 4) After collecting these confidence gaps in different training iterations, the following confidence difference time series is obtained:

[0053]

[0054] The setting range of the training iteration t is generally selected as the initial stage of training. In this embodiment, it is set according to the number of iterations in the initial stage of training, and the specific value can also be set according to the size of the dataset. In this embodiment, it is set to [30, 100].

[0055] 5) Judge the confidence change trend of the confidence difference time series. If the confidence gap time series has an upward trend, that is, the confidence gap between the labeled label and other label categories tends to increase with the training iteration, then the confidence of the labeled label should be the fastest growing. We regard such samples as potentially correctly labeled samples.

[0056] In step 5), the Mann-Kendall trend test method is used to judge the confidence change trend of the confidence difference time series, specifically as follows:

[0057] Use MKTest(·) to represent the Mann-Kendall test process, which takes the confidence difference time series as the input and outputs the standardized test statistic Z. The alternative hypothesis is that the confidence gap time series has an upward trend;

[0058] Formally, our sample selection criterion is:

[0059]

[0060] where α is the selected significance level, and Z 1-α is the 100×(1 - α) percentile of the standard normal distribution. If then it means The probability of no trend is less than α. Therefore, we accept the alternative hypothesis, that is, we believe that there is an upward trend. Therefore, if a sample satisfies the sample selection criterion in the t-th iteration, it indicates that all confidence gaps have an upward trend, and we regard it as a potentially correct label sample, C t is the set of possible correctly labeled samples selected in the t-th iteration.

[0061] A sample selection method based on the changing trend of model prediction confidence, comprising the following steps:

[0062] 1) Obtain all samples data to be selected, and use the small-loss sample selection strategy to select the correct labels;

[0063] Form a k-classification data set containing noisy labels, denoted as where represents the feature of the i-th sample in the data set, is the observed label of the i-th sample;

[0064] is the sample feature space, and the label space is the discrete label set; denote y i as the true label of the sample x i ;

[0065] Without using any noisy label learning strategy, use the ordinary cross-entropy to train the model:

[0066] Use the cross-entropy as the loss function. Denote as the logit output by the model when accepting the sample x i as the input, as the logit output by the model for the class c. The model prediction is the probability distribution This probability distribution is defined by the softmax function:

[0067]

[0068] The formal expression of the cross-entropy loss is as follows:

[0069]

[0070] where, represents the probability that the model predicts the sample x i belongs to the class . In actual training, the stochastic gradient descent process usually uses mini-batch data to update the parameters. In each iteration, the training set is randomly divided into multiple mini-batches Calculate the gradient for each batch and update the parameters. The cross-entropy loss on a mini-batch and the final parameter update rule are as follows:

[0071]

[0072]

[0073] where η is the learning rate, θ t and b t are the model parameters and the sampled mini-batch at time step t, respectively.

[0074] In step 1), the small-loss sample selection strategy can also use sample selection methods based on model predictions and sample selection methods based on training dynamics.

[0075] The above are existing sample selection methods. To alleviate the problem that the small-loss criterion ignores the correct labels in high-loss samples, we further combine a sample selection method based on the trend of model prediction confidence;

[0076] 2) For all the sample data to be selected, the sample selection method based on the trend of model prediction confidence selects the correct labels;

[0077] 2.1) Form a k-class dataset containing noisy labels, denoted as where represents the feature of the i-th sample in the dataset, is the observed label of the i-th sample;

[0078] is the sample feature space, and the label space is the discrete label set; denote y i as the true label of sample x i ;

[0079] 2.2) For each sample in the noisy training set, collect the following confidence gaps:

[0080]

[0081] where represents the confidence difference between the model prediction on sample x i at the t-th training iteration and its annotated label on class c,

[0082] 2.3) After collecting these confidence gaps in different training iterations, obtain the following time series of confidence differences:

[0083]

[0084] 2.4) Determine the confidence change trend of the confidence difference time series. If the confidence gap time series has an upward trend, that is, the confidence gap between the labeled label and other label categories tends to increase with the training iterations, then the confidence of the labeled label should be the fastest growing. Such samples are regarded as potentially correctly labeled samples.

[0085] In step 2.4), the Mann-Kendall trend test method is used to determine the confidence change trend of the confidence difference time series, as follows:

[0086] Use MKTest(·) to represent the Mann-Kendall test process. This process takes the confidence difference time series as input and outputs the standardized test statistic Z. The alternative hypothesis is that the confidence gap time series has an upward trend;

[0087] The sample selection criterion is:

[0088]

[0089] where α is the selected significance level, and Z 1-α is the 100×(1 - α) percentile of the standard normal distribution. If a sample satisfies the sample selection criterion in the t-th iteration, it indicates that all confidence gaps have an upward trend, and it is regarded as a correct label sample. C t is the set of correctly labeled samples selected in the t-th iteration.

[0090] 3) Use the union of the samples selected in steps 1) and 2) as the new training sample set and continue training with the following loss function;

[0091] The loss function used for training is:

[0092]

[0093] where, is the cross-entropy loss

[0094]

[0095] represents the expectation of the sample in the data distribution D, is the indicator function.

[0096] The present invention proposes a sample selection method based on confidence trend tracking, which can accurately distinguish the correct labels and incorrect labels in high-loss data by monitoring the change trend of the model prediction confidence. Even if the loss of correctly labeled data is large, as long as the rising speed of the model prediction confidence on its correct label is the fastest, the method proposed by the present invention can regard this sample as a correctly labeled sample, thus alleviating the problem that the traditional small-loss strategy ignores the correct labels in high-loss samples during sample selection.

[0097] Verification experiment:

[0098] The present invention conducts tests on synthetic noise labels and real noise labels on the CIFAR-10 and CIFAR-100 datasets. The types of synthetic noise labels tested include: (1) Symmetric noise: randomly replacing the labels of a certain proportion of samples with other labels with equal probability; (2) Asymmetric noise: replacing the labels of a certain proportion of samples with specific labels according to their correct labels. In the CIFAR-10 dataset, the asymmetric label replacement rules are: truck → car, bird → plane, deer → horse, cat → dog, dog → cat. In the CIFAR-100 dataset, 100 categories are divided into 20 groups with 5 categories in each group, and the replacement is cycled within each group of labels. For example, categories 1-5 are a group, and the replacement rules are: 1 → 2, 2 → 3, 3 → 4, 4 → 5, 5 → 1. The real labels tested come from the CIFARN dataset, and the noise labels provided by this dataset are labeled through the Amazon Machine Learning crowdsourcing platform. A total of 6 sets of noise labels with different noise rates are provided, and their statistical data are shown in Table 1.

[0099] Table 1 Statistical results of the CIFARN dataset

[0100]

[0101]

[0102] The baseline methods selected for this verification experiment include:

[0103] 1. Do not use any noise label learning strategy: Train the model using ordinary cross-entropy, which provides the reference performance lower limit of the noise label learning method. Denote this method as CE.

[0104] 2. Methods based on model prediction: Co-teaching and CNLCU. Both of these methods train two collaborative networks and mutually select correct samples for training based on the small-loss strategy, which are classic methods in noise label learning.

[0105] 3. Methods based on dynamic prediction: L2D, AUM, HWM, and DIST. Among them, L2D dynamically trains an LSTM-based noisy label recognition network during model training. This network takes the historical prediction sequence of the model for a certain sample as input and outputs the probability that the sample belongs to a clean sample. The noisy label recognition network was trained according to the method in the original L2D paper and used to identify the incorrect labels in the datasets involved in the experiments of the present invention. AUM and HWM rank samples through the average log margin to select the correct labels. DIST uses the momentum average of the maximum confidence of the model prediction for each sample as the threshold. If the prediction confidence of the model for the labeled label is higher than the threshold, it is regarded as a correctly labeled sample. This method sets an independent threshold for each sample, improving the sample selection performance.

[0106] In the experimental part, models were trained on the noisy training set using different noisy label learning methods, and the accuracy on the clean test set was used as the performance metric. To demonstrate the effectiveness of the method proposed in the present invention, in the experimental part, the method proposed in the present invention was combined with GMM, AUM, and DIST (+CT), and the performance before and after combination was compared.

[0107] Table 2 shows the experimental results on the synthetic noisy dataset. Table 3 shows the experimental results on the real-world noisy dataset. The average performance and variance of different methods on the clean test set under 5 random seeds are reported in the table. Sym. represents symmetric noise, and Asym. represents asymmetric noise. The present invention uses the Wilcoxon signed-rank test with a confidence level of 0.05 to compare the performance, and the results with significant improvements after combination (+CT) are shown in bold. The experimental results of both groups show that after combining the method proposed in the present invention, GMM, AUM, and DIST all achieved performance improvements, verifying the effectiveness of the present invention.

[0108] Table 2 Comparison of clean test set accuracies of different noisy label learning methods on synthetic noise

[0109]

[0110]

[0111] Table 3 Comparison of clean test set accuracies of different noisy label learning methods on real noise

[0112]

[0113] To further demonstrate the effectiveness of the method proposed in the present invention, the selection accuracy and recall rate of correct labeled samples for different sample selection methods on different datasets were directly calculated. Assume that the set of correctly labeled samples in the training set is C, and the set of correctly labeled samples identified by the sample selection method is The calculation methods of the sample selection accuracy (Precision) and recall rate (Recall) are as follows:

[0114]

[0115] In the formula, the |·| operator represents the number of elements in the set. Table 4 shows the sample selection accuracy and recall rate of different methods for different types of noisy labels on the CIFAR-100 dataset. Among them, Sym.50% represents 50% symmetric noise, Asym.40% represents 40% symmetric noise, and Real.40% represents the CIFAR-100N-Fine noisy label. The method proposed by the present invention effectively improves the sample selection recall rate and maintains the accuracy under different noisy labels, indicating that the method proposed by the present invention accurately mines the correct labels from the samples rejected by the existing sample selection methods, verifying the effectiveness of the present invention.

[0116] Table 4 Comparison of selection accuracy and recall rate of different sample selection methods on the CIFAR-100 dataset

[0117]

[0118] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A sample selection method based on the trend of model prediction confidence, characterized in that: The following steps are involved: 1) Obtain all sample data to be selected; 2) For the acquired sample data, a k-classification dataset containing noise labels is formed, denoted as in represents the characteristics of the i-th sample in the data set, is the observation label of the i-th sample; is the sample feature space, label space is a discrete label set; y i For sample x i The true label of 3) For each sample in the training set The following confidence gaps are collected: In the formula, Represents the number of samples x in the tth training iteration i , the model predicts its annotation label The difference in confidence with category c, 4) After collecting the confidence gaps in different training iterations, the following confidence difference time series is obtained: 5) Determine the confidence change trend of the confidence difference time series. If the confidence difference time series has an upward trend, that is, the confidence gap between the labeled label and other label categories tends to increase with the training iterations, then this sample is regarded as a correctly labeled sample.

2. The sample selection method based on the model prediction confidence change trend according to claim 1, characterized in that: In step 5), the Mann-Kendall trend test method is used to determine the confidence change trend of the confidence difference time series.

3. The sample selection method based on the model prediction confidence change trend according to claim 1, characterized in that: In step 5), the Mann-Kendall trend test method is used to determine the confidence change trend of the confidence difference time series, as follows: Use MKTest(·) to represent the Mann-Kendall test procedure, which differs the time series with confidence As input, the output is the standardized test statistic Z, and the alternative hypothesis is that the confidence gap time series has an upward trend; The sample selection criteria are: In the formula, α is the selected significance level, Z 1-α is the 100×(1-α) percentile of the standard normal distribution. If a sample If the sample selection criteria are met in the tth iteration, it means that all confidence gaps have an upward trend and are considered as correct label samples. t is the set of correctly labeled samples selected in the tth iteration.

4. The sample selection method based on the model prediction confidence change trend according to claim 1, characterized in that: The setting range of t is [30,100].

5. A sample selection method based on the trend of model prediction confidence, comprising the following steps: 1) Obtain all the sample data to be selected and select the correct label using a small loss sample selection strategy; 2) Obtain all the sample data to be selected and select the correct label based on the sample selection method based on the trend of model prediction confidence change; 2.1) Form a k-classification dataset containing noise labels, denoted as in represents the characteristics of the i-th sample in the data set, is the observation label of the i-th sample; is the sample feature space, label space is a discrete label set; y i For sample x i The true label of 2.2) For each sample in the noise training set The following confidence gaps are collected: In the formula, Represents the number of samples x in the tth training iteration i , the model predicts its annotation label The difference in confidence with category c, 2.3) After collecting these confidence gaps in different training iterations, the following confidence difference time series is obtained: 2.4) Determine the confidence change trend of the confidence difference time series. If the confidence gap time series has an upward trend, that is, the confidence gap between the labeled label and other label categories tends to increase with the training iterations, then this sample is regarded as a potential correctly labeled sample. 3) Use the union of the samples selected in step 1) and step 2) as a new training sample set, and continue training to obtain a correctly labeled sample set.

6. The sample selection method based on the model prediction confidence change trend according to claim 5, characterized in that: In the step 1), the small loss sample selection strategy includes a sample selection method based on model prediction and a sample selection method based on training dynamics.

7. The sample selection method based on the model prediction confidence change trend according to claim 5, characterized in that: In step 3), the loss function used in training is: in, is the cross entropy loss function; Represents a sample in the data distribution D expectations, is the indicator function.

8. The sample selection method based on the model prediction confidence change trend according to claim 5, characterized in that: In step 2.4), the Mann-Kendall trend test method is used to determine the confidence change trend of the confidence difference time series, as follows: Use MKTest(·) to represent the Mann-Kendall test procedure, which differs the time series with confidence As input, the output is the standardized test statistic Z, and the alternative hypothesis is that the confidence gap time series has an upward trend; The sample selection criteria are: In the formula, α is the selected significance level, Z 1-α is the 100×(1-α) percentile of the standard normal distribution. If a sample If the sample selection criteria are met in the tth iteration, it means that all confidence gaps have an upward trend and are considered as correct label samples. t is the set of correctly labeled samples selected in the tth iteration.

9. An electronic device, characterized in that: include: one or more processors; as well as a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are caused to perform the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.