Meta-Tag Self-Correction Method Based on Feature Constraints
By introducing feature constraint loss in the meta-learning process, dividing the data set based on gradient change metrics and training feature extraction and label correction networks, the difficulty of manually specifying clean data in the existing methods is solved, and the accuracy of noise label correction and model generalization performance is improved.
Patent Information
- Application Number
- CN202510170833.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-02-17
AI Technical Summary
The existing meta-learning method requires manual specification of clean data when processing noise labels, which is difficult to precondition and ignores the impact of the overall generalization performance of the model.
By dividing the data set based on the gradient change metric of the classification model, introducing feature constraint loss as a regular term, building a meta-learning model, using the clean sample set and the noise sample set to train the feature extraction network and the label correction network, and performing label self-correction.
It improves the accuracy of meta tag correction, improves the generalization performance of the model, effectively reduces the noise rate of the data set, and improves the performance of the classification model.
Smart Images

Figure CN120107668B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of label correction, and more specifically, relates to a method for self-correcting meta-labels based on feature constraints. Background Art
[0002] Today, with the rapid development of big data technology, the generation, collection, and use of massive data have given rise to various data processing methods. Among them, the research on deep learning methods is particularly hot and widely applied, having an important impact on people's production and life. The training of traditional deep learning models heavily relies on high-quality data and high-quality annotations, and it is difficult to avoid errors during the actual collection, storage, and annotation of data. Therefore, directly using existing noisy label data for training usually results in poor generalization performance of the model, while discarding the labels and using it as unsupervised data will lead to incomplete utilization of information. Therefore, more and more researchers have started to study the problem of learning with noisy labels (LNL).
[0003] According to different processing methods, the methods for dealing with the LNL problem can be roughly divided into the following three categories: 1) Data: Such methods start from the data and add a non-linear mapping layer (i.e., implement a noise conversion matrix) in the network model learning to achieve data adaptation and reduce the impact of noisy labels on model training. There are also methods that correct the loss from the perspective of data loss to alleviate this problem; 2) Loss function: Such methods include starting from the regularization term and proposing to add a suitable regularization term to the loss to alleviate the problem of noisy labels, and there are also methods that propose to design a robust loss function to learn a model with stronger generalization performance; 3) Optimization: Such methods attempt to alleviate this problem from the perspective of optimization. For example, a new way of solving the problem is introduced to make the obtained result closer to the theoretical optimal solution, thereby improving the performance of the model.
[0004] The method for dealing with noisy labels based on the meta-learning method is one of the optimizations. This method usually uses artificially given clean data as the source of meta-data, and then uses the learned meta-knowledge to guide the training of noisy data. The common meta-learning process usually includes two stages, namely the meta-training stage and the meta-testing stage. In the meta-training stage, meta-knowledge is learned through clean data, and then it guides the training of the model (such as a classification model) in the meta-testing stage. For example, Shu et al. proposed to use the loss-weighted network learned through specified clean data in the meta-training stage to provide weight support for the update of the classification network parameters in the meta-testing stage to alleviate the problem of high loss values brought by noisy label samples. There are also researchers who have proposed label correction methods based on meta-learning. For example, Zheng et al. proposed to use the label correction network learned in the meta-training stage to alleviate this problem.
[0005] However, there is still room for improvement in existing meta - learning methods: 1) All of the above - mentioned existing methods require manual specification of clean data, and the pre - conditions are relatively difficult; 2) In the model learning process, only the minimization of the optimization loss is considered, ignoring the impact of the overall generalization performance of the model on the model. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a meta - label self - calibration method based on feature constraints. First, it distinguishes noise and clean - label data in the dataset based on the measurement of the gradient change of the classification model, and then introduces a feature - constraint loss as a regularization term in the label - calibration stage to improve the generalization performance of the meta - learning model, thereby improving the accuracy of meta - label calibration.
[0007] To achieve the above - mentioned invention purpose, the meta - label self - calibration method based on feature constraints of the present invention includes the following steps:
[0008] S1: For the image dataset X that needs to perform label self - calibration, each image sample in it is denoted as I n , with the corresponding original label being y n , n = 1, 2, …, N, where N represents the number of image samples in the image dataset; the image dataset X is divided into several batches, and a preset classification model is used to perform M - round label prediction on the image dataset, and the gradient of the loss function of each image sample in each round of label prediction is calculated m = 1, 2, …, M; the score of each image sample is calculated using the following formula n :
[0009]
[0010] where, Norm() represents normalization, || ||2 represents taking the second - norm, and λ represents a preset weight;
[0011] A separation parameter s is preset, and its value range is s < 0.5. If the score score n of image sample n < s, then it is regarded as a clean sample. If the score s of image sample n ≤ score n < 1 - s, then it is regarded as a complex sample. If the score score n of image sample n ≥ 1 - s, then it is regarded as a noise sample; thus, the image dataset is divided into a clean - sample set a noise - sample set and a complex - sample set
[0012] S2: Two feature - extraction networks with the same structure are set according to actual needs to extract the feature images of the input image samples;
[0013] Set up a label correction network according to actual needs, which is used to receive the feature image output by the first feature extraction network and predict the correction label of the input image sample;
[0014] Set up a target classification model according to actual needs, which is used to receive the feature image output by the second feature extraction network and obtain the classification label of the input image sample;
[0015] Take the learning of the second feature extraction network and the target classification model as the lower-level tasks of meta-learning, and take the first feature extraction network and the label correction network as the upper-level tasks of meta-learning, so as to obtain a meta-learning model;
[0016] Initialize the parameters of the second feature extraction network The parameter w of the target classification model (1) and the parameters of the first feature extraction network The parameter θ of the label correction network (1) ;
[0017] S3: Denote the iteration number t = 1;
[0018] S4: Respectively sample the clean sample set the noise sample set and the complex sample set to obtain the clean sample set X cl 、the noise sample set X ny and the complex sample set X cp ;
[0019] S5: Perform meta-learning for this round based on feature constraints. The specific method is as follows:
[0020] S5.1: Use the current complex sample set X cp to initially train the target classification model, and update the parameters of the second feature extraction network and the target classification model respectively to obtain the parameters of the feature extraction network and the initial training parameters of the target classification model Among them, the loss function of the second feature extraction network adopts the feature constraint loss function, and for each sample x cp in the complex sample set X i the feature constraint loss function is calculated as follows:
[0021]
[0022] Among them, represents the feature extracted by the second feature extraction network with parameters , δ1 represents a |X cp |-dimensional vector, and X cp,k represents the current complex sample set Xcp The sample set of medium category k, where α1 represents the |X cp,k | - dimensional vector, represents the complex sample set X cp the other sample sets in X except for the samples of category k, and β1 represents - dimensional vector;
[0023] During the update of the target classification model parameters, for each sample x in the complex sample set X, the loss function i is calculated as follows:
[0024]
[0025] where, represents the preset classification loss function, represents the label corrected by the label correction network with parameter θ (t) for sample x, i f(x i , w (t) ) represents the label predicted by the target classification model with parameter w (t) for sample x, i γ vc represents the preset balance coefficient;
[0026] S5.2: Use the current clean sample set X cl , the noise sample set X ny to train the label correction network, and update the parameters of the first feature extraction network and the label correction network respectively to obtain the parameters of the first feature extraction network and the training parameters θ (t+1) of the label correction network; among them, the loss function of the first feature extraction network adopts the feature constraint loss function, and the calculation formula of the feature constraint loss function for each complex sample x in the clean sample set X cl and the noise sample set X ny is as follows: j
[0027]
[0028] where, z ∈ {cl, ny}, represents the feature extracted by the first feature extraction network with parameter , δ2 represents the |X z | - dimensional vector, x j represents the j - th sample in the sample set , represents the sample set of category k in the sample set, and α2 represents the |X z,k | - dimensional vector, Denote the sample set The other sample sets except for category k in, and β2 denotes dimensional vector;
[0029] The loss function of each sample x in the label correction network j is as follows The calculation formula of the loss function of is as follows:
[0030]
[0031] Wherein, denotes the preset label correction loss denotes the sample x j current label, denotes the parameter is the target classification model of predicts the label of the sample x j ;
[0032] S5.3: Using the updated label correction network, use the complex sample set to retrain the second feature extraction network and the target classification model, update the parameters of the second feature extraction network and the target classification model respectively, and obtain the parameters of the feature extraction network and the parameter w of the target classification model (t+1) ; Among them, in the update of the feature extraction network parameters, for each sample x in the complex sample set the feature constraint loss function of i is as follows The calculation formula of is as follows:
[0033]
[0034] Wherein, denotes the feature extracted by the second feature extraction network with the parameter ;
[0035] In the update of the target classification model parameters, for each sample x in the complex sample set the loss function of i is as follows The calculation formula of is as follows:
[0036]
[0037] Wherein, denotes the label corrected for the sample x by the label correction network with the parameter θ (t+1) ; i ;
[0038] S6: Determine whether t < t max , t maxIndicates the preset maximum number of iterations. If so, go to step S7; otherwise, go to step S11;
[0039] S7: Determine whether t%Q = 0, where Q represents the preset data set update period. If so, go to step S8; otherwise, directly go to step S9;
[0040] S8: Let t = t + 1, and return to step S5;
[0041] S9: For the clean sample set Let each sample label therein remain unchanged; for the noisy sample set and the complex sample set Input each sample therein into the current first feature extraction network and label correction network, and use the obtained label as the corrected label of the sample and update its current label for it;
[0042] S10: Statistically calculate the gradients of the loss function of each image sample in the most recent Q training times, calculate the average score of each image sample in the most recent Q training times using the same method as in step S1, and re-partition the N image samples to obtain the clean sample set the noisy sample set and the complex sample set Let t = t + 1, and return to step S4;
[0043] S11: Use the labels of the current N image samples as their final labels to complete label self-correction.
[0044] The meta-label self-correction method based on feature constraint in the present invention first uses a preset classification model to perform several rounds of label prediction on an image data set, calculates the score of each image sample according to the gradient of the loss function of each image sample, and accordingly partitions the image data set to obtain a clean sample set, a noisy sample set, and a complex sample set. A label correction network and two feature extraction networks are set. The learning of the second feature extraction network and the target classification model is used as the lower-level task of meta-learning, and the first feature extraction network and the label correction network are used as the upper-level task of meta-learning, so as to obtain a meta-learning model; then the complex sample set is used to train and update the second feature extraction network and the target classification model, and the clean sample set and the noisy sample set are used to train and update the first feature extraction network and the label correction network. During this process, the image data set is regularly re-partitioned, and the current label of each image sample after the iteration is used as its final label to complete label self-correction.
[0045] The present invention introduces a feature constraint loss as a regularization term in the meta-learning process to learn more accurate features, thereby training a more accurate label correction network and improving the accuracy of meta-label correction. Description of the Drawings
[0046] Figure 1 is a flowchart of a specific implementation manner of the meta - label self - calibration method based on feature constraints of the present invention;
[0047] Figure 2 is a flowchart of meta - learning based on feature constraints in the present invention;
[0048] Figure 3 is a comparison heat - map before and after calibration of the present invention and a comparative method on the CIFAR10 dataset with 50% symmetric noise data in this embodiment. Specific Embodiment
[0049] The following describes the specific implementation manner of the present invention with reference to the drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may obscure the main content of the present invention, these descriptions will be omitted here.
[0050] Figure 1 is a flowchart of a specific implementation manner of the meta - label self - calibration method based on feature constraints of the present invention.
[0051] As Figure 1 shown, the meta - label self - calibration method based on feature constraints of the present invention includes the following steps:
[0052] S101: Dataset Partitioning:
[0053] For the image dataset X that needs to perform label self - calibration, each image sample in it is denoted as I n , and the corresponding original label is y n , n = 1, 2, …, N, where N represents the number of image samples in the image dataset. The image dataset X is divided into several batches, and a preset classification model is used to perform M - round label prediction on the image dataset, and the gradient of the loss function of each image sample in each round of label prediction is calculated m = 1, 2, …, M. In this embodiment, the loss function of each image sample in label prediction uses cross - entropy loss, and its calculation formula is as follows:
[0054]
[0055] where y k represents the probability of the original label of the image sample in category k, and p kDenote the probability that the predicted label of the image specimen is in category k, where k = 1, 2, …, K, and K represents the number of categories. Since the dataset labels contain noise, the network learning guided by this loss function will have relatively large gradient changes for the same sample in different rounds. Therefore, the present invention judges the label quality according to the change score of the predicted sample gradient and divides the dataset accordingly. The specific method is as follows:
[0056] Calculate the score of each image sample using the following formula n :
[0057]
[0058] where Norm() represents normalization, || ||2 represents taking the second norm, and λ represents a preset weight used to adjust the parameter to balance the two gradient change value measurement parts.
[0059] In the above formula, the first term is to calculate the mean of the gradient norms of adjacent two rounds and normalize the result, and the second term is to calculate the cosine value of the included angle of the gradient changes of adjacent rounds.
[0060] The larger the score, the more likely the label of the sample is a noise label. Given that the directly learned gradient change values in a high-noise situation cannot quickly distinguish clean and noise labels, the present invention selects to use scoring and ranking to divide the dataset labels. That is, a separation parameter s is preset, and its value range is s < 0.5. If the score score n of the image sample n < s, then it is regarded as a clean sample. If the score s ≤ score n of the image sample n < 1 - s, then it is regarded as a complex sample. If the score score n of the image sample n ≥ 1 - s, then it is regarded as a noise sample. Thus, the image dataset is divided into a clean sample set a noise sample set and a complex sample set
[0061] S102: Construct a meta-learning model:
[0062] Set two feature extraction networks with the same structure according to actual needs to extract the feature images of the input image samples.
[0063] Set a label correction network according to actual needs to receive the feature image output by the first feature extraction network and predict the corrected label of the input image sample.
[0064] Set a target classification model according to actual needs to receive the feature image output by the second feature extraction network and obtain the classification label of the input image sample.
[0065] Taking the learning of the second feature extraction network and the target classification model as the lower-level tasks of meta-learning, and taking the first feature extraction network and the label correction network as the upper-level tasks of meta-learning, a meta-learning model is obtained.
[0066] Initialize the parameters of the second feature extraction network The parameter w of the target classification model (1) and the parameters of the first feature extraction network The parameter θ of the label correction network (1) .
[0067] S103: Denote the iteration number t = 1.
[0068] S104: Sample sampling:
[0069] Perform sample sampling on the clean sample set the noise sample set and the complex sample set respectively to obtain the clean sample set X cl 、the noise sample set X ny and the complex sample set X cp .
[0070] To improve the sample quality of the sample set, the images of the current batch of sample sets can be enhanced. According to the different impacts of different images on the model, in this embodiment, strong enhancement is used for clean samples, such as AutoAugment or RandAugment (in practical applications, the enhancement method can be selected according to different situations), and weak enhancement is used for noise samples and complex samples. The weak enhancement methods include one or several of flipping, rotating, and cropping.
[0071] S105: Meta-learning based on feature constraints:
[0072] To increase the generalization performance of the model, the feature extraction network introduced in the present invention learns constrained features. According to existing research, the generalization error bound holds with a probability not less than 1 - δ, and its specific form is shown in formula (3):
[0073]
[0074] Among them, P() represents probability, E(h) represents the true error of the model, represents the empirical error of the model, and the difference between the two is the generalization error bound of the model. e represents the natural constant, and n represents the number of samples in the data set. It can be seen from formula (3) that the upper bound of the generalization error is bounded by the VC dimension value d of the model VC . Therefore, reducing the value of d VC can reduce the upper bound of the generalization error. Therefore, the generalization error can be minimized by minimizing d VCTo enhance the model generalization performance.
[0075] According to Vapnik's theory, for a data partitioning model, the value of its VC dimension satisfies the following formula:
[0076]
[0077] Among them, R represents the radius of the sphere when the original data is projected onto the sphere, and M is the distance from the partitioning hyperplane to the nearest sample. For a classification task, the data dimension d is usually much larger than R 2 / M 2 . Therefore, the optimization of the VC dimension value can be transformed into the solution of R and M.
[0078] Denote as the model feature extraction parameter, and its value space is Ψ. Then formula (4) can be rewritten as formula (5):
[0079]
[0080] Among them, X + and X - are the data for selecting categories and the sample data of other categories.
[0081] To facilitate the solution of R and M, the present invention introduces the definition of the convex hull ch() by Boyd et al., and its form is shown in formula (6):
[0082]
[0083] Among them, the α dimension is the standard form of |X|-1.
[0084] It can be seen from formula (6) that the convex hull of the data set is the linear weighted sum of the samples in the data set. Since the distance from any point inside the sphere to the center of the sphere is not greater than the radius, there is formula (7):
[0085]
[0086] Among them, ball represents the smallest sphere that can contain all samples X, its radius is r, and the center of the circle is c.
[0087] According to formula (6), it can be known that the convex hull of the data set can be written as the weighting of the data set samples. Therefore, let C(X)=Xη, then there is R 2 is And according to the definition of the convex hull as the smallest convex polygon containing all the data in the data set, the corresponding vector η needs to meet the standard form constraint to make the obtained radius R(X) the smallest. Then there is formula (8):
[0088]
[0089] Define the distance between two convex hulls as where X + and X - represent the data divided into two categories by the hyperplane. Then, by integrating this formula with formula (6), formula (9) can be obtained:
[0090]
[0091] where α and β are learnable vectors to assist in projecting the data into the interior. Finally, substituting formulas (8) and (9) into (5), the final optimization objective formula can be obtained as shown in (10):
[0092]
[0093] where, represents the parameters of the feature extraction network, Ψ represents the value space of the parameters of the feature extraction network. X represents the input sample set, x represents the sample in the input sample set, represents the features extracted by the feature extraction network, δ represents a vector of |X| dimensions, || represents obtaining the number of samples in the sample set. X k represents the sample set of class k in the input sample set X, α represents a vector of |X k | dimensions, represents the sample set other than class D in the input sample set X, β represents dimensional vector.
[0094] Based on the above analysis, the present invention makes the feature constraint loss Then, the feature constraint loss is incorporated into the meta-learning process as a regularization term, and a meta-learning method based on feature constraint is proposed. Figure 2 is the flowchart of the meta-learning based on feature constraint in the present invention. As Figure 2 shown, the specific steps of the meta-learning based on feature constraint in the present invention include:
[0095] S201: Initial training of the target classification model:
[0096] Use the current complex sample set X cp to perform the initial training on the target classification model, and update the parameters of the second feature extraction network and the target classification model respectively, to obtain the parameters of the second feature extraction network and the initial training parameters of the target classification model. Among them, the loss function of the second feature extraction network adopts the feature constraint loss function, and for each sample x cp in the complex sample set X i , the feature constraint loss function is calculated as follows:
[0097]
[0098] Among them, represents the feature extracted by the second feature extraction network with parameter , and δ1 represents the X cp -dimensional vector, where X cp,k represents the current complex sample set X cp and the sample set of class k in it, and α1 represents |X cp,k |-dimensional vector. represents the sample set other than class k in the complex sample set X cp , and β1 represents -dimensional vector.
[0099] At this time, the update formula of the feature extraction network parameters can be expressed as follows:
[0100]
[0101] Among them, η vc represents the learning rate of the feature extraction network.
[0102] In the update of the target classification model parameters, for each sample x in the complex sample set i , the calculation formula of the loss function is as follows:
[0103]
[0104] Among them, represents the preset classification loss function. In this embodiment, the classification loss function uses cross-entropy loss. represents the label corrected for the sample x (t) by the label correction network with parameter θ i , and f(x i , w (t) ) represents the label predicted for the sample x (t) by the target classification model with parameter w i , and γ vc represents the preset balance coefficient, which is used to adjust the weight between the classification loss function and the feature constraint loss function.
[0105] At this time, the update formula of the target classification model parameters can be expressed as:
[0106]
[0107] Among them, η train represents the preset learning rate of the target classification model.
[0108] Since the training of the target classification model depends on correcting the network parameters θ, the initial training update here is not used as the final result of the classification network update in this round.
[0109] S202: Training of the label correction network:
[0110] Next, use the current clean sample set X cl , the noise sample set X ny to train the label correction network, update the parameters of the first feature extraction network and the label correction network respectively, and obtain the parameters of the first feature extraction network and the training parameters θ of the label correction network (t+1) . Among them, the loss function of the first feature extraction network adopts the feature constraint loss function. For each complex sample x cl in the clean sample set X ny and the noise sample set X j , the calculation formula of the feature constraint loss function is as follows:
[0111]
[0112] where z ∈ {cl, ny}, represents the feature extracted by the first feature extraction network with parameters , δ2 represents a |X z [[ID=3(]]| - dimensional vector, x j represents the j - th sample in the sample set , represents the sample set of category k in the sample set , α2 represents a |X z,k | - dimensional vector, represents the sample set of other samples except category k in the sample set , β2 represents - dimensional vector.
[0113] At this time, the update formula of the first feature extraction network parameters can be expressed as follows:
[0114]
[0115] For each sample x j in the label correction network, the loss function of the loss function has the following calculation formula:
[0116]
[0117] where, represents the preset label correction loss, represents the current label of the sample x j , represents with parameters The target classification model for the sample x j The predicted label. In this embodiment, different calculation methods are used for the label correction loss of clean samples and noise samples, and the calculation formula is as follows:
[0118]
[0119] At this time, the update formula for the label correction network parameters is as follows:
[0120]
[0121] Among them, η meta Represents the learning rate of the preset target classification model.
[0122] In order to make more full use of the learned information, in this embodiment, the current label of the sample x j Of From the label of the previous round And the prediction result of the target classification model in this round Are fused to generate a new label to guide the training of the model, and its form is as follows:
[0123]
[0124] Among them, λ y Is a preset weight used to balance the two items, Is the sample x j The initial label y of j , so that the given label information in the dataset is also utilized.
[0125] S203: Retrain the target classification model:
[0126] Since the label contains noise in the initial training of the target classification network in step S201, the gradient obtained by the first update is Not very accurate. Therefore, the present invention uses the updated label correction network and uses the complex sample set To retrain the second feature extraction network and the target classification model, and update the parameters of the second feature extraction network and the target classification model respectively to obtain the parameters of the feature extraction network And the parameters w of the target classification model (t+1) ; Among them, in the update of the feature extraction network parameters, for each sample x in the complex sample set The feature constraint loss function of i The calculation formula is as follows: The calculation formula is as follows:
[0127]
[0128] Among them, Features extracted by the second feature extraction network with parameters .
[0129] At this time, the update formula for the parameters of the second feature extraction network can be expressed as follows:
[0130]
[0131] In the update of the target classification model parameters, for each sample x in the complex sample set i , the loss function is calculated as follows:
[0132]
[0133] where represents the label corrected for the sample x (t+1) by the label correction network with parameters θ i .
[0134] At this time, the update formula for the parameters of the target classification model can be expressed as:
[0135]
[0136] S106: Determine whether t < t max , where t max represents the preset maximum number of iterations. If so, go to step S107; otherwise, go to step S111.
[0137] S107: Determine whether t % Q = 0, where Q represents the preset data set update period. If so, go to step S108; otherwise, go to step S109.
[0138] S108: Let t = t + 1, and return to step S105.
[0139] S109: Label update:
[0140] For the clean sample set , keep each sample label unchanged. For the noisy sample set and the complex sample set , input each sample into the current first feature extraction network and label correction network, and use the obtained label as the corrected label for the sample and update its current label .
[0141] S110: Re-partition the data set:
[0142] Statistically calculate the gradients of the loss function for each image sample in the most recent Q training sessions. Use the same method as in step S101 to calculate the average score for each image sample in the most recent Q training sessions, and re-partition the N image samples to obtain a clean sample set Noisy sample set and complex sample set Let t = t + 1, and return to step S104
[0143] In practical applications, the noise rate of the dataset for initial label correction is relatively high. Therefore, a smaller separation parameter is used to facilitate the separation of noise and clean labels. Then, as the training progresses and deepens, the noise rate of the labels gradually decreases, and the separation parameter can be increased to classify more labels as clean labels. Therefore, when re-partitioning the dataset each time, let the separation parameter s = s + Δs, where Δs represents the preset growth step
[0144] S111: Complete label self-correction
[0145] Use the labels of the current N image samples as their final labels to complete label self-correction. Subsequently, the image samples and the corrected labels can be used as training samples to further update and learn the parameters of the classification network model until the training of the classification network model is completed
[0146] To better illustrate the technical effects of the present invention, specific examples are used to experimentally verify the present invention. In this experimental verification, the most commonly used typical datasets CIFAR10 and CIFAR100 in this field are used as examples, and the present invention is compared and analyzed with other existing methods in this field. Among them, the present invention is implemented using PyTorch programming and the training is completed using NVIDIA GPUs. The detailed experimental information is as follows
[0147] 1. Dataset: The CIFAR10 dataset is a basic image classification dataset. There are a total of 10 categories in this dataset, and there are 60,000 images in the entire dataset. Among them, there are 50,000 images in the training set and 10,000 images in the test set. Each image is in RGB three channels and has a pixel size of 32*32. The CIFAR100 dataset has 100 categories, and each category has 600 images of size 32*32, among which 500 are training samples and 100 are test samples
[0148] 2. Implementation details
[0149] 2.1) Network architecture
[0150] In this embodiment, when dividing the dataset, the classification model uses ResNet18, and the two feature extraction networks also use ResNet18. The label correction network is completed using a single-hidden-layer MLP (feature extraction is not involved here, and the label correction network directly uses the features extracted by the classification network as input for correction). The target classification network trained with noisy data still uses ResNet18 as the backbone network, and then a feature constraint is added on it as a regularization term to improve the accuracy of the extracted features.
[0151] 2.2) Hyperparameter setting
[0152] The value of the noise label separation parameter s is taken from [0.1, 0.2, 0.3, 0.4]. At the beginning of model training, the noise rate of the dataset is relatively high, so a smaller separation value is taken to facilitate the separation of noise and clean labels. Then, as the training progresses and the noise rate of the labels gradually decreases, the separation value is increased to classify more labels as clean labels. The feature learning constraint parameter γ vc Common values according to different datasets are [0.1, 0.01, 0.001]. λ ny Common values are [0.1, 0.01, 0.001]. As the correction process proceeds, the number of noisy labels decreases, gradually reducing the impact of noisy labels on the correction process.
[0153] 2.3) Data preprocessing
[0154] Since the labels of the original CIFAR10 dataset and CIFAR100 dataset are not noisy, in order to better measure the effectiveness of the proposed method, artificial noise is added to the training set. In this experiment, 20%, 40%, 60%, and 80% noise is added to the labels of the dataset to test the present invention. The added noise type is uniform noise, that is, a specified proportion of noise is added to each category so that the added labels still only contain 10 categories and the number of labels in each category remains basically the same.
[0155] 2.4) Comparative algorithms and evaluation metrics
[0156] In this experimental verification, the present invention is compared with previous methods in this field to demonstrate the performance of the proposed method. Specifically, Cross Entropy (CE) is used as a comparison benchmark, and the comparison methods are divided into two categories: non meta-learning-based methods such as Bootstrapping, M-Correction, and Co-teaching, and meta-learning-based methods such as MSLC and MLC. Specifically, the Bootstrapping method gradually reduces the model's dependence on noisy labels by weighted mixing the model predictions and the original labels in a certain proportion, thereby improving the robustness of training. The M-Correction method uses a bias correction module to correct noisy labels, and adjusts the model's prediction results by learning the correction function to reduce the impact of noise. Co-teaching adopts a mechanism of mutual training and sample screening of two models. By allowing the two models to only learn samples with low losses from each other, it gradually filters out possible noisy label samples. Both MSLC and MLC use a meta-learner to learn a correction network to reduce the content of noisy labels in the dataset and thus improve the model performance.
[0157] For a fair comparison, this embodiment uses the classification accuracy as an evaluation index, and its specific calculation formula is as follows:
[0158]
[0159] Where TP represents the positive examples correctly classified, TN represents the positive examples misclassified, FP represents the negative examples correctly classified, and FN represents the negative examples misclassified.
[0160] Table 1 is a comparison table of the classification results of the present invention and the comparison methods in this embodiment on the CIFAR10 dataset and the CIFAR100 dataset under different noise rates in the Symmetric mode.
[0161]
[0162] Table 1
[0163] Table 2 is a comparison table of the classification results of the present invention and the comparison methods in this embodiment on the CIFAR10 dataset and the CIFAR100 dataset under different noise rates in the Asymmetric mode.
[0164]
[0165] Table 2
[0166] The best results in Table 1 and Table 2 are highlighted in bold. As can be seen from Table 1 and Table 2, in the case of Symmetric and Asymmetric noises in the CIFAR10 and CIFAR100 datasets, the present invention has certain improvements compared with all other comparative algorithms in the case of 20%-80% noise inclusion rate. For example, when the noise inclusion rate is 20%, the overall average increase is 8.23%. Generally speaking, from the experimental results, it can be seen that the present invention has better processing ability for noisy labels compared with other comparative algorithms, and the performance of the learned classification model is better.
[0167] Figure 3 It is a comparison heat map before and after correction of the present invention and the comparative method in the present embodiment for the data with 50% symmetric noise in the CIFAR10 dataset. From Figure 3 it can be intuitively seen that the present invention effectively reduces the noise inclusion rate of the dataset, and the average noise inclusion rate of all category noisy labels is reduced by 38.64%, and the quality of the dataset is significantly improved.
[0168] Although the above description of the illustrative specific embodiments of the present invention is for the convenience of those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
Claims
1. A meta-tag self-correction method based on feature constraints, characterized in that It includes the following steps: S1: For the image dataset X that needs to perform label self - calibration, denote each image sample in it as I n , and the corresponding original label is y n , n = 1, 2, …, N, where N represents the number of image samples in the image dataset; divide the image dataset X into several batches, use a preset classification model to perform M - round label prediction on the image dataset, and calculate the gradient of the loss function of each image sample in each round of label prediction m = 1, 2, …, M; calculate the score of each image sample using the following formula n : Among them, Norm() represents normalization, || ||2 represents obtaining the second norm, and λ represents a preset weight; Pre-set the separation parameter s, whose value range is s < 0.
5. If the score of the image sample n n < s, then it is regarded as a clean sample. If the score of the image sample n satisfies s ≤ score n < 1 - s, then it is regarded as a complex sample. If the score of the image sample n n ≥ 1 - s, then it is regarded as a noise sample; thereby, the image dataset is partitioned into a clean sample set a noise sample set and a complex sample set S2: Set two feature extraction networks with the same structure according to actual needs to extract the feature images of the input image samples; Set a label correction network according to actual needs to receive the feature images output by the first feature extraction network and predict the correction labels of the input image samples; Set a target classification model according to actual needs to receive the feature images output by the second feature extraction network and obtain the classification labels of the input image samples; Take the learning of the second feature extraction network and the target classification model as the lower-level tasks of meta-learning, and take the first feature extraction network and the label correction network as the upper-level tasks of meta-learning, so as to obtain a meta-learning model; Initialize the parameters of the second feature extraction network The parameter w of the target classification model (1) and the parameters of the first feature extraction network The parameter θ of the label correction network (1) ; S3: Denote the iteration number t = 1; S4: Respectively sample the clean sample set the noise sample set and the complex sample set to obtain the clean sample set X cl , the noise sample set X ny and the complex sample set X cp ; S5: Perform meta-learning for this round based on feature constraints. The specific method is as follows: S5.1: Use the current complex sample set X cp Conduct the initial training on the target classification model, update the parameters of the second feature extraction network and the target classification model respectively, and obtain the parameters of the feature extraction network and the initial training parameters of the target classification model Among them, the loss function of the second feature extraction network adopts the feature constraint loss function, and for each sample x in the complex sample set X cp The calculation formula of the feature constraint loss function i is as follows: Among them, represents the feature extracted by the second feature extraction network with parameter , δ1 represents the |X cp |-dimensional vector, X cp,k represents the sample set of class k in the current complex sample set X cp , α1 represents the |X cp,k |-dimensional vector, represents the other sample sets in the complex sample set X cp except for class k, β1 represents -dimensional vector; Updating the parameters of the target classification model, complex sample set For each sample x i in the The calculation formula of the loss function is as follows: where l train () represents a preset classification loss function, denotes the label corrected by the label correction network with parameter θ (t) for the sample x i , and f(x i , w (t) ) represents the label predicted by the target classification model with parameter w (t) for the sample x i , and γ vc represents a preset balance coefficient; S5.2: Use the current clean sample set X cl , and the noise sample set X ny to train the label correction network, update the parameters of the first feature extraction network and the label correction network respectively, and obtain the parameters of the first feature extraction network and the training parameters θ of the label correction network (t+1) ; among them, the loss function of the first feature extraction network adopts the feature constraint loss function, and for each complex sample x cl in the clean sample set X ny and the noise sample set X j , the calculation formula of the feature constraint loss function is as follows: where z ∈ {cl, ny}, denotes the feature extracted by the first feature extraction network with parameter , δ2 represents a |X z |-dimensional vector, x j represents the j-th sample in the sample set ; represents the sample set of class k in the sample set , α2 represents a |X z,k |-dimensional vector, represents the sample set of other samples except class k in the sample set , and β2 represents a -dimensional vector; The loss function of each sample x in the label correction network j is as follows The calculation formula is as follows: Among them, l meta () represents the preset label correction loss represents the sample x j the current label, represents that the target classification model with parameters for the sample x j the predicted label; S5.3: Using the updated label correction network, using complex sample sets The second feature extraction network and the target classification model are trained again, and the parameters of the second feature extraction network and the target classification model are updated respectively to obtain the parameters of the feature extraction network. and the parameters w of the target classification model (t+1) ; In the feature extraction network parameter update, the complex sample set For each sample x i The feature-constrained loss function The calculation formula is as follows: Among them, represents the feature extracted by the second feature extraction network with the parameter ; The parameters of the target classification model are being updated, and for the complex sample set each sample x i in the loss function is calculated as follows: Among them, represents the label corrected by the label correction network with parameter θ (t+1) for the sample x i after correction; S6: Determine whether t < t max , t max represents the preset maximum number of iterations. If so, proceed to step S7; otherwise, proceed to step S11; S7: Judge whether t % Q = 0, where Q represents the preset data set update period. If so, enter step S8; otherwise, directly enter step S9; S8: Let t = t + 1 and return to step S5; S9: For the clean sample set Keep each sample label therein unchanged; for the noisy sample set and the complex sample set Input each sample therein into the current first feature extraction network and the label correction network, and use the obtained label as the corrected label of the sample and update its current label thereby. S10: Statistically calculate the gradients of the loss function for each image sample in the most recent Q training sessions, calculate the average scores of each image sample in the most recent Q training sessions using the same method as in step S1, and re-partition the N image samples to obtain a clean sample set noisy sample set and complex sample set Let t = t + 1, and return to step S4; S11: Take the labels of the current N image samples as their final labels to complete label self-correction.
2. The meta-tag self-calibration method according to claim 1, wherein In step S4, the samples in the clean sample set X cl , the noise sample set X ny and the complex sample set X cp are also enhanced. Among them, the clean samples are strongly enhanced, and the noise samples and complex samples are weakly enhanced.
3. The meta-tag self-calibration method according to claim 1, wherein The label correction loss \(l\) meta () in step S5.2 is calculated as follows:
4. The meta-tag self-calibration method according to claim 1, wherein The label in step S5.2 is obtained by using the following formula: Among them, λ y is a preset weight, and x j is the initial label y j of the sample.
5. The meta-tag self-calibration method according to claim 1, wherein When re-partitioning the data set in step S10, let the separation parameter s = s + Δs, where Δs represents the preset growth step.
Citation Information
Patent Citations
Label self-correction method based on meta learning
CN117456306A
Passive domain adaptive image classification method based on class guide meta-learning
CN117523295A