Attribute-dependent tag noise-oriented distribution adaptive identification method

By calculating the robust attribute core and using an adaptive combination of Gaussian mixed model and gamma mixed model, a distribution adaptive identification method for attribute-dependent label noise is solved, and the high stability and performance of the model in complex industrial scenarios are achieved.

CN120217233APending Publication Date: 2025-06-27SHANXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510291803.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with attribute-dependent label noise, resulting in unstable performance of models in complex industrial scenarios, and the robust mechanism of existing models is difficult to adapt to dynamically changing noise distribution patterns.

Method used

A distributed adaptive recognition method for attribute-dependent label noise is proposed. By calculating a robust attribute kernel and using an adaptive combination of Gaussian mixed model and gamma mixed model, noise samples are identified and processed, and model parameters are optimized through semi-supervised learning.

Benefits of technology

Effectively alleviate the interference of abnormal samples on class attribute cores, improve the robustness and generalization capabilities of the model on noise data sets, and significantly improve the stability and performance of the model in complex industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217233A_ABST
    Figure CN120217233A_ABST
Patent Text Reader

Abstract

The invention discloses an attribute-dependent tag noise-oriented distribution adaptive identification method, and belongs to the technical field of machine learning and data mining. Aiming at the problems of class center offset, insufficient noise distribution dynamic adaptability, single model fitting limitation and the like in the prior art, the following core schemes are provided: (1) designing a robust attribute kernel iteration algorithm, dynamically weighting samples through cosine similarity and optimizing a class center, and reducing abnormal sample interference; (2) adaptively distributing model weights based on KL divergence and a training stage by utilizing a distribution adaptive model and combining distribution fitting advantages of a Gaussian mixture model and a gamma mixture model; and (3) fusing a semi-supervised learning framework, and converting the noise sample into unlabeled data for joint training. According to the method, the effectiveness of the method is verified through experiments on CIFAR-10 and CIFAR-100 data sets of which the attributes depend on noise and real noise data sets Animal-10N and Clothing1M, wherein the CIFAR-10 and CIFAR-100 data sets and the real noise data sets Animal-10N and Clothing1M are artificially synthesized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning and data mining, and particularly relates to a distribution adaptation recognition method for attribute-dependent label noise. Background Art

[0002] As the core architecture of modern machine learning systems, the excellent modeling ability of Deep Artificial Neural Networks has been empirically verified in multiple fields. The success of these networks is inseparable from the availability of large-scale datasets used in their training. Obtaining accurately labeled data is expensive and time-consuming, and the time cost of professional annotators is high. To reduce the annotation cost, many researchers rely on the AMT platform (Amazon crowdsourcing annotation system) to obtain the datasets required for training. However, in real application scenarios, obtaining high-quality data with accurate annotations faces significant challenges. Due to the differences among data annotators and the quality problems of the images themselves, even with the use of distributed annotation schemes, it is still difficult to avoid annotation distortion problems.

[0003] The defects in annotation quality have multi-dimensional negative impacts on the model. The first and foremost is the interference with attribute representation learning: when there are systematic biases in the supervision signals, the direction of network parameter updates will deviate from the true data distribution, resulting in a significant decline in the performance of the model on benchmark datasets such as MNIST and CIFAR series. In addition, in complex visual datasets such as ImageNet and industrial application scenarios such as Clothing1M, this distortion of supervision signals will induce the model to establish false attribute associations, thereby damaging its generalization performance. It is worth noting that compared with the data noise at the input level, the impact of annotation distortion on the supervised learning framework is more fundamental - it directly contaminates the objective function for model optimization, leading to systematic biases in the parameter space search process.

[0004] From the analysis of the noise generation mechanism, existing research mainly focuses on two types of annotation distortion patterns: the first type is class-conditional noise, which is characterized by describing the annotation error rules between classes through a probability transition matrix; the second type is more realistic feature-correlated noise, whose misannotation probability is closely related to the distribution of the example attribute space. It should be particularly pointed out that the processing of the second type of noise faces three technical bottlenecks: first, the strong coupling between the noise pattern and the attribute distribution makes traditional probability modeling methods ineffective; second, the heterogeneity of the high-dimensional attribute space makes it difficult to effectively decouple the noise attributes; finally, the robustness mechanisms of existing models often struggle to adapt to the dynamically changing noise distribution patterns.

[0005] In summary, the research on learning with noisy labels for attribute-dependent label noise has significant theoretical and practical value. From the perspective of technological evolution, establishing an effective noise filtering mechanism can not only improve the stability of the model in complex industrial scenarios, but also drive the innovation of data annotation paradigms - by constructing a noise-aware intelligent annotation system, manual annotation resources can be focused on key examples, achieving double optimization of annotation efficiency and quality. This technological breakthrough will strongly promote the practical application of artificial intelligence systems in key fields such as medical image analysis and industrial quality inspection, laying a solid foundation for building trustworthy AI systems. Summary of the Invention

[0006] Aiming at the problems of class center shift, insufficient dynamic adaptability of noise distribution, and limitations of single model fitting in the existing technology, the present invention provides a distribution adaptive recognition method for attribute-dependent label noise.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A distribution adaptive recognition method for attribute-dependent label noise, comprising the following steps:

[0009] Step 1: Extract attributes for each example in the dataset to generate example attribute vectors;

[0010] Step 2: Calculate the robust attribute kernel of each class based on the example attribute vectors and their original labels;

[0011] Further, the specific operation of step 2 is as follows:

[0012] a) Weight the examples of the same class according to the reliability score of the examples;

[0013] b) Update the class attribute kernel through iterative optimization, gradually reducing the weight of abnormal examples;

[0014] First, initialize the class attribute kernel as the mean of the example attribute vectors of the corresponding class:

[0015]

[0016] In the formula: represents the initial attribute kernel of class k, n k represents the number of examples in class k; z i represents the attribute vector of example i with the predicted class ;

[0017] Secondly, the class attribute kernel is updated by weighted iteration according to the example reliability; the update formula of the attribute kernel of class k is:

[0018]

[0019] Where: n k represents the number of examples in class k; z i represents the attribute vector of the example whose predicted class is k; is the weight of example i: The calculation formula is:

[0020]

[0021] The weight of example i calculates the reliability of the example through the cosine similarity between the example and the kernel of the current class attributes. The more similar example i is to the current attribute kernel the more reliable the example label is, and its weight w i is larger; otherwise, the greater the possibility that the example is abnormal or contains noise, and its weight w i is smaller; During the iteration process, the class attribute kernel and the weight will be updated alternately.

[0022] Step 3: Calculate the cosine similarity between the attribute vector of each example and the robust attribute kernel of the corresponding class of its original label;

[0023] Further, the specific operation of step 3 is:

[0024] For each example, according to its original label y i find the corresponding final class attribute kernel Calculate the cosine similarity between the example attribute and the corresponding class attribute kernel to identify noise examples; Since both the attribute vector and the class attribute kernel are normalized, the cosine similarity is equivalent to their dot product, that is:

[0025]

[0026] Where, S i represents the cosine similarity.

[0027] Step 4: Use the distribution adaptation model to fit the cosine similarity distribution and divide clean examples and noise examples;

[0028] Further, the specific operation of step 4 is:

[0029] a) Use Gaussian mixture model and gamma mixture model to model the cosine similarity distribution respectively;

[0030] b) Dynamically adjust the combined weights of the two according to the distribution fitting indexes of the Gaussian mixture model and the gamma mixture model;

[0031] Calculate the goodness-of-fit indicators of the Gaussian mixture model and the gamma mixture model, and use the total distribution of cosine similarity and the Kullback-Leibler (KL) divergence between the total distributions fitted by the Gaussian mixture model and the gamma mixture model to evaluate the goodness-of-fit of the two models; the normalized model density f G f Γ (x), f

[0032]

[0033] In the formula: D KL (·||·) represents calculating the KL divergence of two distributions; h(x) represents the total distribution of cosine similarity, and f G (x) is the distribution density function fitted by the Gaussian mixture model, and f Γ (x) is the distribution density function of the gamma mixture model;

[0034] Allocate dynamic weights to the Gaussian mixture model and the gamma mixture model according to the goodness-of-fit indicators and the number of training rounds:

[0035] Considering that the gamma mixture model has a weak adaptability to data in the initial stage, set the initial weights of the Gaussian mixture model and the gamma mixture model to be and The trend coefficient θ is adjusted with the number of experimental rounds using the Logistic growth function to enable the gamma mixture model to fully exert its classification ability in the later stage. The formula is:

[0036]

[0037] In the formula: σ is the translation parameter of the Logistic function, representing the midpoint of alternating the weights of the two models, that is, the round when the output result of the gamma mixture model dominates. Its value is the round when the KL divergence of the gamma mixture model is first lower than that of the Gaussian mixture model; epoch is the current training round; j is the steepness parameter of the Logistic function, determining the speed of change of the model weights;

[0038] The degree of distribution fitting is evaluated by the KL divergence. The smaller the KL divergence, the better the model fitting effect, and the greater the weight of the model. Therefore, the fitting coefficients of the Gaussian mixture model and the gamma mixture model are defined as:

[0039]

[0040] In the formula, D KL (h, f G ) represents the KL divergence between the total distribution h(x) of cosine similarity and the distribution f G (x) of the Gaussian mixture model; D KL (h, fΓ ) represents the KL divergence between the overall cosine similarity distribution h(x) and the gamma mixture model distribution f Γ (x); overall, the weights of the two models are proportional to their goodness of fit and inversely proportional to their KL divergence;

[0041] The weights of the Gaussian mixture model and the gamma mixture model are finally obtained by combining the iterative trend coefficient and the fitting coefficient; in the early stage of model training, the similarity is more concentrated and symmetric, and the trend coefficient of the Gaussian mixture model is larger; in the later stage of training, the similarity is more dispersed and skewed, and the trend coefficient of the gamma mixture model is larger; in each round of training, weights are assigned to the Gaussian mixture model and the gamma mixture model. In the t-th round, the final weights of the Gaussian mixture model and the gamma mixture model are:

[0042]

[0043] Combine the prediction results of the Gaussian mixture model and the gamma mixture model according to the weights to generate a mixed noise probability score:

[0044] Based on the sample division results of the Gaussian mixture model and the gamma mixture model, the final calculation formula for the label noise probability is:

[0045]

[0046] In the formula: is the prediction result of the Gaussian mixture model, is the prediction result of the gamma mixture model;

[0047] The distribution adaptation model in step 4 includes:

[0048] In the initial stage of model training, the Gaussian mixture model is the dominant model to adapt to the symmetric similarity distribution;

[0049] In the later stage of model training, the gamma mixture model is the dominant model to adapt to the asymmetric similarity distribution.

[0050] Step 5: Use the clean examples as the labeled dataset and the noisy examples as the unlabeled dataset for semi-supervised model training, update the model parameters, and use a weighted loss function to optimize the model during training so that the model can have good performance on the noisy dataset;

[0051] Furthermore, the specific operation of step 5 is:

[0052] Based on the comprehensive prediction probability, the data is divided into a clean data set and a noise data set; the noise data set is regarded as unlabeled data, and the clean data set is regarded as labeled data. A semi-supervised learning strategy is used to jointly train the network model. A weighted loss function is used to fuse the supervised loss and the unsupervised loss, balance the contributions of the two parts of information, and update the model parameters through the loss function. The loss function is as follows:

[0053]

[0054] In the formula: D clean represents the set of clean data, x is the sample, y is the true label, represents calculating the loss for all clean data samples and taking the average, D noisy represents the noise data set, represents the pseudo label, f θ (x) represents the predicted output of the model, represents the cross-entropy loss function, λ is the balance factor, which is used to control the weight of the unsupervised loss, represents calculating the loss for all noise data samples and taking the average; balance the contributions of the two parts of information and update the model parameters through the loss function to make it obtain good robustness and generalization ability on the noise data set.

[0055] Compared with the prior art, the present invention has the following advantages:

[0056] (1) The present invention proposes a robust attribute kernel acquisition method and designs an example attribute similarity index based on the category attribute kernel, which is used as the basis for dividing clean labels and noise labels. It can effectively alleviate the interference of abnormal or unreliable examples on the class attribute kernel when calculating the example attribute kernel, and enhance the quality of classification data;

[0057] (2) The present invention proposes a distribution adaptation model that adaptively combines a Gaussian mixture model and a gamma mixture model. By combining the attribute similarity distribution fitting situations of the two models, it can flexibly adapt to the similarity distributions in different training stages of the deep model, and thus improve the label noise recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a flowchart of a distribution adaptation recognition method for attribute-dependent label noise;

[0059] Figure 2 is a comparison chart of the results of a distribution adaptation recognition method for attribute-dependent label noise and other label noise processing methods on the CIFAR-10 data set with different proportions of artificial noise added;

[0060] Figure 3It is a comparison graph of the results of a distribution adaptation recognition method for attribute-dependent label noise and other label noise processing algorithms on the CIFAR-100 dataset without adding different proportions of artificial noise;

[0061] Figure 4 It is a comparison graph of the results of a distribution adaptation recognition method for attribute-dependent label noise and other label noise processing algorithms on the real-world attribute-dependent noise dataset Clothing1M;

[0062] Figure 5 It is a comparison graph of the results of a distribution adaptation recognition method for attribute-dependent label noise and other label noise processing algorithms on the real-world attribute-dependent noise dataset Animal-10N. Detailed implementation

[0063] To understand the present invention in depth, we will describe it comprehensively and meticulously. However, the present invention has multiple implementation manners and is not limited to the specific examples listed herein. The presentation of these examples aims to deepen the comprehensive understanding of the disclosed content of the present invention.

[0064] As Figure 1 shown, the steps of the distribution adaptation recognition method for attribute-dependent label noise described in this embodiment are as follows:

[0065] A distribution adaptation recognition method for attribute-dependent label noise includes the following steps:

[0066] Step 1: Extract attributes for each example in the dataset to generate an example attribute vector;

[0067] Step 2: Calculate the robust attribute kernel for each category based on the example attribute vector and its original label;

[0068] Further, the specific operation of step 2 is:

[0069] a) Weight the examples of the same category according to the reliability score of the example;

[0070] b) Update the category attribute kernel through iterative optimization to gradually reduce the weight of abnormal examples;

[0071] First, initialize the category attribute kernel as the mean of the example attribute vectors of the corresponding category:

[0072]

[0073] In the formula: represents the initial attribute kernel of category k, and n k represents the number of examples in category k; z i represents the predicted category as The attribute vector of sample i;

[0074] Secondly, the class attribute kernel is updated iteratively with weights according to the sample reliability; the update formula for the attribute kernel of class k is:

[0075]

[0076] In the formula: n k represents the number of samples in class k; z i represents the attribute vector of the sample whose predicted class is k; is the weight of sample i: the calculation formula is:

[0077]

[0078] The weight of sample i calculates the reliability of the sample through the cosine similarity between the sample and the current class attribute kernel. The more similar sample i is to the current attribute kernel , the more reliable the sample label is, and its weight w i is larger; otherwise, the possibility of the sample being abnormal or containing noise is greater, and its weight w i is smaller; during the iteration process, the class attribute kernel and the weight are updated alternately.

[0079] Step 3: Calculate the cosine similarity between the attribute vector of each sample and the robust attribute kernel of the category corresponding to its original label;

[0080] Furthermore, the specific operation of step 3 is:

[0081] For each sample, according to its original label y i find the corresponding final class attribute kernel Calculate the cosine similarity between the sample attribute and the corresponding class attribute kernel to identify noise samples; since both the attribute vector and the class attribute kernel are normalized, the cosine similarity is equivalent to their dot product, that is:

[0082]

[0083] In the formula, S i represents the cosine similarity.

[0084] Step 4: Use the distribution adaptation model to fit the cosine similarity distribution and divide clean samples and noise samples;

[0085] Furthermore, the specific operation of step 4 is:

[0086] a) Model the cosine similarity distribution using Gaussian mixture model and gamma mixture model respectively;

[0087] b) Dynamically adjust the combined weights of the Gaussian mixture model and the gamma mixture model according to the distribution fitting metrics of both models.

[0088] Calculate the goodness-of-fit metrics of the Gaussian mixture model and the gamma mixture model, and use the Kullback-Leibler (KL) divergence between the total distribution of cosine similarity and the total distribution fitted by the Gaussian mixture model and the gamma mixture model to evaluate the goodness-of-fit of the two models; the normalized model densities f G (x), f Γ (x) and the KL divergence of the similarity histogram distribution h are respectively:

[0089]

[0090] In the formula: D KL (·||·) represents calculating the KL divergence of two distributions; h(x) represents the total distribution of cosine similarity, and f G (x) is the distribution density function fitted by the Gaussian mixture model, and f Γ (x) is the distribution density function of the gamma mixture model;

[0091] Allocate dynamic weights to the Gaussian mixture model and the gamma mixture model according to the goodness-of-fit metrics and the number of training rounds:

[0092] Considering that the gamma mixture model has relatively weak adaptability to data in the initial stage, set the initial weights of the Gaussian mixture model and the gamma mixture model to be and The trend coefficient θ is adjusted with the number of experimental rounds using the Logistic growth function to enable the gamma mixture model to fully exert its classification ability in the later stage. The formula is:

[0093]

[0094] In the formula: σ is the translation parameter of the Logistic function, representing the midpoint of alternating the weights of the two models, that is, the round when the output result of the gamma mixture model takes the lead. Its value is the round when the KL divergence of the gamma mixture model is first lower than that of the Gaussian mixture model; epoch is the current training round; j is the steepness parameter of the Logistic function, determining the speed of change of the model weights;

[0095] The degree of distribution fitting is evaluated by the KL divergence. The smaller the KL divergence, the better the model fitting effect, and the greater the weight of the model. Therefore, the fitting coefficients of the Gaussian mixture model and the gamma mixture model are respectively defined as:

[0096]

[0097] In the formula, D KL (h, f G) represents the KL divergence between the total cosine similarity distribution h(x) and the Gaussian mixture model distribution f G (x); D KL (h, f Γ ) represents the KL divergence between the total cosine similarity distribution h(x) and the gamma mixture model distribution f Γ (x); Overall, the weights of the two models are proportional to their goodness of fit and inversely proportional to their KL divergence;

[0098] The weights of the Gaussian mixture model and the gamma mixture model are finally obtained by combining the iterative trend coefficient and the fitting coefficient; In the early stage of model training, the similarity is more concentrated and symmetric, and the trend coefficient of the Gaussian mixture model is larger; In the later stage of training, the similarity is more dispersed and skewed, and the trend coefficient of the gamma mixture model is larger; In each round of training, weights are assigned to the Gaussian mixture model and the gamma mixture model. In the t-th round, the final weights of the Gaussian mixture model and the gamma mixture model are:

[0099]

[0100] Combine the prediction results of the Gaussian mixture model and the gamma mixture model according to the weights to generate a mixed noise probability score:

[0101] Combining the sample partitioning results of the Gaussian mixture model and the gamma mixture model, the final formula for calculating the label noise probability is:

[0102]

[0103] In the formula: is the prediction result of the Gaussian mixture model, is the prediction result of the gamma mixture model;

[0104] The distribution adaptation model in step 4 includes:

[0105] In the initial stage of model training, the Gaussian mixture model is used as the dominant model to adapt to the symmetric similarity distribution;

[0106] In the later stage of model training, the gamma mixture model is used as the dominant model to adapt to the asymmetric similarity distribution.

[0107] Step 5: Use the clean examples as the labeled dataset and the noisy examples as the unlabeled dataset for semi-supervised model training, update the model parameters, and use a weighted loss function to optimize the model during training so that the model can have good performance on the noisy dataset;

[0108] Furthermore, the specific operation of step 5 is:

[0109] Based on the comprehensive prediction probability, the data is divided into a clean data set and a noise data set; the noise data set is regarded as unlabeled data, and the clean data set is regarded as labeled data. A semi-supervised learning strategy is adopted to jointly train the network model. A weighted loss function is used to fuse the supervised loss and the unsupervised loss, balance the contributions of the two parts of information, and update the model parameters through the loss function. The loss function is as follows:

[0110]

[0111] In the formula: D clean represents the set of clean data, x is the sample, y is the true label, represents calculating the loss for all clean data samples and taking the average, D noisy represents the noise data set, represents the pseudo-label, f θ (x) represents the predicted output of the model, represents the cross-entropy loss function, λ is the balance factor, used to control the weight of the unsupervised loss, represents calculating the loss for all noise data samples and taking the average; balance the contributions of the two parts of information and update the model parameters through the loss function to make it obtain good robustness and generalization ability on the noise data set.

[0112] The present invention verifies the method and designs experiments for two aspects: the artificially synthesized attribute-dependent noise environment and the real-world noise environment.

[0113] Experiment (1): Verification experiment under the artificially synthesized attribute-dependent noise environment. To evaluate the robustness of the method of the present invention under the artificially synthesized attribute-dependent noise, an artificial noise test environment is constructed on the CIFAR series data sets: The experiment uses the ResNet-34 architecture as the basic model. The optimization process uses the stochastic gradient descent optimizer (momentum coefficient 0.9, L2 regularization strength 5e-4); the learning rate adopts a two-stage scheduling strategy: maintaining an initial value of 0.02 for the first 150 epochs and then decreasing to 0.002. A learning rate warm-up stage of 10 and 15 epochs is set for CIFAR-10 and CIFAR-100 respectively; attribute-dependent noise generated based on partial dependence with a proportion of 20%, 40%, and 60% is added to the CIFAR-10 and CIFAR-100 data sets respectively; the method of the present invention and other comparison methods are respectively used to train the model on the artificially noise-added data sets, and the classification accuracies of the final models are compared.

[0114] Experiment (2): Validation experiment in a real-world noise environment. The effectiveness of the method is verified by comparing the classification accuracy of the method of the present invention with other methods on a data set. The experimental settings are as follows: On the industrial-level noise data set Clothing1M, a ResNet-50 architecture based on ImageNet pre-trained parameters is deployed using a transfer learning strategy. The training process includes 80 epochs, and the learning rate decays from 0.002 to 0.0002 in the middle of the training (the 40th epoch). This setting can effectively balance the fine-tuning intensity of the feature extractor; for the experiment on the Animal-10N data set, VGG-19 is used as the backbone network and the VGG-19 network is trained for 100 epochs. The initial learning rate is 0.01, and a 5-fold stepwise decrease is performed at the 50th and 75th rounds; by comparing the classification accuracy of this method with that of the leading methods, the generalization ability of the model under the real noise distribution is verified.

[0115] For Experiment (1), Figure 2 and Figure 3 respectively show the accuracy of the method of the present invention and other comparative methods in the CIFAR-10 and CIFAR-100 data sets with attribute-dependent noise generated based on partial dependence with the noise proportion added being 20%, 40%, and 60%. It can be seen from the figure that the method of the present invention is superior to the other methods under various noise ratios. Specifically, as Figure 2 shown, most existing methods perform well at low noise ratios (η = 0.2). For example, the accuracy of methods such as Reweight-R, CORES2, and DivideMix on CIFAR-10 all exceed 90%; when η = 0.6, the overfitting of the model on large-scale noise examples causes the accuracy of most comparative methods to drop to 60%-80%. However, the method of the present invention can still exceed 90% in accuracy and perform optimally in a high-noise environment. Figure 3 shows the accuracy of each method on the more challenging CIFAR-100 data set. Due to the large number of classes and few training examples for each class, the accuracy of most methods is only 50%-70% when η = 0.2, and only the accuracy of TSCLS, SV-learner, and the method of the present invention exceeds 70%; when η = 0.6, the performance of most methods drops below 50%. For the relatively better TSCLS and SV-learner methods in CIFAR-100, the method of the present invention performs better than both under various noise levels. The above results show that the method of the present invention is more effective in dealing with the problem of attribute-dependent noise.

[0116] Figure 4 and Figure 5 present the performance comparison between the method of the present invention and the mainstream methods in a real attribute-dependent noise environment. Figure 4For the accuracy results of the method of the present invention and other methods on the industrial-grade clothing classification dataset Clothing1M in Experiment (2), the method of the present invention shows significant advantages: compared with the current optimal TSCLS model and SV-learner method, performance gains of 0.4% and 0.6% are respectively achieved, and this improvement verifies the effectiveness of the method of the present invention in the clothing classification task. Figure 5 For the comparison of the accuracy of the method of the present invention and other methods on the Animal-10N dataset in Experiment (2). The method of the present invention maintains a leading position with an identification accuracy of 84.21%, which is 0.3 percentage points higher than the sub-optimal NAL method. This result further demonstrates the robustness and generalization ability of the method of the present invention when dealing with real-world noisy data.

[0117] The content not described in detail in the specification of the present invention belongs to the prior art well-known to those skilled in the art. Although the illustrative specific embodiments of the present invention have been described above for the understanding of those skilled in the art of the present technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.

Claims

1. A distribution adaptive recognition method for attribute-dependent label noise, characterized in that: The following steps are involved: Step 1: Extract attributes for each sample in the data set and generate a sample attribute vector; Step 2: Based on the sample attribute vector and its original label, calculate the robust attribute kernel for each category; Step 3: Calculate the cosine similarity between the attribute vector of each sample and the robust attribute kernel of the corresponding category of its original label; Step 4: Use the distribution adaptive model to fit the cosine similarity distribution and divide clean samples into noise samples; Step 5: Use the clean samples as labeled data sets and the noisy samples as unlabeled data sets to perform semi-supervised model training and update the model parameters.

2. According to claim 1, a distribution adaptive recognition method for attribute-dependent label noise is characterized in that: The step 2 calculates the robust attribute kernel of each category based on the sample attribute vector and its original label. The specific operation is: a) Weighting samples of the same category according to their reliability scores; b) Update the category attribute core through iterative optimization and gradually reduce the weight of abnormal samples; First, initialize the category attribute kernel to the mean of the attribute vector of the corresponding category sample: Where: represents the initial attribute kernel of category k, n k represents the number of samples in category k; z i Indicates that the predicted category is k The attribute vector of sample i; Secondly, the category attribute core is updated in a weighted iterative manner according to the sample reliability; the attribute core update formula for category k is: Where: n k represents the number of samples in category k; z i Represents the attribute vector of the sample with predicted category k; is the weight of sample i: the calculation formula is: The weight of sample i is calculated by the cosine similarity between the sample and the current category attribute kernel to calculate the reliability of the sample.

3. According to claim 2, a distribution adaptive recognition method for attribute-dependent label noise is characterized in that: The step 3: calculating the cosine similarity between the attribute vector of each sample and the robust attribute kernel of the category corresponding to its original label, specifically, the operation is: For each example, according to its original label y i Find the corresponding final category attribute core The cosine similarity between the sample attribute and the corresponding category attribute kernel is calculated to identify noise samples. The formula is: In the formula, S i Represents cosine similarity.

4. According to claim 3, a distribution adaptive recognition method for attribute-dependent label noise is characterized in that: The step 4: using the distribution adaptive model to fit the cosine similarity distribution and divide the clean samples and the noise samples, the specific operation is: a) Gaussian mixture model and gamma mixture model are used to model the cosine similarity distribution respectively; b) dynamically adjust the combined weight of the Gaussian mixture model and the gamma mixture model according to their distribution fitting indices; The goodness-of-fit index of the Gaussian mixture model and the gamma mixture model is calculated, and the KL divergence of the total distribution of cosine similarity and the total distribution of the Gaussian mixture model and the gamma mixture model fitting is used to evaluate the goodness-of-fit of the two models; the normalized model density f G (x), f Γ The KL divergence of (x) and the similarity histogram distribution h are: Where: D KL (·||·) represents the calculation of the KL divergence of two distributions; h(x) represents the total distribution of cosine similarity, f G (x) is the distribution density function of the Gaussian mixture model, f Γ (x) is the distribution density function of the gamma mixture model; Dynamic weights are assigned to Gaussian and gamma mixture models based on goodness-of-fit metrics and training rounds: The initial weights of the Gaussian mixture model and the gamma mixture model are set to and The trend coefficient θ is adjusted with the experimental rounds using the Logistic growth function, so that the gamma mixture model can fully exert its classification ability in the later stage. The formula is: Where: σ is the shift parameter of the Logistic function, which represents the midpoint of the weights of the two alternating models, that is, the round in which the output of the gamma mixture model dominates, and its value is the round when the KL divergence of the gamma mixture model is lower than that of the Gaussian mixture model for the first time; epoch is the current training round; j is the steepness parameter of the Logistic function, which determines how fast the model weight changes; The distribution fit is evaluated by KL divergence, so the fitting coefficients of Gaussian mixture model and gamma mixture model are defined as: Where D KL (h,f G ) represents the total distribution of cosine similarity h(x) and the Gaussian mixture model distribution f G (x) KL divergence between; D KL (h,f Γ ) represents the total distribution of cosine similarity h(x) and the gamma mixture model distribution f Γ (x) KL divergence between them; Overall, the weights of the two models are proportional to their goodness of fit and inversely proportional to their KL divergence; The weights of the Gaussian mixture model and the gamma mixture model are finally obtained by combining the iterative trend coefficient and the fitting coefficient. The weights are assigned to the Gaussian mixture model and the gamma mixture model in each round of training. In the tth round, the final weights of the Gaussian mixture model and the gamma mixture model are: The prediction results of the Gaussian mixture model and the gamma mixture model are combined according to the weights to generate a mixed noise probability score: Combining the sample partitioning results of the Gaussian mixture model and the gamma mixture model, the final label noise probability calculation formula is: Where: is the prediction result of the Gaussian mixture model, This is the prediction result of the gamma mixture model.

5. According to claim 4, a distribution adaptive recognition method for attribute-dependent label noise is characterized in that: The distributed adaptive model in step 4 includes: In the early stage of model training, the Gaussian mixture model is used as the dominant model to adapt to the symmetric similarity distribution; In the later stage of model training, the gamma mixture model is used as the dominant model to adapt to the asymmetric similarity distribution.

6. According to claim 5, a distribution adaptive recognition method for attribute-dependent label noise is characterized in that: Step 5: Use the clean samples as labeled data sets and the noise samples as unlabeled data sets to perform semi-supervised model training and update model parameters. The specific operations are: The data is divided into a clean data set and a noisy data set; the noisy data set is regarded as unlabeled data, and the clean data set is regarded as labeled data. A semi-supervised learning strategy is used to jointly train the network model. A weighted loss function is used to fuse the supervised loss with the unsupervised loss, balance the contribution of the two parts of information, and update the model parameters through the loss function. The loss function is as follows: Where: D clean represents a set of clean data, x is a sample, y is the true label, Indicates that the loss is calculated and averaged for all clean data samples, D noisy represents a noisy data set, represents the pseudo label, f θ (x) represents the predicted output of the model, l(·, ·) represents the cross entropy loss function, and λ is a balancing factor used to control the weight of the unsupervised loss. Indicates that the loss is calculated and averaged for all noise data samples.