Noise label cleaning method based on visual language model and fine-grained expansion strategy

By introducing fine-grained augmentation strategy and cyclic MixFix training in the noise label cleaning method, the problem of existing methods performing poorly in high-noise environments is solved, and a more efficient and robust noise label cleaning effect is achieved.

CN120219840APending Publication Date: 2025-06-27JIANGSU OPEN UNIVERSITY (THE CITY VOCATIONAL COLLEGE OF JIANGSU)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510302724.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing noise label cleaning methods based on visual language models perform poorly at high noise ratios, have limited generalization capabilities, and rely too much on prior knowledge of visual language models to perform unstable in some scenarios.

Method used

Using a noise label cleaning method based on visual language model and fine-grained amplification strategy, the initial clean subset is filtered through the pre-trained visual language model, and two DNN models with the same structure but different initialization weights are used for semi-supervised training and cyclic MixFix training to generate a more balanced and high-quality clean subset.

Benefits of technology

It improves the robustness and classification performance of the model in high noise environment, reduces the dependence on prior knowledge of visual language models, and improves the accuracy and efficiency of sample classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219840A_ABST
    Figure CN120219840A_ABST
Patent Text Reader

Abstract

The invention discloses a noise label cleaning method based on a visual language model and a fine-grained expansion strategy, and belongs to the technical field of computer vision, and the method comprises the following steps: for an initial noise data set, screening clean subsets by using the visual language model, and pre-training two DNN models by using the initial noise data set; performing semi-supervised training of a fine-grained expansion strategy on the two pre-trained DNN models; specifically, a historical sequence is updated, and a clean subset and a noise subset of each DNN model are constructed based on the clean subsets, so that semi-supervised training is performed on the two pre-trained DNN models by using the clean subset and the noise subset of the opposite side, and a clean subset DFExp is generated; and based on the clean subset DFExp and the initial noise data set, performing cyclic MixFix training on the two DNN models after semi-supervised training until the trained DNN models are output, and classifying the to-be-processed image. According to the invention, the accuracy and efficiency of image classification can be improved, and the applicability and robustness in a complex noise environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a method for cleaning noisy labels based on a vision language model and a fine-grained augmentation strategy. Background Art

[0002] In recent years, Deep Neural Networks (DNNs) have made breakthrough progress in many downstream tasks, such as classification, segmentation, and object detection. However, the performance of DNNs highly depends on large-scale and high-quality labeled datasets. In practical applications, due to the time-consuming and labor-intensive manual annotation process, the problem of noisy labels is inevitably introduced. These incorrect or inaccurate labels will mislead the training process of machine learning models, especially for deep neural networks, which may lead to a significant decline in model performance.

[0003] To address this challenge, Learning with Noisy Labels (LNL) has gradually become a research hotspot. Its goal is to develop methods that can effectively mitigate the impact of noisy labels, so that the model has better generalization ability on the test set. Currently, the research in the LNL field is very active, and the mainstream directions include label repair, loss function correction, sample selection, and robust loss function design, etc. Among them, the methods based on sample selection have attracted much attention due to their effectiveness in data cleaning. These methods usually select samples with smaller losses as clean samples, or use the memory effect to identify clean samples. Theoretically, DNNs will preferentially fit clean samples in the initial stage of training and then overfit noisy samples. Therefore, the losses of clean samples are usually smaller than those of noisy samples in the early stage.

[0004] However, these methods are prone to self-confirmation bias in the initial stage of training. Self-confirmation bias means that these methods rely entirely on the model discrimination ability learned from the noisy dataset to identify noisy labels. As the training error accumulates, it becomes difficult for the model to distinguish clean samples from noisy samples. To address the self-confirmation bias problem, Han et al. (Co-teaching: Robust Training of Deep Neural Networks with Extremely Noisy Labels) proposed a dual-network training framework, in which two models alternately detect noisy labels for each other, thus reducing the accumulation of self-confirmation errors. This method is called co-teaching. Since a high learning rate can inhibit the fitting of DNN to noisy labels, while a low learning rate may lead to overfitting, Zhang et al. (CJC-net: A cyclical training method with joint loss and co-teaching strategy net for deep learning under noisy labels) combined co-teaching with a cyclical learning rate adjustment strategy to further reduce the accumulation of self-confirmation errors and improve the accuracy of the model. Li et al. (DivideMix: Learning with Noisy Labels as Semi-supervised Learning) introduced semi-supervised learning (SSL) techniques to enhance the robustness of the model, thereby alleviating the impact of label noise. They also introduced Gaussian Mixture Models (GMM) to model the loss of each sample to reduce self-confirmation errors and filter noisy labels. This idea inspired a series of subsequent sample selection methods. For example, the DISC method (DISC: Learning from noisy labels via dynamic instance-specific selection and correction) introduced a confidence quantization DNN to represent the memory strength of samples and proposed a dynamic threshold strategy to classify samples into clean, difficult, and purified subsets, and then applied different regularization techniques to each subset. Zhang et al. (Learning with noisy labels using hyperspherical margin weighting) replaced the small loss criterion with the integrated area margin to improve the recognition accuracy of clean but difficult-to-learn samples and embedded it into the UNICON framework to further enhance performance.In addition, some methods reduce the impact of self-confirmation errors on sample selection by enhancing the robustness of the model. For example, Zhang Qian et al. (A Noise Label Learning Method Based on Class Balance and Cross-Merger Strategy) proposed a robustness enhancement strategy from cross to merger and made progress by combining the GMM-based balanced selection technique. Liu Jinghua et al. (Partial Label Feature Selection Method, Device, Equipment and Medium for Dynamic Flow Labels) dynamically determine the confidence that a candidate label is a clean label based on the probability distribution of samples, thereby achieving the correction of noisy labels.

[0005] Vision-Language Models (VLMs) have recently been widely applied to the noisy label detection task due to their rich knowledge reserve and support for context learning. Different from traditional methods that rely on the training model itself or shallow machine learning models (such as GMM, etc.), some studies have tried to introduce pre-trained VLMs to detect noisy labels, thereby improving the accuracy and efficiency of sample selection. For example, NoiseGPT (NoiseGPT: label noise detection and rectification through probability curvature) observed that in the Clip (Learning transferable visual models from natural language supervision) model, when the sample image matches its label, the output result is relatively stable; while when the image does not match the label, the output result will show significant fluctuations. This characteristic is used by NoiseGPT for sample selection, by detecting the stability of the output to distinguish clean samples and noisy samples. Similarly, the ClipCleaner (CLIPCleaner: cleaning noisy labels with clip) method inputs multiple prompts for each category of image pairs and uses Clip to calculate the confidence score to detect noisy samples. This method makes full use of the powerful cross-modal understanding ability of VLMs and can effectively identify some noisy labels. However, although the VLM-based method performs well in some scenarios, it has a significant problem: the distribution of the clean subset will be strongly affected by the inherent prior knowledge of the VLM, which in turn leads to large differences in the performance of the training model in different scenarios. Specifically, since VLMs have learned a large amount of category-related features and semantic information during the pre-training stage, these prior knowledge may cause the selected clean subset to be biased towards the sample types that the VLM is more familiar with or better at processing, while ignoring other samples that may be equally important. This not only limits the generalization ability of the model but also may lead to a performance decline.

[0006] In summary, although certain progress has been made in the existing methods for cleaning noisy labels, there are still many deficiencies. In particular, there is still room for improvement in terms of applicability and robustness in complex noise environments. Therefore, developing a more efficient and robust method for cleaning noisy labels has important research value and practical significance. Summary of the Invention

[0007] The present invention provides a method for cleaning noisy labels based on a vision-language model and a fine-grained augmentation strategy, which can overcome the problems existing in the existing sample learning method CLIPCleaner based on a vision-language model, such as poor performance under a high noise ratio, limited generalization ability in complex scenarios, and unstable performance in some scenarios due to over-reliance on the prior knowledge of the vision-language model.

[0008] The present invention provides the following technical solutions:

[0009] In a first aspect, there is provided a method for cleaning noisy labels based on a vision-language model and a fine-grained augmentation strategy, including the following steps:

[0010] S1: For the initial noisy dataset D, use the pre-trained vision-language model to screen out the clean subset D clip ;

[0011] S2: Use the initial noisy dataset D to pre-train two DNN models with the same structure and different initial weights. During the pre-training process, generate a state storage sequence corresponding to the training rounds and regarding the posterior probability for each DNN model and save it in the corresponding historical sequence Q m , m = 1, 2;

[0012] S3: Perform semi-supervised training on the two pre-trained DNN models with the fine-grained augmentation strategy; update the historical sequence Q m before each round of training, and based on the clean subset D clip construct the clean subset of each DNN model and the noisy subset to utilize the clean subsets and noisy subsets of each other to perform semi-supervised training on the two pre-trained DNN models to generate the clean subset D FExp ;

[0013] S4: Based on the clean subset D FExp and the initial noisy dataset D, perform cyclic MixFix training on the two DNN models after semi-supervised training until the trained DNN models are output to classify the images to be processed.

[0014] Optionally, step S1 specifically includes:

[0015] S1.1: For the initial noisy dataset (where x n represents the input image features, y n ∈ {1, 2, …, c} 1 represents the corresponding observed label, c represents the total number of dataset categories, and N represents the total number of samples), for each sample, construct a set of prompt words z n according to the category name corresponding to its label using the preset representative feature names;

[0016] S1.2: Use the pre-trained vision-language model to match each sample (x n , y n ) with the constructed prompt words z n to calculate the matching score s n , and use the matching score s n as the estimated value of the conditional probability of this sample (x n , y n );

[0017] S1.3: Normalize the conditional probability of each sample and perform sample confidence detection based on the preset threshold τ cons to form a clean subset O cons by screening samples that exceed the threshold τ cons ;

[0018] S1.4: Calculate the cross-entropy loss of each sample and calculate the posterior probability that each sample contains a clean label using a two-component mixture Gaussian model based on the cross-entropy loss ; Form a clean subset O by samples whose posterior probability of containing a clean label loss exceeds the preset threshold τ loss ;

[0019] S1.5: After taking the intersection of the clean subset O cons and the clean subset O loss , form the clean subset D clip screened by the vision-language model.

[0020] Optionally, step S2 specifically includes:

[0021] S2.1: Based on the initial noisy dataset D, pre-train two DNN models with the same structure and different initial weights using the cross-entropy loss function;

[0022] S2.2: For each DNN model, calculate the cross-entropy loss of each sample at the end of each training round

[0023] S2.3: Cross-entropy loss calculated based on each DNN model Use a two-component Gaussian mixture model to estimate the posterior probability that each sample contains a clean label μ k and σ k respectively represent the normalized mean and variance of the k-th GMM component, where k = {0, 1};

[0024] S2.4: Initialize the state sequences at the current training round t for the two DNN models respectively where is the state value of the n-th sample obtained using the m-th DNN model in the t-th training round;

[0025]

[0026] S2.5: Save the state sequences of each training round in the corresponding historical state sequences Q of the two DNN models in ascending order of training rounds m ;

[0027]

[0028] where T wu is the preset pre-training round number.

[0029] Optionally, step S3 specifically includes:

[0030] S3.1: At the start of the current training round t (T wu < t < T Fexp ) of semi-supervised training, update the historical sequence Q based on the posterior probability, and generate a clean subset m , a clean subset D clip and the updated historical sequence respectively for the two pre-trained DNN models and a noise subset

[0031] S3.2: For each DNN model, use the clean subset and the noise subset of the other party respectively for semi-supervised training in the current training round;

[0032] S3.3: If the training round t reaches the preset fine-grained expansion round number T Fexp , extract the subsequence q m with the largest number of state values of 1 from the last training rounds in the historical state sequences Q of the two DNN models m and q3-m and merge the subsequences q of the two DNN models m and q 3-m in which the samples with state 1 are combined, and the combined samples are the refined clean subset D generated by fine-grained augmentation FExp .

[0033] Optionally, step S3.1 specifically includes:

[0034] S3.1.1: Use the clean subset D clip as the initial clean subset of each DNN model;

[0035] S3.1.2: At the initial moment of the current training round, calculate the cross-entropy loss of all sample pairs in each DNN model, and use the two-component Gaussian mixture model to estimate the posterior probability that each sample contains a clean label;

[0036] S3.1.3: According to the value of the posterior probability, set the state value of the sample, and save the state value of all samples in the current training round as the state storage sequence of the current iteration round to the historical sequence Q m to update the historical state sequence of each DNN model in the current training round;

[0037] S3.1.4: According to the updated historical state sequence corresponding to each DNN model select the samples with state value 1 in a continuous set number of iteration times and put them into the initial clean subset to generate a clean subset Put the remaining samples into the noise subset .

[0038] Optionally, in step S3.2: During semi-supervised training, use the combined loss of the supervised loss L of the comprehensive Mixup entropy sup and the penalty loss L of the KL divergence p as the loss function L FExp ;

[0039] L FExp = L sup + L p

[0040]

[0041] where is the predicted value of the i-th newly input sample pair output by the m-th DNN model, and c is the total number of dataset categories; is the refined pseudo-label;

[0042]

[0043] Among them, λ is a random number between 0 and 1;

[0044] If the sample pair then keep the original observation label If the sample pair then generate a refined pseudo-label using the following formula

[0045]

[0046] Among them, is the input x n After data augmentation, it is the mean of the prediction results of the m-th DNN model; is the input x n After data augmentation, it is the mean of the prediction results of the (3 - m)-th DNN model.

[0047] Optionally, step S4 specifically includes:

[0048] The cyclic MixFix training of the m-th DNN model and the cyclic MixFix training of the (3 - m)-th DNN model; the steps of the cyclic MixFix training of the two DNN models are the same. Among them, when performing cyclic MixFix training on the m-th DNN model,

[0049] S4.1: If the current training round t is equal to the preset semi-supervised training round T Fexp , then use the clean subset D FExp as the clean subset D for this round of training c m , calculate the Mixup entropy loss for backpropagation training;

[0050] S4.2: If the current training round t satisfies T Fexp <t<T CMix , then according to the difference between the predicted value of the m-th DNN model and the original label, perform label renovation and weight coefficient setting on the samples; and combine the samples with the weight coefficient of 1 set by the m-th DNN model with the clean subset D FExp to generate the clean subset of the m-th DNN during the current training round t calculate the Mixup entropy loss for backpropagation training;

[0051] S4.3: If the current training round t satisfies t>T CMix , then the cyclic MixFix training is completed, and the trained DNN model is output for classifying the image to be processed.

[0052] Optionally, in step S4.2, setting the label renovation and weight coefficient for the sample according to the difference between the predicted value of the m-th DNN model and the original label specifically includes:

[0053] Using the m-th DNN to calculate the predicted value for each sample pair (x n , y n ) in the initial noise dataset D

[0054] According to the predicted value and the difference between the original label y n , and the preset threshold τ pr and τ rl , use the following formula to set the label renovation and weight coefficient for all samples;

[0055]

[0056] Wherein, is the weight coefficient, is the renovated label.

[0057] In a second aspect, a computer device is provided, including a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the steps of the noise label cleaning method based on the vision language model and the fine-grained expansion strategy described in any item of the first aspect are implemented.

[0058] In a third aspect, a computer-readable storage medium is provided for storing a computer program; when the computer program is executed by a processor, the steps of the noise label cleaning method based on the vision language model and the fine-grained expansion strategy described in any item of the first aspect are implemented.

[0059] Compared with the prior art, the beneficial effects of the present invention are:

[0060] The present invention uses the existing vision language model Clip to perform the initial clean sample screening operation, giving full play to the powerful discrimination ability and rich feature extraction ability of VLMs; secondly, by applying the pre-trained DNNs to the selected subset, the further fine expansion of its own discrimination ability is performed. The present invention aims to make up for the deviation caused by completely relying on the inherent prior knowledge of the VLM for noise label cleaning by introducing the specific features learned from the data by the trained DNNs, so as to obtain a more balanced and high-quality clean subset; finally, improve the robustness and classification performance of the trained DNN model, and improve the accuracy and efficiency of sample classification. It overcomes the problems existing in the existing sample learning method CLIPCleaner based on the vision language model, such as poor performance under high noise ratios, limited generalization ability in complex scenarios, and unstable performance in some scenarios due to over-reliance on the prior knowledge of the vision language model. Brief Description of the Drawings

[0061] Figure 1 is a flowchart of the noise label cleaning method based on the vision - language model and the fine - grained augmentation strategy of the present invention. Detailed Embodiments

[0062] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.

[0063] It should be noted that the term "including" in the description and claims of the present invention and any of its variations are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0064] As Figure 1 shown, a noise label cleaning method based on a vision - language model and a fine - grained augmentation strategy is provided, including the following steps:

[0065] S1: For the initial noise data set, use the pre - trained vision - language model to screen the clean subset D clip .

[0066] Step S1 specifically includes:

[0067] S1.1: For each sample in the initial noise data set (where x n represents the input image feature, y n ∈{1, 2, …, c} 1 represents the corresponding observed label, c represents the total number of data set categories, and N represents the total number of samples), according to the category name corresponding to its label, use the preset representative feature names to construct the prompt word set z n .

[0068] Specifically, y n is in its corresponding one - hot form. The observed labels of some samples are not equal to their true category labels and are regarded as noise labels; according to the template "A photo of {class name y n}, which is / has / can be {a representative feature of the category y n}." Use the preset J≥4 representative feature names to construct J≥4 prompt words, z n,j ∈zn Among them, each prompt word is z n,j , and z n is a set of prompt words, and the prompt words of each sample within the same category are exactly the same here.

[0069] S1.2: Use the pre-trained vision-language model to match each sample with the constructed prompt words, and calculate the matching score s n , and use the matching score s n as the estimated value of the conditional probability of this sample.

[0070] Specifically, the pre-trained vision-language model can adopt the vision-language model CLIP disclosed in the prior art.

[0071] The calculation formula of the matching score s n is:

[0072]

[0073] S1.3: Normalize the conditional probability of each sample, and perform sample confidence detection based on a preset threshold, so as to form a clean subset O cons .

[0074] The method of confidence detection can adopt the prior art, and the representation formula of the clean subset O cons is:

[0075]

[0076] Among them, is the j-th element of

[0077] S1.4: Calculate the cross-entropy loss of each sample and calculate the posterior probability that each sample contains a clean label using a two-component mixture Gaussian model based on the cross-entropy loss Form a clean subset O from the samples whose posterior probability of containing a clean label exceeds the preset threshold τ loss . loss .

[0078] Specifically, for each sample pair (x n , y n ), use the estimated conditional probability and the given label y n , and calculate the CE (Cross Entropy) loss according to formula (3).

[0079]

[0080] CE loss of all samples based on calculation Calculate the posterior probability containing the clean label for each sample within each category using a two-component Gaussian mixture model (GMM) wherein and respectively represent the normalized mean and variance of the k-th GMM component, and k = {0, 1}. When the given label is provided, for these two GMM components, the present invention selects the output result of the Gaussian distribution where the smaller component is located as the posterior probability of each sample pair (x n , y n )

[0081] Based on a preset threshold τ loss Filter out the samples greater than or equal to the threshold value according to formula (4) to form a clean subset O loss .

[0082]

[0083] S1.5: After taking the intersection of the clean subset O cons and the clean subset O loss , form a clean subset D for visual language model screening clip .

[0084] D clip = O loss ∩ O cons (5)

[0085] S2: Use the initial noisy dataset D to pre-train two DNN models with the same structure and different initial weights. During the pre-training process, generate a state storage sequence corresponding to the training rounds and regarding the posterior probability for each DNN model and save it in the corresponding historical sequence Q m , m = 1, 2

[0086] Step S2 specifically includes:

[0087] S2.1: Based on the initial noisy dataset D, use the cross-entropy loss function to pre-train two DNN models with the same structure and different initial weights

[0088] S2.2: For each DNN model, calculate the cross-entropy loss of each sample at the end of each training round

[0089]

[0090] S2.3: Cross-entropy loss calculated based on each DNN model Use a two-component Gaussian mixture model to estimate the posterior probability that each sample contains a clean label μ k and σ k respectively represent the normalized mean and variance of the k-th GMM component, and k = {0, 1}

[0091] In the present invention, the output result of the Gaussian distribution where the component with the smaller mean is located is also selected as the posterior probability of each sample pair (x n , y n ) at this time

[0092] S2.4: Initialize the state sequences of the two DNN models at the current training round t Among them, is the state value of the n-th sample obtained by using the m-th DNN model in the t-th training round;

[0093] Specifically, according to formula (7), the corresponding states of samples greater than 0.5 are set to 1, and the corresponding states of the remaining samples are set to 0

[0094]

[0095] S2.5: In the order of increasing training rounds, save the state sequence of each training round in the corresponding historical state sequences Q of the two DNN models m ;

[0096]

[0097] Among them, T wu is the preset number of pre-training rounds

[0098] S3: Perform semi-supervised training with a fine-grained augmentation strategy on the two pre-trained DNN models; update the historical sequence Q before each round of training m , and based on the clean subset D clip construct the clean subset of each DNN model and the noise subset to use the clean subsets and noise subsets of each other to perform semi-supervised training on the two pre-trained DNN models to generate the clean subset D FExp .

[0099] Step S3 specifically includes:

[0100] S3.1: At the current training round t (T wu < t < TFexp )'s starting moment, update the historical sequence Q based on the posterior probability m , based on the clean subset D clip and the updated historical sequence Q″ m respectively generate a clean subset for two pre-trained DNN models and a noisy subset

[0101] Step S3.1 specifically includes:

[0102] S3.1.1: Use the clean subset D clip as the initial clean subset for each DNN model.

[0103] S3.1.2: At the initial moment of the current training round, calculate the cross-entropy loss of all sample pairs for each DNN model, and use a two-component Gaussian mixture model to estimate the posterior probability that each sample contains a clean label.

[0104] S3.1.3: According to the value of the posterior probability, set the status value of the sample, and save the status values of all samples in the current training round as the status storage sequence of the current iteration round to the historical sequence Q m to update the historical state sequence of each DNN model in the current training round.

[0105] The processes of Step S3.1.2 and Step S3.1.3 are similar to the processes of Step S2.2 - Step S2.5, that is, the calculation method of the cross-entropy loss refers to formula (6), and use a 2-component GMM to calculate the cross-entropy loss for each model to estimate the posterior probability that the sample (x n , y n ) contains a clean label Re-initialize a status storage sequence for the current t-th epoch for the two models and set the corresponding status of samples greater than greater than 0.5 to 1, and the status of the remaining samples to 0. Finally, in the order of increasing t according to formula (8), save the status sequence to the historical state sequence Q corresponding to the m-th model respectively m .

[0106] That is, the updated historical state sequence Q″ corresponding to each DNN model m ,

[0107] S3.1.4: According to the updated historical state sequence Q″ corresponding to each DNN model m , Select the samples with the state value of 1 in the consecutive set number of iteration times and put them into the initial clean subset to generate a clean subset Put the remaining samples into the noise subset in it

[0108] Specifically, according to the historical state sequence corresponding to each DNN, select the samples with the state value of 1 in the last consecutive v epochs according to formula (9) and put them into the clean subset Put the remaining samples into the noise subset

[0109]

[0110] S3.2: For each DNN model, use the clean subset of the other party and the noise subset to perform semi-supervised training for the current training round

[0111] That is, for the m-th DNN, use the clean subset generated by the (3 - m)-th DNN and the noise subset to perform conventional semi-supervised training. The semi-supervised training methods of the two DNN models are the same

[0112] Among them, during the semi-supervised training process, use the combined supervision loss L of the Mixup entropy sup and the penalty loss L of the KL divergence p as the joint loss as the loss function L FExp ;

[0113] L FExp = L sup + L p (10)

[0114]

[0115] Among them, is the predicted value of the i-th new input sample pair output by the m-th DNN model, c is the total number of dataset categories; is the refined pseudo label;

[0116] The new input sample is generated according to formula (13):

[0117]

[0118] Among them, λ is a random number between 0 and 1

[0119] If the sample pair then keep the original observed label If the sample pair then use formula (14) to generate refined pseudo-labels

[0120]

[0121] where is the input x n After data augmentation, it is the mean of the prediction results of the m-th DNN model; is the input x n After data augmentation, it is the mean of the prediction results of the (3 - m)-th DNN model. The way of data augmentation can refer to the prior art.

[0122] S3.3: If the training epoch t reaches the preset fine-grained augmentation epoch T Fexp then, respectively, extract from the historical state sequences Q of the two DNN models m the last number of training epochs to extract the subsequence q that contains the largest number of state values of 1 m and q 3-m , and merge the samples with state 1 in the subsequences q m and q 3-m of the two DNN models. The merged samples are the refined clean subset D generated by fine-grained augmentation FExp .

[0123] As the training epoch t increments by 1, if the training epoch t reaches the preset fine-grained augmentation epoch T Fexp then, according to the historical state sequences Q corresponding to the two DNNs m , m = 1, 2, extract from the last number of epochs the subsequence q that contains the largest number of states of 1 m and q 3-m , and merge the samples with state 1 in the two subsequences as the refined clean subset D generated by fine-grained augmentation FExp .

[0124] S4: Based on the clean subset D FExp and the initial noise dataset D, perform Cyclical MixFix training (CMix) on the two DNN models after semi-supervised training until the trained DNN models are output, and classify the images to be processed.

[0125] Step S4 specifically includes: the Cyclical MixFix training of the m-th DNN model and the Cyclical MixFix training of the (3 - m)-th DNN model; the steps of the Cyclical MixFix training of the two DNN models are the same.

[0126] Among them, when performing cyclic MixFix training on the m-th DNN model,

[0127] S4.1: If the current training epoch is equal to the preset semi-supervised training epoch T Fexp , then the clean subset D FExp is used as the clean subset for this round of training Calculate the Mixup entropy loss for backpropagation training.

[0128] Specifically, during the cyclic MixFix training process of the m-th DNN in the current t-th epoch, the clean subset D c m generated by the m-th DNN is D FExp ; for the m-th DNN, based on the clean subset generated by the (3 - m)-th DNN Calculate the supervised loss based on Mixup entropy according to the methods of formulas (13)-(14) for backpropagation training; finally, for the (3 - m)-th DNN, based on the clean subset generated by the m-th DNN Calculate the Mixup entropy in the same way for backpropagation training.

[0129] S4.2: If the current training epoch t satisfies T Fexp < t < T CMix , then according to the difference between the predicted value of the m-th DNN model and the original label, perform label refurbishment and weight coefficient setting on the samples; and combine the samples with a sample weight coefficient of 1 set by the m-th DNN model with the clean subset D FExp to generate the clean subset of the m-th DNN during the current training epoch t Calculate the Mixup entropy loss for backpropagation training.

[0130] Specifically, it includes: using the m-th DNN to calculate the predicted value for each sample pair (x n , y n ) in the initial noise dataset D and according to the predicted value and the original label y n , and the preset thresholds τ pr and τ rl , use formula (15) to perform label refurbishment and weight coefficient setting for all samples;

[0131]

[0132] Among them, is the weight coefficient, is the refurbished label.

[0133] That is, for and for the sample pair where y n = j, after refurbishment, its label remains the original label y n and the weight coefficient For and for the sample pair where y n ≠ j, after refurbishment, its label is set to j and the weight coefficient The weight coefficient of the remaining samples

[0134] As shown in formula (16), the weight coefficient of the sample calculated by the m-th DNN with a value of 1 is merged with the refined clean subset D generated by fine-grained augmentation FExp to generate the clean subset of the m-th DNN in the cyclic MixFix training process of the current t-th epoch

[0135]

[0136] where Ι(A) is the indicator function, which returns 1 if and only if A is True, otherwise returns 0.

[0137] For the m-th DNN, based on the clean subset generated by the (3 - m)-th DNN According to the methods of formulas (11), (13)-(14), calculate the Mixup entropy for backpropagation training. For the (3 - m)-th DNN, based on the clean subset generated by the m-th DNN According to the methods of formulas (11), (13)-(14), calculate the Mixup entropy in the same way for backpropagation training.

[0138] S4.3: If the current training round t satisfies t > T CMix , then the cyclic MixFix training is completed, and the trained DNN model is output for classifying the image to be processed.

[0139] That is, when the training round t satisfies t < T wu , the DNN model is in the pre-training stage. When T wu < t < T Fexp , the DNN model is in the semi-supervised training stage of implementing the fine-grained augmentation strategy. When T Fexp ≤ t < T CMix , it is in the cyclic MixFix training stage. T wu is the set number of pre-training rounds, and T Fexp is the set number of semi-supervised training rounds for implementing the fine-grained augmentation strategy, and T CMix is the set number of cyclic MixFix training rounds.

[0140] In another embodiment, the present invention provides a computer device, including a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the steps of the above-mentioned noise label cleaning method based on the vision language model and the fine-grained expansion strategy are implemented.

[0141] For a more specific process of the above method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0142] In another embodiment, the present invention provides a computer-readable storage medium for storing a computer program; when the computer program is executed by a processor, the steps of the above-mentioned noise label cleaning method based on the vision language model and the fine-grained expansion strategy are implemented.

[0143] For a more specific process of the above method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0144] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices and storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0145] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present invention.

[0146] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.

Claims

1. A noise label cleaning method based on visual language model and fine-grained expansion strategy, characterized in that: The following steps are involved: S1: For the initial noisy dataset D, use the pre-trained visual language model to filter the clean subset D clip ; S2: Use the initial noise dataset D to pre-train two DNN models with the same structure and different initialization weights. During the pre-training process, a state storage sequence corresponding to the training round and about the posterior probability is generated for each DNN model. And save it in the corresponding historical sequence Q m ,m=1,2; S3: Perform semi-supervised training of the fine-grained expansion strategy on the two pre-trained DNN models; update the historical sequence Q before each round of training m , and based on the clean subset D clip Build a clean subset of each DNN model and noise subset By using the clean subset and noise subset of each other, the two pre-trained DNN models are semi-supervised trained to generate the clean subset D FExp ; S4: Based on the clean subset D FExp and the initial noise dataset D, perform cyclic MixFix training on the two DNN models after semi-supervised training until the trained DNN model is output to classify the processed images.

2. The noise label cleaning method based on visual language model and fine-grained expansion strategy according to claim 1 is characterized in that: Step S1 specifically includes: S1.1: For the initial noisy dataset (where x n Represents the input image features, y n ∈{1,2,…,c} 1 For each sample in the dataset, we construct a prompt word set z using the preset representative feature name according to the category name corresponding to its label. n ; S1.2: Use the pre-trained visual language model to train each sample (x n ,y n ) and the constructed prompt word z n Perform matching and calculate matching score s n , and the matching score s n As the sample (x n ,y n )The estimated value of the conditional probability S1.3: Normalize the conditional probability of each sample and calculate the probability of each sample based on the preset threshold τ cons Perform sample confidence detection to filter samples that exceed the threshold τ cons The samples constitute a clean subset O cons ; S1.4: Calculate the cross entropy loss for each sample And based on the cross entropy loss Use a two-component Gaussian mixture model to calculate the posterior probability that each sample contains a clean label The posterior probability of containing the clean label Exceeding the preset threshold τ loss The samples constitute a clean subset O loss ; S1.5: Clean subset O cons and the clean subset O loss After taking the intersection, the visual language model is constructed to filter the clean subset D clip .

3. The noise label cleaning method based on visual language model and fine-grained expansion strategy according to claim 1 is characterized in that: Step S2 specifically includes: S2.1: Based on the initial noise dataset D, two DNN models with the same structure and different initialization weights are pre-trained using the cross entropy loss function; S2.2: For each DNN model, calculate the cross entropy loss for each sample at the end of each training round S2.3: Cross entropy loss calculated based on each DNN model Use a two-component Gaussian mixture model to estimate the posterior probability that each sample contains a clean label μ k and σ k Respectively represent the normalized mean and variance of the kth GMM component and k={0,1}; S2.4: Initialize the state sequence at the current training round t for the two DNN models respectively in, The state value of the nth sample obtained using the mth DNN model in the tth training round; S2.5: In the order of increasing training rounds, save the state sequence of each training round in the historical state sequence corresponding to the two DNN models respectively. m ; Among them, T wu is the preset pre-training round.

4. The noise label cleaning method based on visual language model and fine-grained expansion strategy according to claim 1 is characterized in that: Step S3 specifically includes: S3.1: At the current training round t(T wu <t<T Fexp ) at the starting time, update the historical sequence Q based on the posterior probability m , based on the clean subset D clip and the updated historical sequence Q" m Generate clean subsets for the two pre-trained DNN models respectively and noise subset S3.2: For each DNN model, use the clean subset of the other and noise subset Perform semi-supervised training for the current training round; S3.3: If the training round t reaches the preset fine-grained expansion round T Fexp When the historical state sequence Q of the two DNN models is m reciprocal Extract the subsequence q with the largest number of state values ​​1 in the training rounds m and q 3-m , and the subsequences q of the two DNN models are m and q 3-m The samples with status 1 in the above example are merged, and the merged samples are the refined clean subset D generated by fine-grained expansion. FExp .

5. The noise label cleaning method based on visual language model and fine-grained expansion strategy according to claim 4 is characterized in that: Step S3.1 specifically includes: S3.1.1: Clean subset D clip As the initial clean subset for each DNN model; S3.1.2: At the beginning of the current training round, calculate the cross entropy loss of all sample pairs in each DNN model, and use a two-component Gaussian mixture model to estimate the posterior probability that each sample contains a clean label; S3.1.3: According to the value of the posterior probability, set the state value of the sample, and save the state values ​​of all samples in the current training round as the state storage sequence of the current iteration round to the historical sequence Q m In the above example, the historical state sequence of each DNN model in the current training round is updated; S3.1.4: According to the historical state sequence Q corresponding to each updated DNN model m , Select samples whose state values ​​are all 1 in a set number of iterations and put them into the initial clean subset to generate a clean subset The remaining samples are put into the noise subset middle.

6. The noise label cleaning method based on visual language model and fine-grained expansion strategy according to claim 4 is characterized in that: Step S3.2: During the semi-supervised training, the supervised loss L using the comprehensive Mixup entropy is used. sup and the penalty loss L of KL divergence p The joint loss of is used as the loss function L FExp ; L FExp =L sup +L p in, is the i-th new input sample pair The predicted value output by the mth DNN model, c is the total number of categories in the data set; is a refined pseudo-label; in, λ is a random number between 0 and 1; If the sample pair Keep the original observation label If the sample pair Then use the following formula to generate refined pseudo labels in, For input x n The mean of the prediction results of the mth DNN model after data enhancement; For input x n The mean of the prediction results of the 3-mth DNN models after data augmentation.

7. The noise label cleaning method based on visual language model and fine-grained expansion strategy according to claim 1 is characterized in that: Step S4 specifically includes: The cyclic MixFix training of the mth DNN model and the cyclic MixFix training of the 3rd-mth DNN models; the steps of the cyclic MixFix training of the two DNN models are consistent, among which, when the cyclic MixFix training of the mth DNN model is performed, S4.1: If the current training round t is equal to the preset semi-supervised training round T Fexp , then the clean subset D FExp As a clean subset for this round of training Calculate Mixup entropy loss for backpropagation training; S4.2: If the current training round t satisfies T Fexp <t<T CMix , then according to the difference between the predicted value of the mth DNN model and the original label, the sample is labeled and the weight coefficient is set; and the sample with the sample weight coefficient set by the mth DNN model as 1 is compared with the clean subset D FExp Merge to generate a clean subset of the mth DNN in the current training round t Calculate Mixup entropy loss for backpropagation training; S4.3: If the current training round t satisfies t>T CMix , then the MixFix training cycle is completed, and the trained DNN model is output for classifying the processed images.

8. The noise label cleaning method based on visual language model and fine-grained expansion strategy according to claim 7 is characterized in that: In step S4.2, the label renovation and weight coefficient setting of the sample are performed according to the difference between the predicted value of the mth DNN model and the original label, which specifically includes: Use the mth DNN to generate the initial noise dataset D for each sample pair (x n ,y n ) Calculate the predicted value According to the predicted value and the original label y n The difference between pr and τ rl , use the following formula to update labels and set weight coefficients for all samples; in, is the weight coefficient, This is a refurbished label.

9. A computer device, characterized in that: The method comprises a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the method implements the steps of the noise label cleaning method based on the visual language model and the fine-grained expansion strategy as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: Used to store computer programs; when the computer programs are executed by the processor, the steps of the noise label cleaning method based on the visual language model and fine-grained expansion strategy described in any one of claims 1-8 are implemented.

Citation Information

Cited By

  • Method for carrying out noise label learning from code annotation data

    CN121117590A