Model Training Method, Apparatus, Device, and Medium for Long-Tail Noise
By using the text encoder and image encoder of the visual language model on the long-tail noise data set, the label inconsistency and calibrate feature deviations are corrected, which improves the classification accuracy of the model in high-noise environments, and solves the problem of low classification accuracy in long-tail noise tag learning.
Patent Information
- Application Number
- CN202510345862.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The existing long-tail noise label learning method ignores the impact of different noise rates on model training in high-noise environments, resulting in low classification accuracy of head and tail class samples in high-noise scenarios.
The text encoder and image encoder in the visual language model are used to correct label inconsistencies using text prompt words, and the supervision start-stop state is determined by calculating the similarity between text features and image features, and the fine-tuning module is updated based on the target loss function to calibrate the deviation between the learned features and the observed labels.
The classification accuracy of head and tail samples in high noise scenarios is improved, and the robustness of the model is enhanced, especially in high noise ratio scenarios.
Smart Images

Figure CN119888412B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and particularly relates to a method, device, equipment, and medium for model training for long-tail noise. Background Art
[0002] In the field of computer vision, model training is usually based on large-scale public datasets. However, in the real world, datasets often have the problem of class imbalance, showing the characteristics of a long-tail distribution, that is, the head classes contain most of the samples, while the tail classes have insufficient samples; the datasets also have the problem of image mislabeling, called noise labels, and creating a balanced dataset with correctly labeled classes is costly. Therefore, it is usually still necessary to perform model training on datasets with both long-tail distribution and label noise. To solve these problems, a long-tail noise label learning method has been introduced.
[0003] However, the existing long-tail noise label learning methods ignore the impact of different noise rates in the data on model training. In a high-noise environment, the long-tail noise label learning method will be insufficient because the noise labels will damage the reliability of the samples and blur the difference between the noise samples and the tail-class samples, resulting in a low classification accuracy for the head-class and tail-class samples in a high-noise scenario.
[0004] Therefore, the existing technology has defects and needs to be improved and developed. Summary of the Invention
[0005] The present application provides a method, device, equipment, and medium for model training for long-tail noise to solve the technical problem that the classification accuracy of the head-class and tail-class samples is relatively low in a high-noise scenario in the related art.
[0006] To achieve the above object, the present application adopts the following technical solutions:
[0007] A method for model training for long-tail noise, wherein the method includes:
[0008] Input an input image, a text prompt word corresponding to the input image, and an observation label in a target long-tail noise dataset into a pre-trained vision-language model, where the vision-language model includes a text encoder and an image encoder, and a fine-tuning module is provided in the image encoder;
[0009] In the vision-language model, the text prompt word is processed by the text encoder to obtain a text feature, the input image is processed by the image encoder to obtain an image feature, and the input image is processed by the image encoder with the fine-tuning module to obtain the original output values for each category;
[0010] Calculate the text-image similarity between the text features and the image features, and obtain a text prediction label based on the text-image similarity;
[0011] Determine a supervised start / stop state based on the text prediction label and the observed label, determine an objective loss function based on the supervised start / stop state and the original output value, and update the fine-tuning module based on the objective loss function to obtain a trained vision-language model.
[0012] In one embodiment of the present application, determining a supervised start / stop state based on the text prediction label and the observed label includes:
[0013] Calculate the coincidence rate between the text prediction label and the observed label;
[0014] Obtain a preset control threshold, compare the coincidence rate with the control threshold, and obtain a comparison result;
[0015] Determine the supervised start / stop state according to the comparison result.
[0016] In one embodiment of the present application, determining the supervised start / stop state according to the comparison result includes:
[0017] If the comparison result is that the coincidence rate is greater than the control threshold, the supervised start / stop state is the supervised stop state;
[0018] If the comparison result is that the coincidence rate is less than or equal to the control threshold, the supervised start / stop state is the supervised start state.
[0019] In one embodiment of the present application, determining an objective loss function based on the supervised start / stop state and the original output value includes:
[0020] If the supervised start / stop state is the supervised start state, calculate the information divergence based on the text-image similarity and the original output value, and use the information divergence as the first loss function;
[0021] Obtain an image prediction label based on the original output value, and calculate a second loss function according to the image prediction label and the observed label;
[0022] Obtain the objective loss function based on the first loss function and the second loss function.
[0023] In one embodiment of the present application, obtaining the objective loss function based on the first loss function and the second loss function includes:
[0024] Set a first weight for the first loss function and a second weight for the second loss function, and the sum of the first weight and the second weight is 1;
[0025] Weight the first loss function and the second loss function based on the first weight and the second weight, and sum them to obtain a target loss function.
[0026] In an embodiment of the present application, determining the target loss function based on the supervised start-stop state and the original output value further includes:
[0027] If the supervised start-stop state is the supervised stop state, obtain an image prediction label based on the original output value;
[0028] Calculate a second loss function according to the image prediction label and the observation label, and use the second loss function as the target loss function.
[0029] In an embodiment of the present application, obtaining an image prediction label based on the original output value includes:
[0030] Use an activation function to convert the original output value into a probability distribution for each category, and obtain an image prediction label according to the probability distribution.
[0031] The present application also provides a model training device for long-tail noise, wherein the device includes:
[0032] An image input module, configured to input an input image, a text prompt corresponding to the input image, and an observation label in a target long-tail noise dataset into a pre-trained vision-language model, where the vision-language model includes a text encoder and an image encoder, and a fine-tuning module is provided in the image encoder;
[0033] A feature extraction module, configured to, in the vision-language model, obtain a text feature after the text prompt is processed by the text encoder, obtain an image feature after the input image is processed by the image encoder, and obtain an original output value for each category after the input image is processed by the image encoder with the fine-tuning module;
[0034] A text prediction module, configured to calculate the text-image similarity between the text feature and the image feature, and obtain a text prediction label based on the text-image similarity;
[0035] An update module, configured to determine a supervised start-stop state based on the text prediction label and the observation label, determine a target loss function based on the supervised start-stop state and the original output value, and update the fine-tuning module based on the target loss function to obtain a trained vision-language model.
[0036] The present application also provides a device, which includes: a memory, a processor, and a model training program for long-tail noise stored on the memory and executable on the processor. When the model training program for long-tail noise is executed by the processor, the steps of the model training method for long-tail noise described above are implemented.
[0037] The present application also provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the model training method for long-tail noise described above.
[0038] Advantages of the present invention: In the method of the embodiment of the present invention, the input image, the text prompt word corresponding to the input image, and the observation label in the target long-tail noise dataset are input into a pre-trained vision-language model, which includes a text encoder and an image encoder, and a fine-tuning module is provided in the image encoder; in the vision-language model, the text prompt word is processed by the text encoder to obtain a text feature, the input image is processed by the image encoder to obtain an image feature, and the input image is processed by the image encoder with the fine-tuning module to obtain the original output values for each category; calculate the text-image similarity between the text feature and the image feature, obtain a text prediction label based on the text-image similarity; determine a supervision start-stop state based on the text prediction label and the observation label, determine a target loss function based on the supervision start-stop state and the original output value, and update the fine-tuning module based on the target loss function to obtain a trained vision-language model. The present application determines whether text-image alignment prior assistance supervision is required by evaluating the difference between the text prediction label and the observation label, and calibrates the deviation between the learned feature and the observation label, thereby improving the classification accuracy of head-class and tail-class samples in a high-noise scenario. Description of the Drawings
[0039] Figure 1 is a flowchart of a preferred embodiment of the model training method for long-tail noise in the present invention.
[0040] Figure 2 is a model logic block diagram of the model training method for long-tail noise in the present invention.
[0041] Figure 3 is a functional principle block diagram of a preferred embodiment of the model training device for long-tail noise in the present invention.
[0042] Figure 4 is a functional principle block diagram of a preferred embodiment of the device in the present invention. Detailed Embodiments
[0043] To make the objectives, technical solutions and advantages of the present invention more clear and definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0044] The long-tail noise label learning methods can be classified into the following categories: (1) emphasizing the importance of different samples through reweighting or regularization; (2) selecting clean samples according to specific criteria; (3) developing improved representation learning methods. The above methods can effectively improve the robustness of the model on long-tail noise label data. However, they ignore the impact of different noise ratios on model training, resulting in poor performance in high-noise scenarios. It should be noted that low-level label noise has a relatively small impact on model performance. Therefore, it is necessary to specifically handle high-noise ratio scenarios, in which unreliable labels constitute one of the main problems in long-tail data noise label learning. In this case, noise labels introduce a large number of misleading supervision signals, which in turn lead to the inability to effectively distinguish noise samples from clean samples, or the inability to effectively improve feature representations, resulting in the accumulation of feature learning biases and the amplification of the problems of label noise and class imbalance.
[0045] In the prior art, some long-tail noise label learning models ignore the impact of different noise rates in the data on model training and will have deficiencies in high-noise environments because noise labels will undermine the reliability of samples and blur the distinction between noise samples and tail class samples, resulting in unsatisfactory performance in high-noise situations; some long-tail noise label learning models only consider the class imbalance problem in long-tail learning and do not consider the problem of a large number of noise labels in the data set, resulting in a sharp decline in model performance when both the class imbalance problem and the noise label problem exist.
[0046] When the present invention uses large model fine-tuning, it simultaneously considers the class imbalance problem and the noise label problem, takes different correction measures by evaluating the noise rate of the data, pays special attention to the processing in high-noise scenarios, and at the same time improves the classification accuracy of both the head classes and the tail classes in high-noise scenarios.
[0047] The following describes the model training method, device, equipment, and medium for long-tail noise in the embodiments of the present application with reference to the accompanying drawings. Aiming at the problem that the classification accuracy of head class and tail class samples is relatively low in the high-noise scenario in the related art mentioned in the above background art, the present application provides a model training method for long-tail noise. In this method, the input image, the text prompt word corresponding to the input image, and the observation label in the target long-tail noise dataset are input into a pre-trained vision-language model, which includes a text encoder and an image encoder, and a fine-tuning module is set in the image encoder; in the vision-language model, the text prompt word is processed by the text encoder to obtain text features, the input image is processed by the image encoder to obtain image features, and the input image is processed by the image encoder with the fine-tuning module to obtain the original output values for each category; calculate the text-image similarity between the text features and the image features, and obtain the text prediction label based on the text-image similarity; determine the supervision start-stop state based on the text prediction label and the observation label, determine the target loss function based on the supervision start-stop state and the original output values, and update the fine-tuning module based on the target loss function to obtain the trained vision-language model. The present application determines whether text-image alignment prior auxiliary supervision is required by evaluating the difference between the text prediction label and the observation label, and calibrates the deviation between the learned features and the observation label, thereby improving the classification accuracy of head class and tail class samples in the high-noise scenario.
[0048] Please refer to Figure 1 , the model training method for long-tail noise according to the embodiments of the present invention includes the following steps:
[0049] Step S100: Input the input image, the text prompt word corresponding to the input image, and the observation label in the target long-tail noise dataset into a pre-trained vision-language model, which includes a text encoder and an image encoder, and a fine-tuning module is set in the image encoder.
[0050] Specifically, the long-tail noisy dataset includes input images, text prompts, and observation labels, where the observation labels contain class information. Among them, one text prompt corresponds to one observation label, and multiple images can correspond to the same text prompt and observation label. The observation labels may have incorrect class information, while the text prompts do not have the problem of incorrect marking. Therefore, since the class information in the long-tail noisy dataset may be inconsistent with the corresponding images, the embodiments of the present application use the text prompts corresponding to the observation labels to correct the inconsistency. The visual language model can be the CLIP model, which is a large pre-trained text-image model using contrastive learning. The fine-tuning module adds an adapter module to the Transformer layer of the image encoder for fine-tuning specific tasks, which can be called "adapter former".
[0051] That is to say, although the observation labels do not match the input images, they still contain class information, that is, text prompts. The present application uses the text prompts to correct the label-image inconsistency in the long-tail noisy label data. Specifically, the present application uses the text encoder in the pre-trained visual language model to obtain text-based predictions and uses this text-image alignment prior to correct the label-image inconsistency. This supervision signal called weak teacher supervision (WTS) is not always accurate. Therefore, the present application evaluates the difference between the text prediction labels and the observation labels to decide when to activate the weak teacher supervision. By calibrating the deviation between the learned features and the labels, the weak teacher supervision ultimately enhances the robustness of the model.
[0052] As Figure 1 shown, the model training method for long-tail noise further includes the following steps:
[0053] Step S200: In the visual language model, the text prompt is processed by the text encoder to obtain text features, the input image is processed by the image encoder to obtain image features, and the input image is processed by the fine-tuning module to obtain the raw output values for each category.
[0054] Specifically, the raw output values (logits) are the linear outputs of the model for each category, without being converted into a probability distribution. These output values can be regarded as the preliminary judgments or raw output values of the model for each category. The raw output values will be converted into a probability distribution after passing through an activation function (softmax). The text encoder of the visual language model can provide text embeddings for the labels of the categories , and the image encoder can provide image embeddings for the input images , and by comparing and and and The similarity between them can be used to derive text prediction labels from label-based text prompts. The calculation formula for text prediction labels is as follows:
[0055] ;
[0056] ;
[0057] Among them, is the text prediction label of the input image , i is the image number, T represents the text prompt, represents a certain category, represents the text feature of the text prompt corresponding to the category , represents the text-image similarity between the image feature of the input image and the text feature of the text prompt corresponding to the category , represents the number of categories.
[0058] As Figure 1 shown, the model training method for long-tail noise further includes the following steps:
[0059] Step S300, calculate the text-image similarity between the text feature and the image feature, and obtain the text prediction label based on the text-image similarity;
[0060] Specifically, the text prediction label is also inaccurate. For example, its accuracy on the CIFAR-100 test set is only 64.4%, which indicates that the first loss function introduces another form of label noise. To solve this problem, this application uses the text-image similarity between text and image to provide additional supervision. Because it can not only provide a discrete one-hot optimization target, but also provide information about the inter-class relationship. Specifically, the activation function (softmax) is used to convert the text-image similarity into a probability. The formula is:
[0061] ;
[0062] Among them, represents a certain category, represents the probability that the input image belongs to the category , represents the th dimension, that is, the th category. One category corresponds to one text prompt, represents the image feature of the input image and the text feature of the text prompt in the th dimension.
[0063] This application uses the text encoder in a pre-trained vision-language model (VLM) to obtain text-based predictions, and uses this text-image alignment prior to correct the inconsistency between labels and images.
[0064] As Figure 1 shown, the model training method for long-tail noise further includes the following steps:
[0065] Step S400: Determine the supervision start-stop state based on the text prediction label and the observed label, determine the target loss function based on the supervision start-stop state and the original output value, and update the fine-tuning module based on the target loss function to obtain a trained vision-language model.
[0066] However, using the text-image alignment prior as a supervision signal is not always accurate. Therefore, the embodiments of this application evaluate the difference between the text prediction label predicted by the text encoder and the observed label to decide whether to activate this supervision.
[0067] In the embodiments of this application, determining the supervision start-stop state based on the text prediction label and the observed label includes:
[0068] Calculate the overlap rate between the text prediction label and the observed label;
[0069] Obtain a preset control threshold, compare the overlap rate with the control threshold, and obtain a comparison result;
[0070] Determine the supervision start-stop state according to the comparison result.
[0071] Specifically, since the weak teacher is not always accurate and the observed label is reliable in a low-noise scenario, the observed label can be directly used without additional processing in a low-noise scenario. However, since the proportion of noisy labels cannot be determined in advance, an indicator is needed to evaluate when the teacher provides effective supervision. The overlap rate (OR) between the text prediction label and the observed label can be used as this indicator:
[0072] ;
[0073] where represents the text prediction label of a batch of data, represents the observed label of a batch of data.
[0074] This application determines the difference between the text prediction label and the observed label by calculating the overlap rate between the text prediction label and the observed label.
[0075] By training on the target long-tail noise dataset, this application can ultimately obtain a model that can effectively learn long-tail data with noisy labels and has a relatively high classification accuracy for each category. That is to say, even if the training set has a long-tail distribution and noisy labels, a model with a relatively high classification accuracy can still be trained.
[0076] In the embodiment of this application, determining the supervision start-stop state according to the comparison result includes:
[0077] If the comparison result is that the coincidence rate is greater than the control threshold, the supervision start-stop state is the supervision deactivation state;
[0078] If the comparison result is that the coincidence rate is less than or equal to the control threshold, the supervision start-stop state is the supervision activation state.
[0079] Specifically, if the text prediction label predicted by the pre-trained text encoder deviates greatly from the observed label, it means that the text prediction label contains rich and relatively accurate information. Therefore, the text prediction label is used to guide the model training. Existing long-tail learning methods can all effectively apply the method of this application. Since text-based labels usually have lower accuracy than directly fine-tuning the image encoder, the text encoder is regarded as a "weak teacher", and the method of this application is called weak teacher supervision (WTS). Multiple experiments conducted on benchmark datasets with various types of noisy labels and an inherent long-tail distribution show that the weak teacher supervision proposed in this application improves the performance of the strong student, especially in high-noise ratio scenarios. The overall structural block diagram of weak teacher supervision is as Figure 2 shown.
[0080] When the coincidence rate OR is relatively high, it indicates that the text prediction label and the observed label are basically the same, and the information provided by weak teacher supervision is limited. Therefore, weak teacher supervision is selected to be deactivated. On the contrary, when OR is low, the observed label and the text prediction label (visual language alignment prior) are significantly different, indicating that auxiliary supervision is needed.
[0081] In the embodiment of this application, determining the target loss function based on the supervision start-stop state and the original output value includes:
[0082] If the supervision start-stop state is the supervision activation state, calculate the information divergence based on the text-image similarity and the original output value, and use the information divergence as the first loss function;
[0083] Obtain the image prediction label based on the original output value, and calculate the second loss function according to the image prediction label and the observed label;
[0084] Obtain the target loss function based on the first loss function and the second loss function.
[0085] Specifically, convert the text image similarity into a probability distribution, convert the original output value into a probability distribution, and incorporate the text supervision information from the pre-trained vision-language model into the training process by minimizing the divergence between the probability distributions of the image and text predictions. This application uses the Kullback-Leibler divergence (KL), that is, the first loss function is:
[0086] ;
[0087] Among them, is the probability distribution of all categories predicted by the text encoder, represents the category predicted by the text encoder of the probability distribution; is the probability distribution predicted by the image encoder with a fine-tuning module, represents the category predicted by the image encoder with a fine-tuning module of the probability distribution. Since the training parameters of the fine-tuning module are less and it has the ability to quickly adapt to new datasets, this application regards the image encoder fine-tuned on the long-tail noisy label data as a strong student, and the pre-trained vision-language model as a weak teacher, providing weak teacher supervision (WTS).
[0088] The second loss function can adopt the cross-entropy (CE) loss or existing long-tail learning logarithmic adjustment methods, such as the LDAM method and the LA method.
[0089] In the high-noise ratio scenario of this application embodiment, the weak teacher supervision shifts the focus from differentiating noise samples to solving the feature misalignment problem caused by noisy labels. This strategy helps to reduce potential classification errors that may seriously affect the tail categories and improve the performance of all categories, including the head and tail categories.
[0090] In an embodiment of this application, based on the first loss function and the second loss function, a target loss function is obtained, including:
[0091] Set a first weight for the first loss function and a second weight for the second loss function, and the sum of the first weight and the second weight is 1;
[0092] Based on the first weight and the second weight, the first loss function and the second loss function are weighted and summed to obtain the target loss function.
[0093] Specifically, the calculation formula of the target loss function is:
[0094] ;
[0095] Among them, is a hyperparameter, is the second loss function, is the input image and the observed label of. The objective loss function is used to fine-tune the pre-trained vision-language model.
[0096] In the embodiment of the present application, determining the objective loss function based on the supervised start-stop state and the original output value further includes:
[0097] If the supervised start-stop state is the supervised stop state, then an image prediction label is obtained based on the original output value;
[0098] Calculate the second loss function according to the image prediction label and the observed label, and use the second loss function as the objective loss function.
[0099] The supervision switch based on OR can be calculated as:
[0100] ;
[0101] where, is the control threshold of the coincidence rate, represents the random variable the distribution of which follows a beta distribution with parameters and . In the embodiment of the present application, this value is estimated online by calculating the coincidence rate between the text prediction label and the observed label in the mini-batch data, that is, the control threshold is obtained by calculating the coincidence rate between the text prediction label and the observed label of the preset batch data. When weak teacher supervision needs to be turned off, , the objective loss function only includes the second loss function. On the contrary, when weak teacher supervision is turned on, is set to a random number sampled from the beta distribution, and the beta distribution is a continuous probability distribution defined on the interval [0,1] and is often used to describe the probability of a random variable.
[0102] In the low-noise scenario of the embodiment of the present application, the automatic deactivation of weak teacher supervision may introduce incorrect supervision, and only rely on the observed labels that are considered relatively more reliable.
[0103] In an embodiment of the present application, obtaining the image prediction label based on the original output value includes:
[0104] Use the activation function to convert the original output value into the probability distribution of each category, and obtain the image prediction label according to the probability distribution.
[0105] The present application obtains the probability of the input image from the similarity between the fine-tuned image features and the classifier weights, and then obtains the probability distribution of each category.
[0106] The weak teacher supervision (WTS) of this application has the following advantages:
[0107] (1) Feature calibration robust to distribution bias. In high-noise-ratio scenarios, weak teacher supervision shifts the focus from differentiating noise samples to addressing the feature misalignment problem caused by noisy labels. This strategy helps mitigate potential classification errors that may severely affect tail classes and improves the performance of all classes, including head and tail classes. In low-noise scenarios, weak teacher supervision automatically deactivates the supervision that may introduce errors and relies only on the observed labels that are considered relatively more reliable.
[0108] (2) Feature calibration robust to noisy labels. The pre-trained vision-language model is not affected by noisy labels and is aligned with vision-language features, so it is used as a teacher model to correct the bias in the features obtained by the student model. It is worth noting that this application observes that fine-tuning with the second loss function can sometimes outperform text prediction because text prediction is usually less accurate. Nevertheless, weak teacher supervision still provides valuable guidance in correcting the label-image mismatch introduced by noisy labels.
[0109] (3) Efficient training. The computational overhead introduced by weak teacher supervision is extremely small, and the main computational overhead comes from the fine-tuning of the image encoder.
[0110] Table 1 shows the object classification accuracy (%) on the CIFAR-10-LTN and CIFAR-100-LTN datasets (joint noise).
[0111] Table 1
[0112]
[0113] Table 2 shows the object classification accuracy (%) on the CIFAR-10-LTN and CIFAR-100-LTN datasets (symmetric noise, asymmetric noise), with an imbalance rate of 10.
[0114] Table 2
[0115]
[0116] Table 3 shows the classification accuracy (%) on the Webvision-50 dataset (real noise).
[0117] Table 3
[0118]
[0119] Table 4 shows the classification accuracy (%) on the Red mini-ImageNet dataset (real noise).
[0120] Table 4
[0121]
[0122] Among them, CE (Cross Entropy) in the table refers to the cross entropy loss function, which is used for classification tasks to measure the difference between the predicted distribution and the true distribution. LDAM (Label-Distribution-Aware Margin Loss) is a margin adjustment loss function that considers label distribution and aims to improve long-tail learning. LDAM-DRW (Label Distribution Aware Margin with Deferred Re-Weighting) refers to the combination of LDAM and delayed reweighting, which addresses the long-tail distribution problem by adjusting the weights of different categories. NCM (Nearest Class Mean) refers to the nearest class mean classifier, which classifies by calculating the mean of each class and is suitable for processing unbalanced data. MiSLAS (Mixup Shifted Label-Aware Smoothing model) is a label smoothing learning method that enhances classifier learning and calibration by dealing with different degrees of class overconfidence. Co-teaching is a collaborative training method in which two models share data during training but make predictions and updates independently to reduce the impact of noise labels on the model. CDR (Combating noisylabels with Different Update Rules) uses different update strategies for different types of model parameters to reduce the impact of noisy labels. Sel-CL+ (Selective-supervised Contrastive Learning) performs supervised contrastive learning by selecting paired samples with high confidence to obtain robust pre-trained representations. RoLT (Robust Long-Tailed Learning under Label Noise) establishes a new prototype noise detection method to overcome the limitations of the small loss strategy, thereby improving the model's ability to detect noisy labels under long-tail distributions. RoLT-DRW refers to the combination of RoLT and delayed reweighting, which further improves the performance of the model by emphasizing the tail categories. HAR-DRW (Heteroskedastic Adaptive Regularization with Deferred Re-Weighting) combines delayed reweighting to perform stronger regularization on data points with high uncertainty and low density. RCAL (RepresentationCALibration) enhances the robustness of the model by calibrating the image representation obtained by self-supervised contrastive learning.ECBS (Extracting Clean and Balanced Subset for noisy long-tailed classification) proposes a pseudo-labeling framework for extracting clean and class-balanced subsets from long-tailed noisy datasets. ECBS-Res32 is the ECBS model based on the 32-layer residual network; ECBS-Res18 is the ECBS model based on the 18-layer residual network. LA (LogitAdjustment) conducts long-tailed learning by adjusting the logits of different classes. IB (Influence-Balanced Loss) adaptively assigns different weights to samples according to the influence of samples on the decision boundary. DivdeMix is a framework for learning with noisy labels using semi-supervised learning techniques. UNICON (UNIform selection and CONtrastivelearning) unifies the selection mechanism to ensure class balance among the selected clean samples. MW-Net (Meta-Weight-Net) automatically learns an explicit loss weight function in a meta-learning manner. HAR (HeteroskedasticAdaptive Regularization) performs stronger regularization on data points with high uncertainty and low density. ULC (Uncertainty-aware Label Correction framework) combines uncertainty-aware per-class noise modeling and uncertainty-aware learning for label noise. TABASCO (Two-Stage Bi-DimensionalSample Selection) is a two-stage noise sample separation strategy using two separate metrics. CT (Co-teaching) is a co-training method where two models share data during training but make predictions and updates independently of each other to reduce the impact of noisy labels on the models. MentorNet is a method for training a supervised base deep network (student network). ELR+ (Early-Learning Regularization) utilizes early learning through regularization to prevent the model from fitting noisy samples; MoPro (Momentum Prototypes) is a simple contrastive learning method that enables online noisy label correction, out-of-distribution sample removal, and representation learning. NGC (Noisy Graph Cleaning) collects clean samples by leveraging the geometric structure of the data and the model prediction confidence. RCAL+ is the improved version of RCAL.LADE (LAbel distribution DisEntangling) refers to directly decomposing the source label distribution during the training phase, enabling the model to effectively adapt to any target label distribution.
[0123] Tables 1 to 4 compare the classification accuracies of the method of this application and conventional methods. The experimental results show that the weak teacher supervision method proposed in this application can better improve the classification accuracies of each dataset on the basis of CLIP fine-tuning. For all datasets that require manually injecting noisy labels, the strategy of first sampling to grow the long-tail distribution and then injecting the required noisy labels according to the noise rate is adopted. This application has conducted experiments on loss functions such as CE loss, LA loss, and LDAM loss, and has adopted conventional joint noise, symmetric noise, asymmetric noise, and real noise as the types of noisy labels. Sufficient experiments demonstrate the effectiveness of the method of this application.
[0124] In one embodiment, as Figure 3 shown, based on the above-mentioned model training method for long-tail noise, the present invention also correspondingly provides a model training device for long-tail noise, including:
[0125] An image input module 100, configured to input the input image, the text prompt corresponding to the input image, and the observed label in the target long-tail noise dataset into a pre-trained vision-language model, where the vision-language model includes a text encoder and an image encoder, and a fine-tuning module is provided in the image encoder;
[0126] A feature extraction module 200, configured to, in the vision-language model, obtain text features after the text prompt is processed by the text encoder, obtain image features after the input image is processed by the image encoder, and obtain the original output values for each category after the input image is processed by the image encoder with the fine-tuning module;
[0127] A text prediction module 300, configured to calculate the text-image similarity between the text features and the image features, and obtain a text prediction label based on the text-image similarity;
[0128] An update module 400, configured to determine the supervision start-stop state based on the text prediction label and the observed label, determine a target loss function based on the supervision start-stop state and the original output values, and update the fine-tuning module based on the target loss function to obtain a trained vision-language model.
[0129] It should be noted that the foregoing explanations of the embodiments of the model training method for long-tail noise also apply to the model training device for long-tail noise in this embodiment, and will not be elaborated here.
[0130] The present invention discloses a model training device for long-tail noise. By inputting the input image, the corresponding text prompt word and the observation label in the target long-tail noise dataset into a pre-trained vision-language model, the vision-language model includes a text encoder and an image encoder, and a fine-tuning module is set in the image encoder; in the vision-language model, the text prompt word is processed by the text encoder to obtain a text feature, the input image is processed by the image encoder to obtain an image feature, and the input image is processed by the image encoder with the fine-tuning module to obtain the original output values for each category; calculate the text-image similarity between the text feature and the image feature, and obtain a text prediction label based on the text-image similarity; determine a supervised start-stop state based on the text prediction label and the observation label, determine a target loss function based on the supervised start-stop state and the original output values, and update the fine-tuning module based on the target loss function to obtain a trained vision-language model. This application determines whether text-image alignment prior assistance supervision is required by evaluating the difference between the text prediction label and the observation label, and calibrates the deviation between the learned features and the observation label, thereby improving the classification accuracy of head-class and tail-class samples in a high-noise scenario.
[0131] Figure 4 FIG. 4 is a schematic structural diagram of the device provided in the embodiment of the present application. The device may include:
[0132] A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.
[0133] When the processor 502 executes the program, it implements the model training method for long-tail noise provided in the above embodiment.
[0134] Further, the device further includes:
[0135] A communication interface 503 for communication between the memory 501 and the processor 502.
[0136] The memory 501 is used to store a computer program executable on the processor 502.
[0137] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0138] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 can be interconnected via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0139] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a single chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other via an internal interface.
[0140] The processor 502 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0141] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned model training method for long-tail noise is implemented.
[0142] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0143] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0144] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code that includes one or N executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of this application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of this application belong.
[0145] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can read and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with such instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or N wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0146] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0147] Those of ordinary skill in the art can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0148] In addition, in each embodiment of the present application, the functional units can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0149] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, or the like. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for training a model for long-tail noise, characterized in that, The method includes: Inputting the input image, the text prompt corresponding to the input image, and the observation label in the target long-tail noise dataset into a pre-trained vision-language model, where the vision-language model includes a text encoder and an image encoder, and a fine-tuning module is provided in the image encoder; In the vision-language model, the text prompt is processed by the text encoder to obtain text features, the input image is processed by the image encoder to obtain image features, and the input image is processed by the image encoder with the fine-tuning module to obtain the original output values for each category; Calculating the text-image similarity between the text features and the image features, and obtaining a text prediction label based on the text-image similarity; Determining the supervision start-stop state based on the text prediction label and the observation label, determining a target loss function based on the supervision start-stop state and the original output values, and updating the fine-tuning module based on the target loss function to obtain a trained vision-language model; Determining the supervision start-stop state based on the text prediction label and the observation label, including: Calculating the coincidence rate between the text prediction label and the observation label; Obtaining a preset control threshold, comparing the coincidence rate with the control threshold, and obtaining a comparison result; Determining the supervision start-stop state according to the comparison result; Determining the supervision start-stop state according to the comparison result, including: If the comparison result is that the coincidence rate is greater than the control threshold, the supervision start-stop state is the supervision stop state; If the comparison result is that the coincidence rate is less than or equal to the control threshold, the supervision start-stop state is the supervision start state.
2. The model training method for long-tail noise according to claim 1, wherein Determining the target loss function based on the supervision start-stop state and the original output values, including: If the supervision start-stop state is the supervision start state, calculating the information divergence based on the text-image similarity and the original output values, and taking the information divergence as the first loss function; Obtaining an image prediction label based on the original output values, and calculating a second loss function according to the image prediction label and the observation label; Obtaining the target loss function based on the first loss function and the second loss function.
3. The model training method for long-tail noise according to claim 2, wherein, Obtaining the target loss function based on the first loss function and the second loss function, including: Setting a first weight for the first loss function and a second weight for the second loss function, and the sum of the first weight and the second weight is 1; Weighted summing the first loss function and the second loss function based on the first weight and the second weight to obtain the target loss function.
4. The model training method for long-tail noise according to claim 1, characterized in that Determining the target loss function based on the supervision start-stop state and the original output values, further including: If the supervision start-stop state is the supervision stop state, obtaining an image prediction label based on the original output values; Calculating a second loss function according to the image prediction label and the observation label, and taking the second loss function as the target loss function.
5. The model training method for long-tail noise according to claim 2 or 4, characterized in that Obtaining an image prediction label based on the original output values, including: Using an activation function to convert the original output values into a probability distribution for each category, and obtaining an image prediction label according to the probability distribution.
6. A model training device for long-tail noise, characterized in that, The device includes: An image input module, configured to input an input image, a text prompt corresponding to the input image, and an observation label in a target long-tail noise dataset into a pre-trained vision-language model, where the vision-language model includes a text encoder and an image encoder, and a fine-tuning module is provided in the image encoder; A feature extraction module, configured to obtain a text feature after the text prompt is processed by the text encoder in the vision-language model, obtain an image feature after the input image is processed by the image encoder, and obtain an original output value for each category after the input image is processed by the image encoder with the fine-tuning module; A text prediction module, configured to calculate the text-image similarity between the text feature and the image feature, and obtain a text prediction label based on the text-image similarity; An update module, configured to determine a supervision start-stop state based on the text prediction label and the observation label, determine a target loss function based on the supervision start-stop state and the original output value, and update the fine-tuning module based on the target loss function to obtain a trained vision-language model; Determining the supervision start-stop state based on the text prediction label and the observation label includes: Calculating the coincidence rate between the text prediction label and the observation label; Obtaining a preset control threshold, comparing the coincidence rate with the control threshold, and obtaining a comparison result; Determining the supervision start-stop state according to the comparison result; Determining the supervision start-stop state according to the comparison result includes: If the comparison result is that the coincidence rate is greater than the control threshold, the supervision start-stop state is the supervision stop state; If the comparison result is that the coincidence rate is less than or equal to the control threshold, the supervision start-stop state is the supervision start state.
7. A device, characterized in that, Including: A memory, a processor, and a model training program for long-tail noise stored on the memory and executable on the processor, where when the model training program for long-tail noise is executed by the processor, the steps of the model training method for long-tail noise according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the model training method for long-tail noise according to any one of claims 1 to 5.
Citation Information
Patent Citations
Target detection method based on mixed supervised training
CN114154563A
Water level monitoring method based on deep learning
CN118887374A