Pseudo-label enhancement training method using predicted inconsistent samples

By using the pseudo-label enhancement training method that predicts inconsistent samples, the problem of poor calibration of pseudo-label selection indicators in deep learning models is solved, and a more stable and efficient model training effect is achieved.

CN120087502AActive Publication Date: 2025-06-03XI AN JIAOTONG UNIV

Patent Information

Application Number
CN202510525771.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-06-03
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

During the training process of existing deep learning models, pseudo-label selection indicators such as confidence scores have poor calibration problems, resulting in mispredictions and the impact of adversarial samples, affecting the model training effect.

Method used

By using the pseudo-label enhancement training method to predict inconsistent samples, it is determined whether the historical prediction data set of labelless image data meets the set conditions, that is, the label prediction is stabilized on two different categories of labels, and then the pseudo-label is determined and an enhanced training set is constructed to train the label prediction model.

Benefits of technology

Through this method, the training effect of the label prediction model can be further enhanced based on the traditional pseudo-label set, and the performance and stability of the model can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087502A_ABST
    Figure CN120087502A_ABST
Patent Text Reader

Abstract

The invention provides a pseudo-label enhancement training method using predicted inconsistent samples, and relates to the technical field of deep learning, the method comprises the following steps: a first stage, training an initial model by a real label training set to obtain a first model; in the second stage, a baseline pseudo-label training set is constructed through a baseline pseudo-label selection method, the baseline pseudo-label training set and a real label training set form a target training set to train the first model to obtain a second model, and prediction distribution of the first model after each round of training on unlabeled data is stored as historical prediction data; in the third stage, non-label data with prediction inconsistency are screened from historical prediction data to serve as pseudo-label data, pseudo labels of the non-label data are the category predicted by the second model, and the prediction inconsistency means that prediction labels of the non-label data are stabilized in two categories in sequence; and forming an enhanced training set by the pseudo label data and the target training set to train the initial model, and obtaining a final model. The objective of the invention is to provide a model trained by various feature pseudo-label data so as to improve the final model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and particularly to a method for enhancing training using pseudo-labels of prediction-inconsistent samples. Background Art

[0002] In the field of deep learning technology, the training of some models requires a large amount of labeled sample data with labels. In practical applications, the cost of obtaining a large amount of labeled sample data is high, and semi-supervised learning can use unlabeled data to improve the model performance with limited labeled data. As a common method in semi-supervised learning, the pseudo-labeling technology is widely used in semi-supervised learning. The core of the pseudo-labeling technology is pseudo-label selection, and the goal of pseudo-label selection is to determine which samples should be assigned pseudo-labels.

[0003] Currently, the most commonly used pseudo-label selection metric is the confidence score obtained from the softmax distribution. However, this metric has the problem of poor calibration, that is, high confidence scores are often assigned to mispredicted labels, resulting in sample data having incorrect pseudo-labels. Using such pseudo-label sample data for training the model will affect the model training effect, and at the same time, the confidence score is also vulnerable to manipulation by adversarial samples. Summary of the Invention

[0004] In view of this, this application provides a method for enhancing training using pseudo-labels of prediction-inconsistent samples, aiming to provide pseudo-label data with different features to train the model and improve the model training effect.

[0005] In the first aspect of this application, a method for enhancing training using pseudo-labels of prediction-inconsistent samples is provided. The method includes: Determine the pseudo-labels of unlabeled image data through a baseline pseudo-label selection method to construct a baseline pseudo-label training set; Train a first model with a target training set composed of a true-label training set and a baseline pseudo-label training set, and use the first model after each training round to predict the labels of unlabeled image data to obtain corresponding prediction distribution results. All the prediction distribution results obtained by the unlabeled image data in all training rounds form a corresponding historical prediction data set; When the first model is trained to be qualified to obtain a corresponding second model, determine whether the historical prediction data set of unlabeled image data meets the set condition. The set condition is that the label predictions of unlabeled image data are successively stable on two different category labels during the training process; Determine the unlabeled image data that meets the set condition as pseudo-label image data, and determine the category label on which the final label prediction is stable as the pseudo-label of the pseudo-label image data; Compose an augmented training set with the obtained large amount of pseudo-labeled image data and the target training set, and train the constructed initial model through the augmented training set to obtain a qualified label prediction model.

[0006] Optionally, training the constructed initial model through the augmented training set to obtain a qualified label prediction model includes: When training the constructed initial model through the augmented training set and a qualified label prediction model cannot be obtained after a preset number of trainings, determine the initial model after the preset number of trainings as the third model; Train the third model through the target training set, and use the third model after each training round to predict the labels of new unlabeled image data to obtain corresponding prediction distribution results. All the prediction distribution results obtained from the new unlabeled image data in all training rounds form the corresponding historical prediction data set; When the third model is trained qualified, determine whether the historical prediction data set of the new unlabeled image data meets the set conditions; Determine the new unlabeled image data that meets the set conditions as pseudo-labeled image data, and determine the class label at which the final label prediction stabilizes as the pseudo-label of the pseudo-labeled image data; Compose a new augmented training set with the obtained large amount of new pseudo-labeled image data and the augmented training set, and train the third model after the preset number of trainings through the new augmented training set to obtain a qualified label prediction model.

[0007] Optionally, the method further includes: Test the trained qualified label prediction model with the labeled image data in the test set to evaluate the performance of the label prediction model on unseen image data; Determine whether to perform new training on the label prediction model according to the performance evaluation result.

[0008] Optionally, determining the pseudo-labels of unlabeled image data through a baseline pseudo-label selection method to construct a baseline pseudo-label training set includes: Predict the labels of unlabeled image data through the first model; Determine the predicted label categories that meet the baseline pseudo-label selection method in the label prediction results as the pseudo-labels of the unlabeled image data; Construct a baseline pseudo-label training set based on the determined unlabeled image data and their respective pseudo-labels.

[0009] Optionally, before predicting the labels of unlabeled image data through the first model, the method further includes: Pre-train the initial model with the labeled image data in the real label training set; Determine the initial model that passes the pre-training as the first model.

[0010] Optionally, constructing the initial model includes: Based on the selected type of neural network model, construct the basic structure of the model, and the basic structure includes: input layer, convolutional layer, pooling layer, activation function, and fully connected layer; Add a softmax layer after the fully connected layer in the basic structure to convert the model output into a probability distribution; Set the loss function of the model to construct and obtain the final initial model.

[0011] Optionally, when the first model is trained successfully to obtain the corresponding second model, determine whether the historical prediction dataset of the unlabeled image data meets the set conditions, including: When the first model is trained successfully to obtain the corresponding second model, calculate the historical prediction dataset of the unlabeled image data through the prediction distribution mean algorithm to obtain the mean vector of the prediction distribution of the unlabeled image data; The calculation expression of the prediction distribution mean algorithm is:

[0012] Among them, is the prediction distribution result obtained by the i-th unlabeled image data in the t-th training round, and this prediction distribution result records the probabilities of the i-th unlabeled image data belonging to various category labels in the t-th training round; represents the total number of training rounds; is the mean vector of the prediction distribution of the i-th unlabeled image data; Calculate the obtained mean vector of the prediction distribution through the target entropy algorithm to obtain the spatial features of the unlabeled image data; The calculation expression of the target entropy algorithm is:

[0013] Where is except all other category labels except the category label corresponding to the maximum value in; is the mean of the prediction probabilities of the i-th unlabeled image data under the category label c; is the spatial feature of the i-th unlabeled image data; is the entropy after removing the maximum value; Calculate the historical prediction dataset of the unlabeled image data through the change trend algorithm to obtain the change trend vector of the prediction distribution of the unlabeled image data; The calculation expression of the change trend algorithm is:

[0014] where, and are the consecutive training rounds of the unlabeled image data, is the next training round; and respectively represent the prediction distribution results of the i-th unlabeled image data in the training rounds and respectively, represents the change trend vector of the prediction distribution of the i-th unlabeled image data; Calculate the obtained change trend vector through the first algorithm to obtain the first operator value and the second operator value of the unlabeled image data; The calculation expression of the first algorithm is:

[0015] where, is the first operator of the i-th unlabeled image data; is the second operator of the i-th unlabeled image data; is the final predicted category of the unlabeled image data during training, and this predicted category is defined as the predicted category in the second stage; is except, the predicted category corresponding to the maximum change trend in the change trend vector of the prediction distribution of the unlabeled image data, and this predicted category is defined as the predicted category in the first stage; is the change trend of the predicted probability of the i-th unlabeled image data on the predicted category in the first stage; is the change trend of the predicted probability of the i-th unlabeled image data on the predicted category in the second stage; is the change trend of the prediction distribution of the i-th unlabeled image data on other predicted categories except the predicted categories in the first stage and the second stage; Calculate the first operator value and the second operator value through the second algorithm to obtain the time feature of the unlabeled image data; The calculation expression of the second algorithm is:

[0016]

[0017]

[0018]

[0019]

[0020] wherein and are thresholds; is the change index of the i-th unlabeled image data; is the direction index of the i-th unlabeled image data; T is the last training epoch in the training process; is the prediction probability of the i-th unlabeled image data on the predicted class in the second stage in the t-th training epoch; is the prediction probability of the i-th unlabeled image data on the predicted class in the first stage in the t-th training epoch; is the time feature of the i-th unlabeled image data; and are weights; Determine whether the prediction distribution result of the unlabeled image data meets the set conditions according to the spatial feature and time feature of the unlabeled image data.

[0021] Optionally, determining whether the prediction distribution result of the unlabeled image data meets the set conditions according to the spatial feature and time feature of the unlabeled image data includes: Calculating the spatial feature and time feature of the unlabeled image data through a prediction inconsistency index algorithm to obtain the value of the prediction inconsistency index of the unlabeled image data; The calculation expression of the prediction inconsistency index algorithm is:

[0022] wherein and are weights; Determine whether the prediction distribution result of the unlabeled image data meets the set conditions according to the relationship between the value of the prediction inconsistency index and the set threshold.

[0023] For the prior art, the present application has the following advantages: A method for enhancing training by using pseudo - labels of prediction - inconsistent samples provided by an embodiment of the present application. First, through a baseline pseudo - label selection method (such as selecting the class label with the highest confidence obtained by the model's prediction of unlabeled image data as the pseudo - label of the unlabeled image data), the pseudo - labels of the unlabeled image data are determined to construct a baseline pseudo - label training set; the first model is trained with a target training set composed of a true - label training set and a baseline pseudo - label training set, and the first model after each training round is used to predict the labels of the unlabeled image data to obtain the corresponding prediction distribution results. All the prediction distribution results obtained by the unlabeled image data in all training rounds form the corresponding historical prediction data set, and all the prediction distribution results of a single unlabeled image data itself in all training rounds form its own corresponding historical prediction data set; when the first model is trained to be qualified to obtain the corresponding second model, it is determined whether the historical prediction data set of the unlabeled image data itself meets the set conditions. The set conditions are that the label predictions of the unlabeled image data are successively stable on two different class labels during the training process, that is, stable on one of the two class labels in the earlier training rounds during the training process, and stable on the other of the two class labels in the later training rounds during the training process; the unlabeled image data whose own historical prediction data set meets the set conditions is determined as pseudo - label image data, and the class label on which the final label prediction is stable is determined as the pseudo - label of the pseudo - label image data; a large number of obtained pseudo - label image data and the target training set are used to form an enhanced training set, and the initial model constructed is trained with the enhanced training set to obtain a trained and qualified label prediction model. Thus, the present application adopts a three - stage training method (pre - training the model, determining the baseline pseudo - labels based on the pre - trained model, and determining the prediction - inconsistent pseudo - labels), so that the data used for training the model includes not only the image data with manually labeled true labels, but also the image data with baseline pseudo - labels and the image data with prediction - inconsistent pseudo - labels. The pseudo - label image data obtained by these two different pseudo - label determination methods have different characteristics (for example, the characteristics of the baseline pseudo - label image data are mainly reflected in the confidence level, and the characteristics of the new pseudo - label image data selected by the present application through prediction inconsistency are mainly reflected in the inconsistency of the predicted categories before and after), which can further enhance the training effect of the label prediction model on the basis of the traditional pseudo - label set.

[0024] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are hereinafter specifically exemplified. Brief Description of the Drawings

[0025] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art.

[0026] Figure 1 It is a flowchart of a pseudo-label enhancement training method using predicted inconsistent samples provided by an embodiment of the present application; Figure 2 It is a schematic diagram of the relationship between the training rounds and predicted categories of unlabeled image data in a pseudo-label enhancement training method using predicted inconsistent samples provided by an embodiment of the present application; Figure 3 It is another schematic diagram of the relationship between the training rounds and predicted categories of unlabeled image data in a pseudo-label enhancement training method using predicted inconsistent samples provided by an embodiment of the present application. Detailed implementation manners

[0027] The following will describe the exemplary embodiments of the present application in more detail with reference to the accompanying drawings.

[0028] Figure 1 It is a flowchart of a pseudo-label enhancement training method using predicted inconsistent samples provided by an embodiment of the present application. As Figure 1 shown, the method includes: Step S1: Determine the pseudo-labels of the unlabeled image data through a baseline pseudo-label selection method to construct a baseline pseudo-label training set.

[0029] In this embodiment, a pseudo-label enhancement training method using predicted inconsistent samples provided by the present application is applied to an image classification task, that is, training and applying a neural network model for image classification. The image classification task includes a single-label binary classification task and a single-label multi-classification task. The single-label binary classification task refers to the task of classifying an image into two categories. The number of labels or categories of all images in the dataset is only two. The image input into the neural network model only carries one category label, and the neural network model needs to identify the image as one of the two categories based on the learned features. For example, the cat-dog classification task, that is, each image can only be a cat or a dog, and the neural network model identifies each image as one of the two categories of cats and dogs. The single-label multi-classification task refers to the case where the number of labels or categories of all images in the dataset has multiple categories (more than two), but each image has only one label or category, and the neural network model needs to identify the image as one of the multiple categories based on the learned features. For example, the handwritten digit recognition task, each image can only be any one of the handwritten digits from 0 to 9, and the neural network model identifies each image as one of the 10 categories of handwritten digits. Another example is the fruit classification task, such as classifying an image into one of multiple categories such as apples, bananas, and oranges.

[0030] In this embodiment, in the field of deep learning, an image is usually a three-dimensional tensor, including three dimensions: height, width, and channels, containing information such as space and color. An image is the input of a deep learning model, and the model processes these tensors to understand and analyze the image content and complete various computer vision tasks. A label is the target output or correct answer corresponding to the input image data in supervised learning. For example, in an image classification task, the label is the category of the image. A pseudo-label refers to generating pseudo-labels for unlabeled data using the results predicted by the model when there is no or only a small amount of labeled data. Training dynamics refers to the deep learning model iteratively updating gradients to optimize model parameters, and training dynamics studies the laws and behaviors of model parameters, loss functions, gradients, prediction results, etc. changing over time (or training epochs) during this process. Prediction-inconsistent pattern training dynamics: This type of training dynamics has obvious prediction-inconsistent characteristics, that is, in the first stage, the model prediction stabilizes in one category, and in the second stage, the model prediction mainly stabilizes in another category, that is, the prediction results show inconsistent predictions before and after, first stabilizing in one category during training and then stabilizing in another category. And this application mainly proposes a prediction-inconsistent index based on the prediction-inconsistent pattern training dynamics. This prediction-inconsistent index forms a quantitative index (i.e., the prediction-inconsistent index) for screening pseudo-labels of unlabeled image data by quantifying the unique characteristics of unlabeled image data in the training dynamics of prediction-inconsistent labels.

[0031] In this embodiment, first, an image data set is collected, which includes a small amount of labeled image data with real labels and a large amount of unlabeled image data. The small amount of labeled image data is used to divide the real label training set and the test set, and at the same time, ensure that images with different category labels are evenly distributed in these sets during the division. The real label training set is used to train the first model and the initial model, and the real label test set is used to test the finally trained and qualified label prediction model to evaluate its performance on unseen image data, such as metrics like accuracy, precision, recall, and F1 value in an image classification task. The initial model is an initial model constructed for the image classification task, and the model type of this initial model can be CNN, ResNet, etc.

[0032] In this embodiment, an optional implementation of the baseline pseudo-label selection method is to use a model trained to a certain extent to predict labels for unlabeled image data, and select the prediction category with the highest confidence in the obtained label prediction results as the pseudo-label of the unlabeled image data, and determine the unlabeled image data with this pseudo-label as the baseline pseudo-label image data. Through this baseline pseudo-label selection method, a large number of baseline pseudo-label image data are determined to construct a baseline pseudo-label training set. Among them, determining the pseudo-label of unlabeled image data by the method with the highest confidence is only an optional implementation of the baseline pseudo-label selection method, and this baseline pseudo-label selection method can be other baseline pseudo-label selection methods, which are not specifically limited herein.

[0033] Step S2: Train the first model with a target training set composed of a true label training set and a baseline pseudo-label training set, and use the first model after each training round to predict labels for unlabeled image data to obtain corresponding prediction distribution results. All the prediction distribution results obtained by the unlabeled image data in all training rounds constitute the corresponding historical prediction data set.

[0034] In this embodiment, the present application pre-establishes a true label training set, which records labeled image data with true labels. After obtaining the baseline pseudo-label training set through step S1, the true label training set and the baseline pseudo-label training set are merged to jointly form a target training set. The target training set formed is used to train the first model, and the first model is preferably the model that predicts labels for unlabeled image data in step S1.

[0035] In this embodiment, the training process of training the first model with the constructed target training set is as follows: Input the image data in the target training set into the first model for training, and calculate the feature representation of the image data input into the first model through forward propagation. Calculate the training loss function using the label information of the input image data, and update the model parameters of the first model according to the loss value obtained from the loss function. Then calculate the gradient through the backpropagation algorithm, and use an optimization algorithm (such as SGD Stochastic Gradient Descent, Adam Adaptive Moment Estimation, etc.) to update the model parameters. During the training process, when the loss value obtained by calculating the loss function meets the corresponding set conditions, such as being lower than a certain set threshold, it is determined that the first model is qualified for training, and the corresponding second model is obtained. Here, the set threshold can be set according to the actual application scenario and will not be specifically limited here. During the entire training process of the first model to obtain the finally qualified second model, multiple training rounds are required to obtain the finally qualified second model. One training round means that each time the loss value is calculated and the model parameters of the first model are updated based on the calculated loss value, it is counted as one training round.

[0036] In this embodiment, after each training round, the first model after the current training round is used to predict the labels of the unlabeled image data, and the prediction distribution result of the unlabeled image data in the current training round is obtained. A prediction distribution result records the probabilities that the first model predicts the unlabeled image data belongs to various label categories after one training round. Through the same implementation method, each unlabeled image data in a large number of unlabeled image data will have multiple prediction distribution results with the same number as the total number of training rounds. The multiple prediction distribution results of a single unlabeled image data with the same number as the total number of training rounds constitute the historical prediction data set of the single unlabeled image data. The historical prediction data sets of each unlabeled image data are used to determine whether the unlabeled image data itself belongs to pseudo-labeled image data and for pseudo-label selection of pseudo-labeled image data.

[0037] Exemplarily, since the first model predicts labels for unlabeled image data in the same way after each training round during the training process, an unlabeled image data a is taken as an example for illustration here. Assuming that the number of training rounds from the first model to the finally trained and qualified second model is 2 times, this 2 times is only an exemplary explanation for easy understanding, and the actual number of training rounds is generally much larger than 2 times. Correspondingly, after the first training round, the first model will perform a label prediction on the unlabeled image data a to obtain a corresponding prediction distribution result A1, and after the second training round, the first model will perform a label prediction on the unlabeled image data a to obtain a corresponding prediction distribution result A2. That is, for however many training rounds (such as 100 times), an unlabeled image data will have the corresponding number of prediction distribution results (such as 100).

[0038] Step S3: When the first model is trained and qualified to obtain the corresponding second model, determine whether the historical prediction dataset of the unlabeled image data meets the set condition, where the set condition is that the label predictions of the unlabeled image data during the training process are successively stable on two different category labels.

[0039] In this embodiment, after the first model is trained and qualified to obtain the corresponding second model, each unlabeled image data will have a historical prediction dataset corresponding to itself. Then, when the first model is trained and qualified to obtain the corresponding second model, at this time, it is determined whether the historical prediction datasets of each unlabeled image data meet the set condition respectively. The set condition is that the label predictions of the unlabeled image data during the training process are successively stable on two different category labels, that is, in the earlier training rounds during the training process, it is stable on one of the two category labels, and in the later training rounds during the training process, it is stable on the other of the two category labels. As Figure 2 shown Figure 2 in, the abscissa represents the number of training rounds, and the ordinate represents the label category to which the label prediction result obtained by the unlabeled image data in the corresponding training round belongs. The label category to which it belongs refers to the label category with the highest confidence score in the corresponding prediction distribution result. Figure 2 The historical prediction dataset of the unlabeled image data of the first Node in meets the set condition. Not only are the label predictions during the training process stable on two category labels (category 1 and category 2 respectively), but also the label predictions are predicted as one category (category 1) in the earlier training rounds during the training process and as the other category (category 2) in the later training rounds.

[0040] Step S4: Determine the unlabeled image data that meets the set conditions as pseudo-labeled image data, and determine the class label at which the final label prediction stabilizes as the pseudo-label of the pseudo-labeled image data.

[0041] In this embodiment, after determining the historical prediction data set that meets the set conditions through step S3, determine the unlabeled image data corresponding to the historical prediction data set that meets the set conditions as pseudo-labeled image data. At the same time, determine the class label at which the label prediction manifested in the historical prediction data set of this pseudo-labeled image data finally stabilizes as the pseudo-label of this pseudo-labeled image data, label the pseudo-labeled image data with this pseudo-label, and use the labeled pseudo-labeled image data to participate in the subsequent training of the label prediction model. As Figure 2 The historical prediction data set of the unlabeled image data of the first Node in shows that it meets the set conditions. Therefore, this Node can be determined as pseudo-labeled image data. At the same time, the label prediction of the unlabeled image data of this Node stabilizes at class 2. Therefore, class 2 is used as the pseudo-label of this pseudo-labeled image data.

[0042] Step S5: Use the obtained large amount of pseudo-labeled image data and the target training set to form an enhanced training set, and train the constructed initial model through the enhanced training set to obtain a qualified label prediction model.

[0043] In this embodiment, after obtaining a large amount of pseudo-labeled image data labeled with their own pseudo-labels through the same implementation methods as steps S2 to S4, the pseudo-label training set composed of these labeled large amounts of pseudo-labeled image data and the target training set used for the initial training of the first model together form an enhanced training set. Then, use this enhanced training set to train the initial model pre-constructed for the image classification task to obtain a qualified label prediction model that can finally be used to perform the image classification task.

[0044] In this embodiment, an optional implementation manner is that the label prediction model is used for the scenario of identifying the intersection categories of roads in the field of autonomous driving. In this application scenario, the image data used for training the model is the image data of the road intersections collected, and the label categories of these image data include crossroads category, Y-shaped intersection category, T-shaped intersection category, X-shaped intersection category, roundabout intersection category, etc. The trained label prediction model is installed in the vehicle terminal to identify the intersection categories of the collected image data during autonomous driving to guide the autonomous driving of the vehicle. It should be understood that a method for enhancing training with pseudo-labels using prediction-inconsistent samples provided in this application can also be used in other application scenarios for performing image classification tasks.

[0045] In this embodiment, the difference between the training of the first model and the training of the initial model in this application is that the training of the first model refers to predicting labels for unlabeled image data, and using the historical prediction dataset composed of all the prediction distribution results corresponding to the unlabeled image data itself for the subsequent determination of pseudo-labeled image data and the training process of pseudo-label selection for the pseudo-labeled image data. The training of the initial model is to train the model in order to obtain a label prediction model that can ultimately be used for image classification.

[0046] A method for enhancing training using pseudo-labels of prediction-inconsistent samples provided by an embodiment of the present application. First, through a baseline pseudo-label selection method (such as selecting the class label with the highest confidence predicted by the model for unlabeled image data as the pseudo-label of the unlabeled image data), the pseudo-labels of the unlabeled image data are determined to construct a baseline pseudo-label training set; the first model is trained with a target training set composed of a true-label training set and a baseline pseudo-label training set, and the first model after each training round is used to predict the labels of the unlabeled image data to obtain corresponding prediction distribution results. All the prediction distribution results obtained by the unlabeled image data in all training rounds form a corresponding historical prediction data set, and all the prediction distribution results of the unlabeled image data itself in all training rounds form its own corresponding historical prediction data set; when the first model is trained successfully to obtain a corresponding second model, it is determined whether the historical prediction data set of the unlabeled image data itself meets the set conditions. The set conditions are that the label predictions of the unlabeled image data are successively stable on two different class labels during the training process, that is, stable on one of the two class labels in the earlier training rounds during the training process, and stable on the other of the two class labels in the later training rounds during the training process; the unlabeled image data whose own historical prediction data set meets the set conditions is determined as pseudo-label image data, and the class label on which the final label prediction is stable is determined as the pseudo-label of the pseudo-label image data; a large number of obtained pseudo-label image data and the target training set are used to form an enhanced training set, and the constructed initial model is trained with the enhanced training set to obtain a trained label prediction model. Thus, the present application adopts a three-stage training method (pre-training model, determining baseline pseudo-labels based on the pre-training model, and determining prediction-inconsistent pseudo-labels), so that the data used for training the model includes not only the image data with manually labeled true labels, but also the image data with baseline pseudo-labels and the image data with prediction-inconsistent pseudo-labels. The pseudo-label image data obtained by these two different pseudo-label determination methods have different characteristics (such as the characteristics of the baseline pseudo-label image data are mainly reflected in confidence, and the characteristics of the new pseudo-label image data selected by the present application through prediction inconsistency are mainly reflected in the inconsistency of the predicted categories before and after), which can further enhance the training effect of the label prediction model on the basis of the traditional pseudo-label set.

[0047] Combined with the above embodiments, in one implementation manner, the embodiment of the present application further provides a method for enhancing training using pseudo-labels of prediction-inconsistent samples. In this method for enhancing training using pseudo-labels of prediction-inconsistent samples, step S5 may include steps S51 to S55: Step S51: Train the constructed initial model with the enhanced training set. When a qualified label prediction model cannot be obtained after a preset number of training sessions, determine the initial model after the preset number of training sessions as the third model.

[0048] In this embodiment, the initial model is trained with the enhanced training set, and the number of training sessions is specified. After the initial model is trained with the enhanced training set for a preset number of times, if it is still not qualified based on the corresponding loss function, in order to improve the training efficiency and performance of the model, this application will provide more pseudo-label image data to participate in the subsequent training process of the initial model after the preset number of training sessions.

[0049] Specifically, when the initial model is trained with the enhanced training set and a qualified label prediction model cannot be obtained after a preset number of training sessions, the initial model after the preset number of training sessions is determined as the third model.

[0050] Step S52: Train the third model with the target training set, and use the third model after each training round to predict the labels of new unlabeled image data, obtaining the corresponding prediction distribution results. All the prediction distribution results obtained from the new unlabeled image data in all training rounds form the corresponding historical prediction data set.

[0051] In this embodiment, for the third model determined in step S51, the third model is trained with the target training set used to train the first model. At the same time, the third model after each training round in the current training process is used to predict the labels of new unlabeled image data, obtaining the corresponding prediction distribution results. All the prediction distribution results obtained from the new unlabeled image data in all training rounds form the historical prediction data set corresponding to the new unlabeled image data. The purpose of training the third model with the target training set at this time is to obtain the historical prediction data set of the new unlabeled image data for determining the pseudo-labels of the new unlabeled image data, rather than training to obtain the final label prediction model.

[0052] Step S53: When the third model is trained qualified, determine whether the historical prediction data set of the new unlabeled image data meets the set conditions.

[0053] In this embodiment, when the third model is successfully trained, each new unlabeled image data will have a corresponding historical prediction dataset of its own. Then, when the third model is successfully trained, at this time, it is determined whether the historical prediction dataset of each new unlabeled image data meets the set conditions. The set condition is that the historical prediction dataset of the unlabeled image data shows that the label prediction results during the training process are successively stable on two category labels, that is, stable on one of the two category labels in the earlier training rounds during the training process, and stable on the other of the two category labels in the later training rounds during the training process.

[0054] Step S54: Determine the new unlabeled image data that meets the set conditions as pseudo-labeled image data, and determine the category label on which the final label prediction stabilizes as the pseudo-label of the pseudo-labeled image data.

[0055] In this embodiment, after determining whether the historical prediction dataset meets the set conditions through step S53, the new unlabeled image data corresponding to the historical prediction dataset that meets the set conditions is determined as pseudo-labeled image data. At the same time, the category label on which the label prediction manifested in the historical prediction dataset of the pseudo-labeled image data stabilizes is determined as the pseudo-label of the pseudo-labeled image data. The pseudo-labeled image data is labeled with this pseudo-label and the labeled pseudo-labeled image data is used to participate in the subsequent training of the label prediction model.

[0056] Step S55: Use the obtained large number of new pseudo-labeled image data and the enhanced training set to form a new enhanced training set, and train the third model that has been trained a preset number of times through the new enhanced training set to obtain a successfully trained label prediction model.

[0057] In this embodiment, after obtaining a large number of new pseudo-labeled image data through step S54, the training set composed of this large number of new pseudo-labeled image data and the previous enhanced training set are combined to form a new enhanced training set, and the currently trained third model is continuously trained through this new enhanced training set to obtain a finally successfully trained label prediction model.

[0058] Combining the above embodiments, in one implementation manner, the embodiments of the present application further provide a pseudo-label enhancement training method using prediction-inconsistent samples. In this pseudo-label enhancement training method using prediction-inconsistent samples, the method further includes: testing the successfully trained label prediction model with the labeled image data in the test set to evaluate the performance of the label prediction model on unseen image data; and determining whether to perform new training on the label prediction model according to the performance evaluation result.

[0059] In this embodiment, after training to obtain a qualified label prediction model, the trained and qualified label prediction model is tested with the labeled image data with true labels in the test set to evaluate the performance of the label prediction model on unseen image data. According to the performance evaluation results, it is determined whether a new training of the label prediction model is required. For example, in the case where the evaluation results do not meet the corresponding performance conditions, the label prediction model is newly trained. The performance conditions include whether the accuracy rate reaches the corresponding threshold set in advance, whether the precision rate reaches the corresponding threshold set in advance, whether the recall rate reaches the corresponding threshold set in advance, and whether the F1 value meets the corresponding requirements.

[0060] Combined with the above embodiments, in one implementation, the embodiments of the present application further provide a pseudo-label enhanced training method using prediction-inconsistent samples. In this pseudo-label enhanced training method using prediction-inconsistent samples, step S1 may include: performing label prediction on unlabeled image data through a first model; determining the predicted label category that meets the baseline pseudo-label selection method in the label prediction results as the pseudo-label of the unlabeled image data; constructing a baseline pseudo-label training set based on the determined unlabeled image data and their respective pseudo-labels.

[0061] In this embodiment, the baseline pseudo-label selection method is preferably to select the predicted label category with the highest confidence of the unlabeled image data as the pseudo-label of the unlabeled image data. First, label prediction is performed on the unlabeled image data through the first model to obtain the confidence of the unlabeled image data belonging to various label categories. Then, the predicted label category that meets the baseline pseudo-label selection method (i.e., the predicted label category with the highest confidence) in the label prediction results is determined as the pseudo-label of the unlabeled image data; the unlabeled image data and its own pseudo-label form baseline pseudo-label image data, and a large number of baseline pseudo-label image data constitute a baseline pseudo-label training set.

[0062] Combined with the above embodiments, in one implementation, the embodiments of the present application further provide a pseudo-label enhanced training method using prediction-inconsistent samples. In this pseudo-label enhanced training method using prediction-inconsistent samples, before performing label prediction on the unlabeled image data through the first model, the method further includes: pre-training the constructed initial model with the labeled image data in the true label training set; determining the pre-trained and qualified initial model as the first model.

[0063] In this embodiment, before determining the pseudo-labels of the unlabeled image data through the baseline pseudo-label selection method to construct the baseline pseudo-label training set, the initial model constructed is pre-trained with the labeled image data in the true label training set to obtain a pre-trained qualified initial model. This pre-trained qualified initial model is determined as the first model, and this first model is used to predict the label categories of the unlabeled image data in step S1.

[0064] Combined with the above embodiments, in one implementation manner, the embodiments of the present application further provide a pseudo-label enhanced training method using prediction-inconsistent samples. In this pseudo-label enhanced training method using prediction-inconsistent samples, constructing an initial model includes: based on the selected type of neural network model, constructing the basic structure of the model, and the basic structure includes: an input layer, a convolutional layer, a pooling layer, an activation function, and a fully connected layer; adding a softmax layer after the fully connected layer in the basic structure to convert the model output into a probability distribution; setting the loss function of the model to construct and obtain the final initial model.

[0065] In this embodiment, an appropriate type of neural network model is selected based on the requirements of the image classification task, such as CNN, ResNet, etc. Then, the basic structure of the corresponding type of neural network model is constructed, including the input layer, convolutional layer, and pooling layer of the model for extracting image features, activation function, fully connected layer, etc. For the classification task, it is necessary to convert the model output into a probability distribution. Therefore, a softmax layer is added after the fully connected layer in the basic structure of the neural network model to convert the model output into a probability distribution. Then, the loss function of the model is set to obtain the final initial model.

[0066] Combined with the above embodiments, in one implementation manner, the embodiments of the present application further provide a pseudo-label enhanced training method using prediction-inconsistent samples. In this pseudo-label enhanced training method using prediction-inconsistent samples, step S3 may include steps S31 to S36: Step S31: When the first model is trained to be qualified to obtain the corresponding second model, calculate the historical prediction data set of the unlabeled image data through the prediction distribution mean algorithm to obtain the mean vector of the prediction distribution of the unlabeled image data; The calculation expression of the prediction distribution mean algorithm is:

[0067] Where is the prediction distribution result obtained by the i-th unlabeled image data in the t-th training round, and this prediction distribution result records the probabilities of the i-th unlabeled image data belonging to various category labels in the t-th training round; represents the total number of training rounds; is the mean vector of the predicted distribution of the i-th unlabeled image data.

[0068] In this embodiment, a pseudo-label enhancement training method using prediction inconsistent samples provided by the present application mainly determines unlabeled image data whose corresponding historical prediction data set meets the set conditions as pseudo-label image data and determines the pseudo-labels of the pseudo-label image data. For example, Figure 2 the performance of the historical prediction data set of the first Node in Figure 2 meets the set conditions. And for the first to fourth Nodes in Figure 3 as shown in Figure 3 the first column in Figure 3 shows the relationship between the training rounds of each of the four unlabeled image data and the predicted label categories obtained by prediction; the second column shows the mean values of the prediction probabilities of each of the four unlabeled image data under different predicted label categories. For unlabeled image data whose corresponding historical prediction data set meets the set conditions, it shows an obvious bimodal feature throughout the training process. That is, during the training process, the mean values of the predicted distributions of the unlabeled image data under two predicted label categories will be significantly higher. For this performance, the present application designs a spatial feature to capture the unlabeled image data with this performance and the corresponding pseudo-labels, but this performance cannot obtain accurate unlabeled image data that meet the set conditions and the corresponding pseudo-labels. For example, Figure 3 both the first Node and the fourth Node in Figure 3 have such a performance. Therefore, the present application also introduces another performance. For example, Figure 3 the third column in Figure 3 shows the change trend values of the predicted distributions of each of the four unlabeled image data under different predicted label categories. For unlabeled image data whose corresponding historical prediction data set meets the set conditions, during the entire training process, there will also be an obvious larger change trend value in the two predicted categories with the largest mean values of the predicted distributions. However, for the fourth Node that obviously does not meet the set conditions (that is, the unlabeled image data does not stabilize from the initial one predicted category to another predicted category during the subsequent training process), it will also have such a performance, but this performance will be particularly obvious. Therefore, while the present application designs a time feature to capture the unlabeled image data with this performance and the corresponding pseudo-labels, it filters out unlabeled image data whose change trend exceeds a certain value (that is, to filter out unlabeled image data whose predicted categories fluctuate continuously between two categories during the training process, similar to Figure 3 the unlabeled image data of the fourth Node in Figure 3 ). Thus, for the screening of unlabeled image data whose corresponding historical prediction data set meets the set conditions, the present application provides two features for screening, namely the spatial feature and the time feature These two features are calculated from the historical prediction dataset of unlabeled image data. For each unlabeled image data, pseudo-labels with inconsistent training dynamics in predictive patterns are screened from both spatial and temporal aspects, and the unlabeled image data to which the pseudo-labels belong participate in subsequent model training. Specifically, the calculation and screening process is to calculate the historical prediction dataset of each unlabeled image data respectively to obtain the values of each unlabeled image data on these two features, and then determine which unlabeled image data satisfy the corresponding conditions on these two features, so as to determine the unlabeled image data whose values on these two features satisfy the corresponding conditions as pseudo-label image data. Among them, Figure 2 and Figure 3 the first Node in Figure 3 is Node876 in the figure, the second Node is Node121 in the figure, the third Node is Node1016 in the figure, and the fourth Node is Node2359 in the figure.

[0069] Specifically, since the calculation of the spatial and temporal features of each unlabeled image data is the same, an unlabeled image data is taken as an example for illustration. In the case where the first model is trained successfully to obtain the corresponding second model, the historical prediction dataset of the unlabeled image data is calculated by the predictive distribution mean algorithm to obtain the mean vector of the predictive distribution of the unlabeled image data. The calculation expression of the predictive distribution mean algorithm is:

[0070] where, represents the predictive distribution result obtained by the i-th unlabeled image data in the t-th training round. The predictive distribution result records the probabilities of the i-th unlabeled image data belonging to various class labels in the t-th training round. For example, when the class labels include class 1, class 2, and class 3, the predictive distribution result of the i-th unlabeled image data obtained in each training round will record the probability of the i-th unlabeled image data belonging to class 1, the probability of belonging to class 2, and the probability of belonging to class 3 in the corresponding training round; represents the total number of training rounds. For example, when the first model obtains a successfully trained second model after 100 training rounds, the total number of training rounds is 100; is the mean vector of the predictive distribution of the i-th unlabeled image data. The mean vector of the predictive distribution records the mean of the predictive probabilities of the unlabeled image data under each class label. For example, when the class labels include class 1, class 2, and class 3, then The mean of the predicted probabilities of the i-th unlabeled image data under class 1, the mean of the predicted probabilities of the i-th unlabeled image data under class 2, and the mean of the predicted probabilities of the i-th unlabeled image data under class 3 recorded in

[0071] Step S32: Calculate the mean vector of the obtained predicted distribution through the target entropy algorithm to obtain the spatial features of the unlabeled image data; The calculation expression of the target entropy algorithm is:

[0072] where is all other class labels except the class label corresponding to the maximum value in ; is the mean of the predicted probabilities of the i-th unlabeled image data under class label c; is the spatial feature of the i-th unlabeled image data; is the entropy after removing the maximum value.

[0073] In this embodiment, after obtaining the mean vector of the predicted distribution of the unlabeled image data through step S31, calculate the mean vector of the predicted distribution of the unlabeled image data through the target entropy algorithm to obtain the value of the spatial features of the unlabeled image data. The calculation expression of the target entropy algorithm is:

[0074] where is all other class labels except the class label corresponding to the maximum value in the mean vector of the predicted distribution of the i-th unlabeled image data . For example, when the class labels include class 1, class 2, and class 3, and at the same time the mean of the predicted probabilities of the i-th unlabeled image data under class 1 in is the largest, then the class labels included in at this time are class 2 and class 3; is the mean of the predicted probabilities of the i-th unlabeled image data under class label c; is the entropy after removing the maximum value.

[0075] Step S33: Calculate the historical prediction data set of the unlabeled image data through the change trend algorithm to obtain the change trend vector of the predicted distribution of the unlabeled image data; The calculation expression of the change trend algorithm is:

[0076] Among them, and are consecutive training rounds of unlabeled image data, is the next training round of and respectively represent the prediction distribution results of the i-th unlabeled image data in the training rounds and respectively, represents the change trend vector of the prediction distribution of the i-th unlabeled image data.

[0077] In this embodiment, when the first model is trained successfully to obtain the corresponding second model, the historical prediction data set of the unlabeled image data is calculated by the change trend algorithm at the same time to obtain the change trend vector of the prediction distribution of the unlabeled image data. The calculation expression of the change trend algorithm is:

[0078] Among them, and are consecutive training rounds of unlabeled image data, is the next training round of and respectively represent the prediction distribution results of the i-th unlabeled image data in the training rounds and respectively, represents the change trend vector of the prediction distribution of the i-th unlabeled image data. The change trend vector of the prediction distribution records the change trend values of the prediction distribution of the unlabeled image data under each class label. For example, when the class labels include class 1, class 2, and class 3, then records the change trend value of the prediction distribution of the i-th unlabeled image data under class 1, the change trend value of the prediction distribution of the i-th unlabeled image data under class 2, and the change trend value of the prediction distribution of the i-th unlabeled image data under class 3.

[0079] Step S34: Calculate the obtained change trend vector through the first algorithm to obtain the first operator value and the second operator value of the unlabeled image data; The calculation expression of the first algorithm is:

[0080] Among them, is the first operator of the i-th unlabeled image data; is the second operator for the i-th unlabeled image data; is the final predicted class of the unlabeled image data during training, and this predicted class is defined as the predicted class in the second stage; is except Among them, the predicted class corresponding to the largest change trend in the change trend vector of the predicted distribution of the unlabeled image data is defined as the predicted class in the first stage; is the change trend of the predicted probability of the i-th unlabeled image data on the predicted class in the first stage; is the change trend of the predicted probability of the i-th unlabeled image data on the predicted class in the second stage; is the change trend of the predicted distribution of the i-th unlabeled image data on other predicted classes except the predicted classes in the first stage and the second stage.

[0081] In this embodiment, after calculating the change trend vector of the predicted distribution of the unlabeled image data through step S33, the obtained change trend vector of the predicted distribution of the unlabeled image data is calculated by the first algorithm to obtain the first operator value and the second operator value of the unlabeled image data. The calculation expression of this first algorithm is:

[0082] where is the first operator of the i-th unlabeled image data, which is used to capture the changes in two classes of the bimodal; is the second operator of the i-th unlabeled image data, which is used to capture the changes in non-bimodal classes; is the final predicted class of the unlabeled image data during training, and this predicted class is defined as the predicted class in the second stage; is except Among them, the predicted class corresponding to the largest change trend in the change trend vector of the predicted distribution of the unlabeled image data is defined as the predicted class in the first stage; is the value of the change trend of the predicted probability of the i-th unlabeled image data on the predicted class in the first stage is the value of the change trend of the predicted probability of the i-th unlabeled image data on the predicted class in the second stage; is the value of the change trend of the predicted distribution of the i-th unlabeled image data on other predicted classes except the predicted classes in the first stage and the second stage.

[0083] In this embodiment, there is a special case where the predicted class label in the second stage of an unlabeled image data (such as unlabeled image data A) is not only the final predicted class of the unlabeled image data A during the training process, but also the predicted class corresponding to the largest change trend in the change trend vector of the predicted distribution of the unlabeled image data A. In this case, it will no longer be the predicted class corresponding to the largest change trend in the change trend vector of the predicted distribution of the unlabeled image data A, but will use the predicted class corresponding to the second largest change trend in the change trend vector of the predicted distribution of the unlabeled image data A as the predicted class of the unlabeled image data A, because the predicted class corresponding to the largest trend has been occupied by the predicted class in the second stage of the unlabeled image data A.

[0084] Step S35: Calculate the first operator value and the second operator value through a second algorithm to obtain the time feature of the unlabeled image data; The calculation expression of the second algorithm is:

[0085]

[0086]

[0087]

[0088]

[0089] where and are thresholds; is the change index of the i-th unlabeled image data; is the direction index of the i-th unlabeled image data; T is the last training round during the training process; is the predicted probability of the i-th unlabeled image data on the predicted class in the second stage in the t-th training round; is the predicted probability of the i-th unlabeled image data on the predicted class in the first stage in the t-th training round; is the time feature of the i-th unlabeled image data; and are weights.

[0090] In this embodiment, after calculating the first operator value and the second operator value of the unlabeled image data through step S34, calculate the first operator value and the second operator value of the unlabeled image data through a second algorithm to obtain the time feature of the unlabeled image data. The calculation expression of the second algorithm is a system of equations, as follows:

[0091]

[0092]

[0093]

[0094]

[0095] wherein and are two preset thresholds; is the change index of the i-th unlabeled image data; is the direction index of the i-th unlabeled image data; T is the last training epoch in the training process; is the predicted probability of the i-th unlabeled image data on the predicted class in the second stage in the t-th training epoch; is the predicted probability of the i-th unlabeled image data on the predicted class in the first stage in the t-th training epoch; is the time feature of the i-th unlabeled image data; and are weights. Through the same implementation manner, each unlabeled image data can calculate its own corresponding time feature based on its own historical prediction data set.

[0096] In this embodiment, for the sample data with the characteristic of inconsistent predictions, since the predicted class in the first stage is , and the final predicted class is , then at certain moments in the training process, the probability of is greater than , is negative, and at the last moment, is positive, and the difference between the two can reflect this change in sign, that is, the larger, the more likely the corresponding sample data is the sample data with the characteristic of inconsistent predictions. To calculate , this application takes the difference between the change difference of the bimodal class at the last step and the minimum value of in the training process to obtain an index reflecting whether there is a class change.

[0097] Step S36: Determine whether the predicted distribution result of the unlabeled image data meets the set conditions according to the spatial feature and the time feature of the unlabeled image data.

[0098] In this embodiment, based on the values of the spatial features and the temporal features of the unlabeled image data obtained by calculation, it is determined whether the historical prediction data set of the unlabeled image data meets the set conditions, so as to determine whether it can be determined as pseudo-labeled image data.

[0099] Combined with the above embodiments, in one implementation, the embodiments of the present application further provide a pseudo-label enhanced training method using prediction inconsistent samples. In this pseudo-label enhanced training method using prediction inconsistent samples, step S36 may include: calculating the spatial features and the temporal features of the unlabeled image data through a prediction inconsistency metric algorithm to obtain the value of the prediction inconsistency metric of the unlabeled image data; The calculation expression of the prediction inconsistency metric algorithm is:

[0100] where, and are weights; According to the relationship between the value of the prediction inconsistency metric and the set threshold, it is determined whether the prediction distribution result of the unlabeled image data meets the set conditions.

[0101] In this embodiment, the present application defines a prediction inconsistency metric for unlabeled image data, which is calculated based on the spatial features and the temporal features of the unlabeled image data. Based on whether the value of the prediction inconsistency metric of the unlabeled image data is lower than the set threshold, it is determined whether the unlabeled image data meets the set conditions. Among them, the set threshold is preferably the mean value of the prediction inconsistency metric on the unlabeled image data set. This is only a preferred value, and other preset values can also be used. Specifically, after calculating the spatial features and the temporal features of the unlabeled image data through steps S31 to S36, the spatial features and the temporal features are brought into the prediction inconsistency metric algorithm for calculation to obtain the value of the prediction inconsistency metric of the unlabeled image data. Then, the value of the prediction inconsistency metric is compared with the set threshold. When the value of the prediction inconsistency metric does not exceed the set threshold, it is determined that the unlabeled image data meets the set conditions. At this time, the unlabeled image data (such as image data a) is determined as pseudo-labeled image data (i.e., image data a), and at the same time, the prediction category of the second stage of the unlabeled image data (i.e., image data a) is determined as the pseudo-label of the pseudo-labeled image data (i.e., image data a).

[0102] In this embodiment, the prediction inconsistency metric proposed by the pseudo-label enhancement training method using prediction-inconsistent samples provided by the present application describes the characteristics of the training dynamics with a prediction-inconsistent pattern in space and time. Through this prediction inconsistency metric, this novel type of pseudo-label that conforms to the prediction-inconsistent dynamics can be screened out, which is a new type of prediction label with a prediction-inconsistent training dynamics pattern. This pseudo-label and the corresponding unlabeled image data can help the model learn the correct correlations while eliminating the incorrect correlations, complementing the deficiency of screening pseudo-labels by confidence scores.

[0103] It should be noted that for method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequence, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present application.

[0104] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0105] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0106] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.

[0107] The above has introduced in detail a pseudo-label enhancement training method using prediction-inconsistent samples provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A pseudo-label enhancement training method using prediction inconsistent samples, characterized in that: The method comprises: Determine the pseudo labels of the unlabeled image data through the baseline pseudo label selection method to construct a baseline pseudo label training set; The first model is trained by a target training set consisting of a real label training set and a baseline pseudo-label training set, and the first model after each training round is used to perform label prediction on the unlabeled image data to obtain a corresponding prediction distribution result, wherein all prediction distribution results obtained for the unlabeled image data in all training rounds constitute a corresponding historical prediction data set; When the first model is trained successfully to obtain the corresponding second model, determining whether the historical prediction data set of the unlabeled image data meets a set condition, wherein the set condition is that the label prediction of the unlabeled image data during the training process is successively stabilized on two different category labels; Determine the unlabeled image data that meets the set conditions as pseudo-labeled image data, and determine the category label that the final label prediction stabilizes on as the pseudo-label of the pseudo-labeled image data; An enhanced training set is formed by the obtained large amount of pseudo-label image data and the target training set, and the constructed initial model is trained by the enhanced training set to obtain a qualified label prediction model.

2. The pseudo-label enhancement training method using prediction inconsistent samples according to claim 1, characterized in that: The constructed initial model is trained by the enhanced training set to obtain a qualified label prediction model, including: The constructed initial model is trained by the enhanced training set, and if a qualified label prediction model cannot be obtained after a preset number of trainings, the initial model after the preset number of trainings is determined to be the third model; The third model is trained by the target training set, and the label prediction is performed on the new unlabeled image data by the third model after each training round to obtain the corresponding prediction distribution result, and all the prediction distribution results obtained for the new unlabeled image data in all training rounds constitute the corresponding historical prediction data set; If the third model is trained successfully, determining whether the historical prediction data set of the new unlabeled image data meets the set conditions; Determine new unlabeled image data that meets the set conditions as pseudo-labeled image data, and determine the category label that the final label prediction stabilizes on as the pseudo-label of the pseudo-labeled image data; A new enhanced training set is formed by obtaining a large amount of new pseudo-label image data and the enhanced training set, and the third model after a preset number of trainings is trained by the new enhanced training set to obtain a qualified label prediction model.

3. The pseudo-label enhancement training method using prediction inconsistent samples according to claim 1, characterized in that: The method further comprises: Testing the trained label prediction model using labeled image data in the test set to evaluate the performance of the label prediction model on unseen image data; According to the performance evaluation result, it is determined whether to perform new training on the label prediction model.

4. The pseudo-label enhancement training method using prediction inconsistent samples according to claim 1, characterized in that: The pseudo labels of the unlabeled image data are determined by the baseline pseudo label selection method to construct a baseline pseudo label training set, including: Perform label prediction on unlabeled image data through the first model; Determine the predicted label category that satisfies the baseline pseudo label selection method in the label prediction result as the pseudo label of the unlabeled image data; Based on the determined unlabeled image data and their respective pseudo labels, a baseline pseudo-label training set is constructed.

5. The method for pseudo-label enhancement training using prediction inconsistent samples according to claim 4, characterized in that: Before performing label prediction on the unlabeled image data by the first model, the method further includes: Pre-train the constructed initial model using labeled image data from the real label training set; The pre-trained qualified initial model is determined as the first model.

6. The pseudo-label enhancement training method using prediction inconsistent samples according to claim 1, characterized in that: Build an initial model, including: Based on the type of the selected neural network model, construct a basic structure of the model, wherein the basic structure includes: an input layer, a convolution layer and a pooling layer, an activation function, and a fully connected layer; Adding a softmax layer after the fully connected layer in the basic structure converts the model output into a probability distribution; Set the loss function of the model to construct the final initial model.

7. The method for pseudo-label enhancement training using prediction inconsistent samples according to claim 1, characterized in that: In the case where the first model is trained successfully and the corresponding second model is obtained, determining whether the historical prediction data set of the unlabeled image data meets the set conditions includes: When the first model is trained to obtain the corresponding second model, a historical prediction data set of the unlabeled image data is calculated by a prediction distribution mean algorithm to obtain a mean vector of the prediction distribution of the unlabeled image data; The calculation expression of the prediction distribution mean algorithm is: in, is the predicted distribution result obtained for the i-th unlabeled image data in the t-th training round. The predicted distribution result records the probability of the i-th unlabeled image data belonging to various category labels in the t-th training round. represents the total number of training rounds; is the mean vector of the predicted distribution of the i-th unlabeled image data; Calculating the mean vector of the obtained predicted distribution by a target entropy algorithm to obtain the spatial features of the unlabeled image data; The calculation expression of the target entropy algorithm is: in For All other category labels except the category label corresponding to the maximum value; is the mean of the predicted probability of the i-th unlabeled image data under the category label c; is the spatial feature of the i-th unlabeled image data; To remove the maximum entropy; Calculating the historical prediction data set of the unlabeled image data by using a change trend algorithm to obtain a change trend vector of the predicted distribution of the unlabeled image data; The calculation expression of the change trend algorithm is: in, and is the number of consecutive training rounds on unlabeled image data, yes The next training round of and Respectively represent the i-th unlabeled image data in the training round and The predicted distribution results when Represents the change trend vector of the predicted distribution of the i-th unlabeled image data; Calculating the obtained change trend vector by using a first algorithm to obtain a first operator value and a second operator value of the unlabeled image data; The calculation expression of the first algorithm is: in, is the first operator of the i-th unlabeled image data; is the second operator of the i-th unlabeled image data; The final predicted category of the unlabeled image data during the training process is defined as the predicted category of the second stage; For In addition, the prediction category corresponding to the maximum change trend in the change trend vector of the prediction distribution of the unlabeled image data is defined as the prediction category of the first stage; is the changing trend of the predicted probability of the i-th unlabeled image data in the predicted category in the first stage; is the changing trend of the predicted probability of the i-th unlabeled image data in the predicted category in the second stage; is the changing trend of the predicted distribution of the i-th unlabeled image data in other predicted categories except the predicted categories in the first and second stages; Calculating the first operator value and the second operator value by a second algorithm to obtain the time feature of the unlabeled image data; The calculation expression of the second algorithm is: in and is the threshold value; is the change index of the i-th unlabeled image data; is the direction indicator of the i-th unlabeled image data; T is the last training round in the training process; is the predicted probability of the i-th unlabeled image data on the predicted category in the second stage in the t-th training round; is the predicted probability of the i-th unlabeled image data on the predicted category in the first stage in the t-th training round; is the temporal feature of the i-th unlabeled image data; and is the weight; According to the spatial characteristics and temporal characteristics of the unlabeled image data, it is determined whether the predicted distribution result of the unlabeled image data meets the set conditions.

8. The method for pseudo-label enhancement training using prediction inconsistent samples according to claim 7, characterized in that: Determining whether the predicted distribution result of the unlabeled image data meets a set condition according to the spatial characteristics and the temporal characteristics of the unlabeled image data includes: Calculating the spatial features and temporal features of the unlabeled image data by using a prediction inconsistency index algorithm to obtain a prediction inconsistency index value of the unlabeled image data; The calculation expression of the prediction inconsistency index algorithm is: in, and is the weight; According to the relationship between the prediction inconsistency index value and the set threshold, it is determined whether the predicted distribution result of the unlabeled image data meets the set condition.

Citation Information

Patent Citations

  • Semi-supervised learning method based on pseudo label weighting

    CN112232416A

  • Semi-supervised model training method, image recognition method and device

    CN116935146A

  • Lotus phenotype identification method and device based on pseudo tag algorithm and MobileNetV2 network

    CN117953281A

  • Image pseudo label labeling method based on virtual adversarial training

    CN119418338A

  • Method and device for training semi-supervised object detection model, and object detection method and device

    WO2024222444A1

Cited By

  • Label propagation based method for few-shot data annotation expansion and noise label suppression

    CN122615428A