A Pseudo-Label Enhancement Training Method Using Prediction-Inconsistent Samples
Through the three-stage training method, combining baseline pseudo-label and predicting inconsistent pseudo-label image data, the problem of misleading confidence scores in pseudo-label selection is solved, and the training effect of the model and the recognition ability of unseen data is improved.
Patent Information
- Application Number
- CN202510525771.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-25
AI Technical Summary
In the existing pseudo-label selection methods, high confidence scores are often assigned to mispredicted labels, resulting in poor model training results and susceptible to manipulation of adversarial samples. The prior art is difficult to effectively use unlabeled data to improve model performance.
Using a three-stage training method, the enhanced training set is constructed through baseline pseudo-label selection and prediction of inconsistent pseudo-label determination, including manual annotation of real tags and image data for predicting inconsistent pseudo-labels, and the model training is performed using the features of predicting inconsistent samples.
The training effect of the model is improved, the performance of the label prediction model is enhanced, especially in image classification tasks, and the performance of the model on unseen data is improved.
Smart Images

Figure CN120087502B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology, and in particular, to a pseudo-label enhancement training method using prediction-inconsistent samples. Background Art
[0002] In the field of deep learning technology, the training of some models requires a large amount of labeled sample data with labels. In practical applications, the cost of obtaining a large amount of labeled sample data is high, while semi-supervised learning can utilize unlabeled data to improve the model performance with limited labeled data. As a commonly used method in semi-supervised learning, the pseudo-label technology is widely used in semi-supervised learning. The core of the pseudo-label technology is pseudo-label selection, and the goal of pseudo-label selection is to determine which samples should be assigned pseudo-labels.
[0003] Currently, the most commonly used pseudo-label selection metric is the confidence score obtained from the softmax distribution. However, this metric has the problem of poor calibration, that is, high confidence scores are often assigned to mispredicted labels, resulting in sample data having incorrect pseudo-labels. Using such pseudo-label sample data for training the model will affect the model training effect, and at the same time, the confidence score is also vulnerable to manipulation by adversarial samples. Summary of the Invention
[0004] In view of this, this application provides a pseudo-label enhancement training method using prediction-inconsistent samples, aiming to provide pseudo-label data with different features to train the model to improve the model training effect.
[0005] In the first aspect of this application, a pseudo-label enhancement training method using prediction-inconsistent samples is provided, and the method includes:
[0006] Determine the pseudo-labels of unlabeled image data through a baseline pseudo-label selection method to construct a baseline pseudo-label training set;
[0007] Train the first model with a target training set composed of a true-label training set and a baseline pseudo-label training set, and use the first model after each training round to predict the labels of unlabeled image data to obtain corresponding prediction distribution results. All the prediction distribution results obtained by the unlabeled image data in all training rounds form a corresponding historical prediction data set;
[0008] When the first model is trained to be qualified to obtain a corresponding second model, determine whether the historical prediction data set of unlabeled image data meets the set condition, where the set condition is that the label predictions of unlabeled image data during training are successively stable on two different category labels;
[0009] Determine the unlabeled image data that meets the set conditions as pseudo-labeled image data, and determine the class label at which the final label prediction stabilizes as the pseudo-label of the pseudo-labeled image data;
[0010] Use the obtained large amount of pseudo-labeled image data and the target training set to form an augmented training set, and train the constructed initial model through the augmented training set to obtain a qualified label prediction model.
[0011] Optionally, training the constructed initial model through the augmented training set to obtain a qualified label prediction model includes:
[0012] Train the constructed initial model through the augmented training set. If a qualified label prediction model cannot be obtained after a preset number of trainings, determine the initial model after the preset number of trainings as the third model;
[0013] Train the third model through the target training set, and use the third model after each training round to perform label prediction on new unlabeled image data to obtain corresponding prediction distribution results. All the prediction distribution results obtained by the new unlabeled image data in all training rounds form the corresponding historical prediction data set;
[0014] When the third model is trained qualified, determine whether the historical prediction data set of the new unlabeled image data meets the set conditions;
[0015] Determine the new unlabeled image data that meets the set conditions as pseudo-labeled image data, and determine the class label at which the final label prediction stabilizes as the pseudo-label of the pseudo-labeled image data;
[0016] Use the obtained large amount of new pseudo-labeled image data and the augmented training set to form a new augmented training set, and train the third model after the preset number of trainings through the new augmented training set to obtain a qualified label prediction model.
[0017] Optionally, the method further includes:
[0018] Test the trained qualified label prediction model with the labeled image data in the test set to evaluate the performance of the label prediction model on unseen image data;
[0019] Determine whether to perform new training on the label prediction model according to the performance evaluation results.
[0020] Optionally, determine the pseudo-labels of the unlabeled image data through a baseline pseudo-label selection method to construct a baseline pseudo-label training set, including:
[0021] Predict labels for unlabeled image data through a first model;
[0022] Determine the predicted label categories that meet the baseline pseudo-label selection method in the label prediction results as the pseudo-labels of the unlabeled image data;
[0023] Construct a baseline pseudo-label training set based on the determined unlabeled image data and their respective pseudo-labels.
[0024] Optionally, before predicting labels for unlabeled image data through the first model, the method further includes:
[0025] Pre-train the constructed initial model with the labeled image data in the true label training set;
[0026] Determine the initial model that passes the pre-training as the first model.
[0027] Optionally, constructing an initial model includes:
[0028] Based on the selected type of neural network model, construct the basic structure of the model, and the basic structure includes: an input layer, a convolutional layer and a pooling layer, an activation function, and a fully connected layer;
[0029] Add a softmax layer after the fully connected layer in the basic structure to convert the model output into a probability distribution;
[0030] Set the loss function of the model to construct and obtain the final initial model.
[0031] Optionally, when the first model is trained successfully to obtain the corresponding second model, determine whether the historical prediction data set of the unlabeled image data meets the set conditions, including:
[0032] When the first model is trained successfully to obtain the corresponding second model, calculate the historical prediction data set of the unlabeled image data through the prediction distribution mean algorithm to obtain the mean vector of the prediction distribution of the unlabeled image data;
[0033] The calculation expression of the prediction distribution mean algorithm is:
[0034]
[0035] Where, is the prediction distribution result obtained by the i-th unlabeled image data in the t-th training round, and this prediction distribution result records the probabilities of the i-th unlabeled image data belonging to various category labels in the t-th training round; represents the total number of training rounds; is the mean vector of the prediction distribution of the i-th unlabeled image data;
[0036] Calculate the mean vector of the obtained prediction distribution through the target entropy algorithm to obtain the spatial features of the unlabeled image data;
[0037] The calculation expression of the target entropy algorithm is:
[0038]
[0039] where is all other class labels except the class label corresponding to the maximum value in ; is the mean of the predicted probabilities of the i-th unlabeled image data under the class label c; is the spatial feature of the i-th unlabeled image data; is the entropy after removing the maximum value;
[0040] Calculate the historical prediction dataset of the unlabeled image data through the change trend algorithm to obtain the change trend vector of the prediction distribution of the unlabeled image data;
[0041] The calculation expression of the change trend algorithm is:
[0042]
[0043] where, and are consecutive training rounds of the unlabeled image data, is 's next training round; and respectively represent the prediction distribution results of the i-th unlabeled image data in the training rounds and ; represents the change trend vector of the prediction distribution of the i-th unlabeled image data;
[0044] Calculate the obtained change trend vector through the first algorithm to obtain the first operator value and the second operator value of the unlabeled image data;
[0045] The calculation expression of the first algorithm is:
[0046]
[0047] where, is the first operator of the i-th unlabeled image data; is the second operator of the i-th unlabeled image data; Define the predicted category obtained during the training process for the unlabeled image data as the predicted category for the second stage; Excluding Define the predicted category corresponding to the maximum change trend in the change trend vector of the predicted distribution of the unlabeled image data as the predicted category for the first stage; is the change trend of the predicted probability of the i-th unlabeled image data on the predicted category for the first stage; is the change trend of the predicted probability of the i-th unlabeled image data on the predicted category for the second stage; is the change trend of the predicted distribution of the i-th unlabeled image data on other predicted categories except the predicted categories for the first stage and the second stage;
[0048] Calculate the first operator value and the second operator value through the second algorithm to obtain the time feature of the unlabeled image data;
[0049] The calculation expression of the second algorithm is:
[0050]
[0051]
[0052]
[0053]
[0054]
[0055] where and are thresholds; is the change index of the i-th unlabeled image data; is the direction index of the i-th unlabeled image data; T is the last training round during the training process; is the predicted probability of the i-th unlabeled image data on the predicted category for the second stage in the t-th training round; is the predicted probability of the i-th unlabeled image data on the predicted category for the first stage in the t-th training round; is the time feature of the i-th unlabeled image data; and are weights;
[0056] Determine whether the predicted distribution result of the unlabeled image data meets the set conditions according to the spatial feature and the time feature of the unlabeled image data.
[0057] Optionally, determining whether the prediction distribution result of the unlabeled image data meets the set conditions according to the spatial and temporal characteristics of the unlabeled image data includes:
[0058] Calculating the spatial and temporal characteristics of the unlabeled image data through a prediction inconsistency index algorithm to obtain the prediction inconsistency index value of the unlabeled image data;
[0059] The calculation expression of the prediction inconsistency index algorithm is:
[0060]
[0061] Wherein, and are weights;
[0062] Determining whether the prediction distribution result of the unlabeled image data meets the set conditions according to the relationship between the prediction inconsistency index value and the set threshold.
[0063] For the prior art, the present application has the following advantages:
[0064] A method for enhancing training by using pseudo-labels of prediction-inconsistent samples provided by an embodiment of the present application. First, through a baseline pseudo-label selection method (such as selecting the class label with the highest confidence obtained by the model's prediction of unlabeled image data as the pseudo-label of the unlabeled image data), the pseudo-labels of the unlabeled image data are determined to construct a baseline pseudo-label training set; the first model is trained with a target training set composed of a true-label training set and a baseline pseudo-label training set, and the first model after each training round is used to predict the labels of the unlabeled image data to obtain corresponding prediction distribution results. All the prediction distribution results obtained by the unlabeled image data in all training rounds form a corresponding historical prediction data set, and all the prediction distribution results of a single unlabeled image data itself in all training rounds form its own corresponding historical prediction data set; when the first model is trained successfully to obtain a corresponding second model, it is determined whether the historical prediction data set of the unlabeled image data itself meets the set conditions. The set conditions are that the label predictions of the unlabeled image data are successively stable on two different class labels during the training process, that is, stable on one of the two class labels in the earlier training rounds during the training process, and stable on the other of the two class labels in the later training rounds during the training process; the unlabeled image data whose own historical prediction data set meets the set conditions is determined as pseudo-label image data, and the class label on which the final label prediction is stable is determined as the pseudo-label of the pseudo-label image data; a large number of obtained pseudo-label image data and the target training set are used to form an enhanced training set, and the initial model constructed is trained with the enhanced training set to obtain a trained label prediction model. Thus, the present application adopts a three-stage training method (pre-training model, determining baseline pseudo-labels based on the pre-training model, and determining prediction-inconsistent pseudo-labels), so that the data used for training the model includes not only the image data with true labels manually annotated, but also the image data with baseline pseudo-labels and the image data with prediction-inconsistent pseudo-labels. The pseudo-label image data obtained by these two different pseudo-label determination methods have different characteristics (for example, the characteristics of the baseline pseudo-label image data are mainly reflected in the confidence, and the characteristics of the new pseudo-label image data selected by the present application through prediction inconsistency are mainly reflected in the inconsistency of the predicted categories before and after), which can further enhance the training effect of the label prediction model on the basis of the traditional pseudo-label set.
[0065] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically gives the specific implementation manners of the present application. Description of the Drawings
[0066] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art.
[0067] Figure 1 It is a flowchart of a pseudo-label enhancement training method using predicted inconsistent samples provided by an embodiment of the present application;
[0068] Figure 2 It is a schematic diagram of the relationship between the training rounds and the predicted categories of unlabeled image data in a pseudo-label enhancement training method using predicted inconsistent samples provided by an embodiment of the present application;
[0069] Figure 3 It is another schematic diagram of the relationship between the training rounds and the predicted categories of unlabeled image data in a pseudo-label enhancement training method using predicted inconsistent samples provided by an embodiment of the present application. Detailed implementation manners
[0070] The following will describe the exemplary embodiments of the present application in more detail with reference to the accompanying drawings.
[0071] Figure 1 It is a flowchart of a pseudo-label enhancement training method using predicted inconsistent samples provided by an embodiment of the present application. As Figure 1 shown, the method includes:
[0072] Step S1: Determine the pseudo-labels of the unlabeled image data through a baseline pseudo-label selection method to construct a baseline pseudo-label training set.
[0073] In this embodiment, a method for enhancing training using pseudo-labels of predicted inconsistent samples provided by this application is applied to an image classification task, that is, training and applying a neural network model for image classification. The image classification task includes single-label binary classification tasks and single-label multi-classification tasks. A single-label binary classification task refers to a task of classifying images into two categories. The number of labels or categories of all images in the dataset is only two. The image input into the neural network model carries only one category label, and the neural network model needs to identify the image as one of the two categories based on the learned features. For example, in the cat-dog classification task, that is, each image can only be a cat or a dog, and the neural network model identifies each image as one of the two categories of cats and dogs. A single-label multi-classification task means that the number of labels or categories of all images in the dataset has multiple categories (more than two), but each image has only one label or category, and the neural network model needs to identify the image as one of the multiple categories based on the learned features. For example, in the handwritten digit recognition task, each image can only be any one of the handwritten digits from 0 to 9, and the neural network model identifies each image as one of the 10 categories of handwritten digits. Another example is the fruit classification task, such as classifying an image into one of multiple categories such as apples, bananas, and oranges.
[0074] In this embodiment, in the field of deep learning, an image is usually a three-dimensional tensor, including three dimensions: height, width, and channels, containing information such as space and color. An image is the input of a deep learning model, and the model processes these tensors to understand and analyze the image content and complete various computer vision tasks. A label is the target output or correct answer corresponding to the input image data in supervised learning. For example, in an image classification task, the label is the category of the image. A pseudo-label refers to generating a pseudo-label for unlabeled data using the result predicted by the model when there is no or only a small amount of labeled data. Training dynamics refers to the process by which a deep learning model iteratively updates the gradient to optimize the model parameters, and training dynamics studies the laws and behaviors of model parameters, loss functions, gradients, prediction results, etc. changing over time (or training epochs). Prediction-inconsistent pattern training dynamics: This type of training dynamics has obvious prediction-inconsistent characteristics, that is, in the first stage, the model prediction stabilizes in one category, and in the second stage, the model prediction mainly stabilizes in another category, that is, the prediction results show inconsistent predictions before and after, stabilizing in one category first and then in another category during the training process. And this application mainly proposes a prediction-inconsistent index based on the prediction-inconsistent pattern training dynamics. This prediction-inconsistent index forms a quantization index (i.e., the prediction-inconsistent index) for screening pseudo-labels of unlabeled image data by quantifying the unique characteristics of the training dynamics of unlabeled image data with prediction-inconsistent labels during the training process.
[0075] In this embodiment, first, an image data set is collected, which includes a small amount of labeled image data with true labels and a large amount of unlabeled image data. The small amount of labeled image data is used to divide the true label training set and the test set, and at the same time, it is ensured that images with different category labels are evenly distributed in these sets during the division. The true label training set is used to train the first model and the initial model, and the true label test set is used to test the finally trained label prediction model to evaluate its performance on unseen image data, such as metrics like accuracy, precision, recall, and F1 value in an image classification task. The initial model is an initial model constructed for the image classification task, and the model type of the initial model can be CNN, ResNet, etc.
[0076] In this embodiment, an alternative implementation of the baseline pseudo-label selection method is to use the model after a certain degree of training to predict the labels of the unlabeled image data, and select the prediction category with the highest confidence in the obtained label prediction results as the pseudo-label of the unlabeled image data, and determine the unlabeled image data with this pseudo-label as the baseline pseudo-label image data. Through this baseline pseudo-label selection method, a large number of baseline pseudo-label image data are determined to construct a baseline pseudo-label training set. Among them, determining the pseudo-label of the unlabeled image data by the highest confidence is only an alternative implementation of the baseline pseudo-label selection method, and the baseline pseudo-label selection method can be other baseline pseudo-label selection methods, which are not specifically limited here.
[0077] Step S2: Train the first model with the target training set composed of the true label training set and the baseline pseudo-label training set, and use the first model after each training round to predict the labels of the unlabeled image data to obtain the corresponding prediction distribution results. All the prediction distribution results obtained by the unlabeled image data in all training rounds constitute the corresponding historical prediction data set.
[0078] In this embodiment, the present application pre-establishes a true label training set, which records the labeled image data with true labels. After obtaining the baseline pseudo-label training set through step S1, the true label training set and the baseline pseudo-label training set are merged to jointly form a target training set. The target training set thus formed is used to train the first model, and the first model is preferably the model that predicts the labels of the unlabeled image data in step S1.
[0079] In this embodiment, the training process of training the first model with the constructed target training set is as follows: Input the image data in the target training set into the first model for training, and calculate the feature representation of the image data input into the first model through forward propagation. Calculate the training loss function using the label information of the input image data, and update the model parameters of the first model according to the loss value obtained from the loss function. Then calculate the gradient through the backpropagation algorithm, and use an optimization algorithm (such as SGD (Stochastic Gradient Descent), Adam (Adaptive Moment Estimation), etc.) to update the model parameters. During the training process, when the loss value obtained from calculating the loss function meets the corresponding set conditions, such as being lower than a certain set threshold, it is determined that the first model is qualified for training, and the corresponding second model is obtained. Here, the set threshold can be set according to the actual application scenario and will not be specifically limited herein. During the entire training process of the first model to obtain the finally qualified second model, multiple training rounds are required to obtain the finally qualified second model. One training round means that each time the loss value is calculated and the model parameters of the first model are updated based on the calculated loss value, it is counted as one training round.
[0080] In this embodiment, after each training round, the first model after the current training round is used to predict the labels of the unlabeled image data, and the prediction distribution result of the unlabeled image data in the current training round is obtained. A prediction distribution result records the probabilities that the unlabeled image data predicted by the first model after one training round belongs to various label categories. Through the same implementation method, each unlabeled image data in a large number of unlabeled image data will have multiple prediction distribution results with the same number as the total number of training rounds. The multiple prediction distribution results of a single unlabeled image data with the same number as the total number of training rounds constitute the historical prediction data set of the single unlabeled image data. The historical prediction data sets of each unlabeled image data are used to determine whether the unlabeled image data itself belongs to pseudo-labeled image data and for pseudo-label selection of pseudo-labeled image data.
[0081] Exemplarily, since the first model predicts the labels of unlabeled image data in the same way after each training round during the training process, an unlabeled image data a is taken as an example for illustration here. Assuming that the number of training rounds from the first model to the finally trained and qualified second model is 2 times, this 2 times is only an exemplary explanation for easy understanding, and the actual number of training rounds is generally much larger than 2 times. Correspondingly, after the first training round, the first model will make a label prediction for the unlabeled image data a once, obtaining a corresponding prediction distribution result A1, and after the second training round, the first model will make a label prediction for the unlabeled image data a once, obtaining a corresponding prediction distribution result A2. That is, no matter how many training rounds are carried out (such as 100 times), there will be corresponding numbers of prediction distribution results for an unlabeled image data (such as 100).
[0082] Step S3: When the first model is trained and qualified to obtain the corresponding second model, determine whether the historical prediction dataset of the unlabeled image data meets the set condition, where the set condition is that the label predictions of the unlabeled image data during the training process are successively stable on two different category labels.
[0083] In this embodiment, after the first model is trained and qualified to obtain the corresponding second model, each unlabeled image data will have a historical prediction dataset corresponding to itself. Then, when the first model is trained and qualified to obtain the corresponding second model, at this time, it is determined whether the historical prediction datasets of each unlabeled image data meet the set condition respectively. The set condition is that the label predictions of the unlabeled image data during the training process are successively stable on two different category labels, that is, in the earlier training rounds during the training process, it is stable on one of the two category labels, and in the later training rounds during the training process, it is stable on the other of the two category labels. As Figure 2 shown Figure 2 in, the abscissa represents the number of training rounds, and the ordinate represents the label category to which the label prediction result obtained by the unlabeled image data in the corresponding training round belongs. The label category to which it belongs refers to the label category with the highest confidence score in the corresponding prediction distribution result. Figure 2 The historical prediction dataset of the unlabeled image data of the first Node in meets the set condition. Not only are the label predictions during the training process stable on two category labels (category 1 and category 2 respectively), but also the label predictions during the training process are predicted as one category (category 1) in the earlier training rounds and as the other category (category 2) in the later training rounds.
[0084] Step S4: Determine the unlabeled image data that meets the set conditions as pseudo-labeled image data, and determine the class label at which the final label prediction stabilizes as the pseudo-label of the pseudo-labeled image data.
[0085] In this embodiment, after determining the historical prediction dataset that meets the set conditions through step S3, the unlabeled image data corresponding to the historical prediction dataset that meets the set conditions is determined as pseudo-labeled image data. At the same time, the class label at which the label prediction manifested in the historical prediction dataset of this pseudo-labeled image data finally stabilizes is determined as the pseudo-label of this pseudo-labeled image data, label this pseudo-labeled image data with this pseudo-label, and use the labeled pseudo-labeled image data to participate in the subsequent training of the label prediction model. As Figure 2 the historical prediction dataset of the unlabeled image data of the first Node in shows that it meets the set conditions, so this Node can be determined as pseudo-labeled image data. At the same time, the label prediction of the unlabeled image data of this Node has stabilized at class 2, so class 2 is used as the pseudo-label of this pseudo-labeled image data.
[0086] Step S5: Use the obtained large amount of pseudo-labeled image data and the target training set to form an augmented training set, and train the constructed initial model through the augmented training set to obtain a qualified label prediction model.
[0087] In this embodiment, after obtaining a large amount of pseudo-labeled image data labeled with their own pseudo-labels through the same implementation manner as steps S2 to S4, the pseudo-label training set composed of these labeled large amounts of pseudo-labeled image data and the target training set used for the initial training of the first model together form an augmented training set. Then, use this augmented training set to train the initial model pre-constructed for the image classification task to obtain a qualified label prediction model that can finally be used to perform the image classification task.
[0088] In this embodiment, an optional implementation manner is that the label prediction model is used for the scenario of identifying the intersection categories of roads in the field of autonomous driving. In this application scenario, the image data used for training the model is the image data of the road intersections collected, and the label categories of these image data include crossroads category, Y-shaped intersection category, T-shaped intersection category, X-shaped intersection category, roundabout intersection category, etc. The trained label prediction model is installed in the vehicle terminal to identify the intersection categories of the collected image data during autonomous driving to guide the autonomous driving of the vehicle. It should be understood that a method for enhancing training with pseudo-labels using prediction-inconsistent samples provided in this application can also be used in other application scenarios for performing image classification tasks.
[0089] In this embodiment, the difference between the training of the first model and the training of the initial model in this application is that the training of the first model refers to predicting labels for unlabeled image data, and using the historical prediction data set composed of all the prediction distribution results corresponding to the unlabeled image data itself for the subsequent determination of pseudo-labeled image data and the training process of pseudo-label selection for the pseudo-labeled image data, while the training of the initial model is to train the model in order to obtain a label prediction model that can ultimately be used for image classification.
[0090] A method for enhancing training by using pseudo - labels of prediction - inconsistent samples provided by an embodiment of the present application. First, through a baseline pseudo - label selection method (such as selecting the class label with the highest confidence obtained by the model's prediction of unlabeled image data as the pseudo - label of the unlabeled image data), the pseudo - labels of the unlabeled image data are determined to construct a baseline pseudo - label training set; the first model is trained with a target training set composed of a true - label training set and a baseline pseudo - label training set, and the first model after each training round is used to predict the labels of the unlabeled image data to obtain the corresponding prediction distribution results. All the prediction distribution results obtained by the unlabeled image data in all training rounds form the corresponding historical prediction data set, and all the prediction distribution results of a single unlabeled image data itself in all training rounds form its own corresponding historical prediction data set; when the first model is trained successfully to obtain the corresponding second model, it is determined whether the historical prediction data set of the unlabeled image data itself meets the set condition. The set condition is that the label predictions of the unlabeled image data are successively stable on two different class labels during the training process, that is, stable on one of the two class labels in the earlier training rounds during the training process, and stable on the other of the two class labels in the later training rounds during the training process; the unlabeled image data whose own historical prediction data set meets the set condition is determined as pseudo - label image data, and the class label on which the final label prediction is stable is determined as the pseudo - label of the pseudo - label image data; a large number of obtained pseudo - label image data and the target training set are used to form an enhanced training set, and the constructed initial model is trained with the enhanced training set to obtain a trained and qualified label prediction model. Thus, the present application adopts a three - stage training method (pre - training the model, determining the baseline pseudo - labels based on the pre - trained model, and determining the prediction - inconsistent pseudo - labels), so that the data used for training the model includes not only the image data with manually labeled true labels, but also the image data with baseline pseudo - labels and the image data with prediction - inconsistent pseudo - labels. The pseudo - label image data obtained by these two different pseudo - label determination methods have different characteristics (for example, the characteristics of the baseline pseudo - label image data are mainly reflected in the confidence level, and the characteristics of the new pseudo - label image data selected by the present application through prediction inconsistency are mainly reflected in the inconsistency of the predicted classes before and after), which can further enhance the training effect of the label prediction model on the basis of the traditional pseudo - label set.
[0091] Combined with the above embodiments, in one implementation manner, the embodiment of the present application also provides a method for enhancing training by using pseudo - labels of prediction - inconsistent samples. In this method for enhancing training by using pseudo - labels of prediction - inconsistent samples, step S5 may include steps S51 to S55:
[0092] Step S51: Train the constructed initial model with the enhanced training set. When a qualified label prediction model cannot be obtained after a preset number of training sessions, determine the initial model after the preset number of training sessions as the third model.
[0093] In this embodiment, the constructed initial model is trained with the enhanced training set, and the number of training sessions is specified. After the initial model is trained with the enhanced training set for a preset number of times, if it is still not qualified based on the corresponding loss function, in order to improve the training efficiency and performance of the model, this application will provide more pseudo-label image data to participate in the subsequent training process of the initial model after the preset number of training sessions.
[0094] Specifically, when the constructed initial model is trained with the enhanced training set and a qualified label prediction model cannot be obtained after a preset number of training sessions, determine the initial model after the preset number of training sessions as the third model.
[0095] Step S52: Train the third model with the target training set, and use the third model after each training round to predict the labels of new unlabeled image data to obtain the corresponding prediction distribution results. All the prediction distribution results obtained from the new unlabeled image data in all training rounds form the corresponding historical prediction data set.
[0096] In this embodiment, for the third model determined in step S51, train the third model with the target training set used to train the first model. At the same time, use the third model after each training round in the current training process to predict the labels of new unlabeled image data to obtain the corresponding prediction distribution results. All the prediction distribution results obtained from the new unlabeled image data in all training rounds form the historical prediction data set corresponding to the new unlabeled image data. At this time, the purpose of training the third model with the target training set is to obtain the historical prediction data set of the new unlabeled image data for determining the pseudo-labels of the new unlabeled image data, rather than training to obtain the final label prediction model.
[0097] Step S53: When the third model is trained qualified, determine whether the historical prediction data set of the new unlabeled image data meets the set conditions.
[0098] In this embodiment, when the third model is trained successfully, each new unlabeled image data will have a corresponding historical prediction dataset of its own. Then, when the third model is trained successfully, at this time, it is determined whether the historical prediction dataset of each new unlabeled image data meets the set conditions. The set condition is that the historical prediction dataset of the unlabeled image data shows that the label prediction results during the training process are successively stable on two category labels, that is, stable on one of the two category labels in the earlier training rounds during the training process, and stable on the other of the two category labels in the later training rounds during the training process.
[0099] Step S54: Determine the new unlabeled image data that meets the set conditions as pseudo-labeled image data, and determine the category label on which the final label prediction is stable as the pseudo-label of the pseudo-labeled image data.
[0100] In this embodiment, after determining whether the historical prediction dataset meets the set conditions through step S53, the new unlabeled image data corresponding to the historical prediction dataset that meets the set conditions is determined as pseudo-labeled image data. At the same time, the category label on which the label prediction shown in the historical prediction dataset of the pseudo-labeled image data is stable is determined as the pseudo-label of the pseudo-labeled image data, and the pseudo-labeled image data is labeled with this pseudo-label and used to participate in the subsequent training of the label prediction model.
[0101] Step S55: Use the obtained large number of new pseudo-labeled image data and the enhanced training set to form a new enhanced training set, and train the third model that has been trained a preset number of times through the new enhanced training set to obtain a successfully trained label prediction model.
[0102] In this embodiment, after obtaining a large number of new pseudo-labeled image data through step S54, the training set composed of the large number of new pseudo-labeled image data and the previous enhanced training set are combined to form a new enhanced training set, and the current trained third model is continuously trained through the new enhanced training set to obtain a finally successfully trained label prediction model.
[0103] Combined with the above embodiments, in one implementation manner, the embodiments of the present application further provide a pseudo-label enhanced training method using prediction-inconsistent samples. In this pseudo-label enhanced training method using prediction-inconsistent samples, the method further includes: testing the successfully trained label prediction model with the labeled image data in the test set to evaluate the performance of the label prediction model on unseen image data; and determining whether to perform new training on the label prediction model according to the performance evaluation result.
[0104] In this embodiment, after training to obtain a qualified label prediction model, the trained qualified label prediction model is tested with the labeled image data with true labels in the test set to evaluate the performance of the label prediction model on unseen image data. According to the performance evaluation results, it is determined whether a new training of the label prediction model is required. For example, in the case where the evaluation results do not meet the corresponding performance conditions, the label prediction model is newly trained. The performance conditions include whether the accuracy rate reaches the corresponding threshold set in advance, whether the precision rate reaches the corresponding threshold set in advance, whether the recall rate reaches the corresponding threshold set in advance, and whether the F1 value meets the corresponding requirements.
[0105] Combined with the above embodiments, in one implementation, the embodiments of the present application further provide a pseudo-label enhanced training method using prediction-inconsistent samples. In this pseudo-label enhanced training method using prediction-inconsistent samples, step S1 may include: performing label prediction on unlabeled image data through a first model; determining the predicted label category that meets the baseline pseudo-label selection method in the label prediction result as the pseudo-label of the unlabeled image data; constructing a baseline pseudo-label training set based on the determined unlabeled image data and their respective pseudo-labels.
[0106] In this embodiment, the baseline pseudo-label selection method is preferably to select the predicted label category with the highest confidence of the unlabeled image data as the pseudo-label of the unlabeled image data. First, label prediction is performed on the unlabeled image data through the first model to obtain the confidence of the unlabeled image data belonging to various label categories. Then, the predicted label category that meets the baseline pseudo-label selection method (i.e., the predicted label category with the highest confidence) in the label prediction result is determined as the pseudo-label of the unlabeled image data; the unlabeled image data and its own pseudo-label form the baseline pseudo-label image data, and a large number of baseline pseudo-label image data constitute the baseline pseudo-label training set.
[0107] Combined with the above embodiments, in one implementation, the embodiments of the present application further provide a pseudo-label enhanced training method using prediction-inconsistent samples. In this pseudo-label enhanced training method using prediction-inconsistent samples, before performing label prediction on the unlabeled image data through the first model, the method further includes: pre-training the constructed initial model with the labeled image data in the true label training set; determining the pre-trained qualified initial model as the first model.
[0108] In this embodiment, before determining the pseudo-labels of the unlabeled image data through the baseline pseudo-label selection method to construct the baseline pseudo-label training set, the initial model constructed is pre-trained with the labeled image data in the true label training set to obtain a pre-trained qualified initial model, and this pre-trained qualified initial model is determined as the first model, which is used to predict the label categories of the unlabeled image data in step S1.
[0109] Combined with the above embodiments, in one implementation manner, the embodiments of the present application further provide a pseudo-label enhanced training method using prediction-inconsistent samples. In this pseudo-label enhanced training method using prediction-inconsistent samples, constructing an initial model includes: based on the selected type of neural network model, constructing the basic structure of the model, and the basic structure includes: an input layer, a convolutional layer, a pooling layer, an activation function, and a fully connected layer; adding a softmax layer after the fully connected layer in the basic structure to convert the model output into a probability distribution; setting the loss function of the model to construct and obtain the final initial model.
[0110] In this embodiment, an appropriate type of neural network model is selected based on the requirements of the image classification task, such as CNN, ResNet, etc. Then, the basic structure of the corresponding type of neural network model is constructed, including the input layer, convolutional layer, and pooling layer of the model for extracting image features, activation function, fully connected layer, etc. For the classification task, it is necessary to convert the model output into a probability distribution. Therefore, a softmax layer is added after the fully connected layer in the basic structure of the neural network model to convert the model output into a probability distribution. Then, the loss function of the model is set to obtain the final initial model.
[0111] Combined with the above embodiments, in one implementation manner, the embodiments of the present application further provide a pseudo-label enhanced training method using prediction-inconsistent samples. In this pseudo-label enhanced training method using prediction-inconsistent samples, step S3 may include steps S31 to S36:
[0112] Step S31: When the first model is trained to be qualified to obtain the corresponding second model, calculate the historical prediction data set of the unlabeled image data through the prediction distribution mean algorithm to obtain the mean vector of the prediction distribution of the unlabeled image data;
[0113] The calculation expression of the prediction distribution mean algorithm is:
[0114]
[0115] Where, $\hat{p}_{i,t}$ is the predicted distribution result obtained for the $i$-th unlabeled image data in the $t$-th training round, and this predicted distribution result records the probabilities of the $i$-th unlabeled image data belonging to various class labels in the $t$-th training round; denotes the total number of training rounds; $\mu_{i,t}$ is the mean vector of the predicted distribution of the $i$-th unlabeled image data.
[0116] In this embodiment, a method for enhancing training using pseudo-labels of prediction-inconsistent samples provided by the present application mainly determines unlabeled image data whose corresponding historical prediction data set meets the set conditions as pseudo-label image data and determines the pseudo-labels of this pseudo-label image data. For example, Figure 2 the performance of the historical prediction data set of the first Node in Figure 2 meets the set conditions. And for Figure 3 the first to fourth Nodes in Figure 3 as shown in Figure 3 the first column shows the relationship between the training rounds of each of the four unlabeled image data and the predicted label classes obtained by prediction; the second column shows the mean values of the prediction probabilities of each of the four unlabeled image data under different predicted label classes. For unlabeled image data whose corresponding historical prediction data set meets the set conditions, it exhibits an obvious bimodal feature throughout the training process. That is, during the training process, the predicted distribution means of the unlabeled image data under two predicted label classes are significantly higher throughout the training process. For this performance, the present application designs a spatial feature to capture the unlabeled image data with this performance and the corresponding pseudo-labels, but this performance cannot obtain accurate unlabeled image data that meets the set conditions and the corresponding pseudo-labels. For example, Figure 3 both the first Node and the fourth Node in Figure 3the unlabeled image data of the fourth Node in []. Thus, for the screening of unlabeled image data that meet the set conditions in the corresponding historical prediction dataset, the present application provides two features for screening, namely spatial features and temporal features . These two features are calculated through the historical prediction dataset of unlabeled image data. For each unlabeled image data, pseudo-labels with inconsistent prediction pattern training dynamics are screened from both spatial and temporal aspects, and the unlabeled image data to which the pseudo-labels belong participate in subsequent model training. The specific calculation and screening process is to calculate the historical prediction dataset of each unlabeled image data respectively to obtain the values of each unlabeled image data on these two features, and then determine which unlabeled image data meet the corresponding conditions on these two features, so as to determine the unlabeled image data that meet the corresponding conditions on these two features as pseudo-label image data. Among them, Figure 2 and Figure 3 the first Node in [] is Node876 in the figure, the second Node is Node121 in the figure, the third Node is Node1016 in the figure, and the fourth Node is Node2359 in the figure.
[0117] Specifically, since the calculation of the spatial feature and the temporal feature of each unlabeled image data is the same, an example of an unlabeled image data is used for illustration here. In the case where the first model is trained successfully to obtain the corresponding second model, the historical prediction dataset of the unlabeled image data is calculated by the prediction distribution mean algorithm to obtain the mean vector of the prediction distribution of the unlabeled image data. The calculation expression of the prediction distribution mean algorithm is:
[0118]
[0119] Among them, represents the prediction distribution result obtained by the i-th unlabeled image data in the t-th training round. The prediction distribution result records the probabilities of the i-th unlabeled image data belonging to various category labels in the t-th training round. For example, when the category labels include category 1, category 2, and category 3, the prediction distribution result of the i-th unlabeled image data obtained in each training round will record the probability of the i-th unlabeled image data belonging to category 1, the probability of belonging to category 2, and the probability of belonging to category 3 in the corresponding training round; represents the total number of training rounds. For example, when the first model obtains a qualified second model after training for 100 training rounds, the total number of training rounds is 100; is the mean vector of the prediction distribution of the i-th unlabeled image data. The mean vector of the prediction distribution records the mean of the prediction probabilities of the unlabeled image data under each class label. For example, when the class labels include Class 1, Class 2, and Class 3, then records the mean of the prediction probabilities of the i-th unlabeled image data under Class 1, the mean of the prediction probabilities of the i-th unlabeled image data under Class 2, and the mean of the prediction probabilities of the i-th unlabeled image data under Class 3.
[0120] Step S32: Calculate the obtained mean vector of the prediction distribution through the target entropy algorithm to obtain the spatial features of the unlabeled image data;
[0121] The calculation expression of the target entropy algorithm is:
[0122]
[0123] where is all other class labels except the class label corresponding to the maximum value in ; is the mean of the prediction probability of the i-th unlabeled image data under the class label c; is the spatial feature of the i-th unlabeled image data; is the entropy after removing the maximum value.
[0124] In this embodiment, after obtaining the mean vector of the prediction distribution of the unlabeled image data through Step S31, calculate the mean vector of the prediction distribution of the unlabeled image data through the target entropy algorithm to obtain the value of the spatial feature of the unlabeled image data. The calculation expression of the target entropy algorithm is:
[0125]
[0126] where is all other class labels except the class label corresponding to the maximum value in the mean vector of the prediction distribution of the i-th unlabeled image data . For example, when the class labels include Class 1, Class 2, and Class 3, and at the same time the mean of the prediction probability of the i-th unlabeled image data under Class 1 in is the largest, then the includes the class labels Class 2 and Class 3; is the mean of the prediction probability of the i-th unlabeled image data under the class label c; To remove the entropy of the maximum value.
[0127] Step S33: Calculate the historical prediction dataset of the unlabeled image data through the change trend algorithm to obtain the change trend vector of the prediction distribution of the unlabeled image data;
[0128] The calculation expression of the change trend algorithm is:
[0129]
[0130] Where and are consecutive training rounds of the unlabeled image data, is the next training round; and respectively represent the prediction distribution results of the i-th unlabeled image data in the training rounds and respectively, represents the change trend vector of the prediction distribution of the i-th unlabeled image data.
[0131] In this embodiment, when the first model is trained qualified to obtain the corresponding second model, at the same time, calculate the historical prediction dataset of the unlabeled image data through the change trend algorithm to obtain the change trend vector of the prediction distribution of the unlabeled image data. The calculation expression of the change trend algorithm is:
[0132]
[0133] Where and are consecutive training rounds of the unlabeled image data, is the next training round; and respectively represent the prediction distribution results of the i-th unlabeled image data in the training rounds and respectively, represents the change trend vector of the prediction distribution of the i-th unlabeled image data. The values recorded in the change trend vector of the prediction distribution are the change trend values of the prediction distribution of the unlabeled image data under each class label. For example, when the class labels include class 1, class 2, and class 3, then records the change trend value of the prediction distribution of the i-th unlabeled image data under class 1, the change trend value of the prediction distribution of the i-th unlabeled image data under class 2, and the change trend value of the prediction distribution of the i-th unlabeled image data under class 3.
[0134] Step S34: Calculate the obtained change trend vector through the first algorithm to obtain the first operator value and the second operator value of the unlabeled image data;
[0135] The calculation expression of the first algorithm is:
[0136]
[0137] Where, is the first operator of the i-th unlabeled image data; is the second operator of the i-th unlabeled image data; is the final predicted category of the unlabeled image data during training, and this predicted category is defined as the predicted category in the second stage; is except Outside, the predicted category corresponding to the maximum change trend in the change trend vector of the predicted distribution of the unlabeled image data is defined as the predicted category in the first stage; is the change trend of the predicted probability of the i-th unlabeled image data on the predicted category in the first stage; is the change trend of the predicted probability of the i-th unlabeled image data on the predicted category in the second stage; is the change trend of the predicted distribution of the i-th unlabeled image data on other predicted categories except the predicted categories in the first stage and the second stage.
[0138] In this embodiment, after calculating the change trend vector of the predicted distribution of the unlabeled image data through step S33, calculate the obtained change trend vector of the predicted distribution of the unlabeled image data through the first algorithm to obtain the first operator value and the second operator value of the unlabeled image data. The calculation expression of the first algorithm is:
[0139]
[0140] Where, is the first operator of the i-th unlabeled image data, used to capture the changes on the two categories of the bimodal; is the second operator of the i-th unlabeled image data, used to capture the changes on the non-bimodal categories; is the final predicted category of the unlabeled image data during training, and this predicted category is defined as the predicted category in the second stage; is except Outside, the predicted category corresponding to the maximum change trend in the change trend vector of the predicted distribution of the unlabeled image data is defined as the predicted category in the first stage; is the value of the change trend of the predicted probability of the i-th unlabeled image data on the predicted category in the first stage It is the value of the change trend of the prediction probability on the prediction category of the i-th unlabeled image data in the second stage; It is the value of the change trend of the prediction distribution on the prediction categories other than the prediction categories in the first stage and the second stage of the i-th unlabeled image data.
[0141] In this embodiment, there is a special case where the prediction category label in the second stage of an unlabeled image data (such as unlabeled image data A) is not only the final prediction category of the unlabeled image data A during the training process, but also the prediction category corresponding to the largest change trend in the change trend vector of the prediction distribution of the unlabeled image data A. It will no longer be the prediction category corresponding to the largest change trend in the change trend vector of the prediction distribution of the unlabeled image data A, but will use the prediction category corresponding to the second largest change trend in the change trend vector of the prediction distribution of the unlabeled image data A as the , because the prediction category corresponding to the largest trend has been occupied by the prediction category in the second stage of the unlabeled image data A.
[0142] Step S35: Calculate the first operator value and the second operator value through a second algorithm to obtain the time feature of the unlabeled image data;
[0143] The calculation expression of the second algorithm is:
[0144]
[0145]
[0146]
[0147]
[0148]
[0149] where and are thresholds; is the change index of the i-th unlabeled image data; is the direction index of the i-th unlabeled image data; T is the last training round in the training process; is the prediction probability of the i-th unlabeled image data on the prediction category in the second stage in the t-th training round; is the prediction probability of the i-th unlabeled image data on the prediction category in the first stage in the t-th training round; is the time feature of the i-th unlabeled image data; and is the weight.
[0150] In this embodiment, after obtaining the first operator value and the second operator value of the unlabeled image data through step S34, the second algorithm is used to calculate the first operator value and the second operator value of the unlabeled image data to obtain the time feature of the unlabeled image data. The calculation expression of the second algorithm is a system of equations, as follows:
[0151]
[0152]
[0153]
[0154]
[0155]
[0156] where and are two preset thresholds; is the change index of the i-th unlabeled image data; is the direction index of the i-th unlabeled image data; T is the last training round in the training process; is the prediction probability of the i-th unlabeled image data on the predicted class in the second stage in the t-th training round; is the prediction probability of the i-th unlabeled image data on the predicted class in the first stage in the t-th training round; is the time feature of the i-th unlabeled image data; and are weights. Through the same implementation manner, each unlabeled image data can calculate its corresponding time feature based on its own historical prediction data set.
[0157] In this embodiment, for the sample data with the characteristic of inconsistent prediction, since the predicted class in the first stage is , and the final predicted class is , then at some moments during the training process, the probability of is greater than , is negative, and at the last moment, the difference between the two can reflect this change in sign, that is, the larger is, the more likely the corresponding sample data is the sample data with the characteristic of inconsistent prediction. To calculate , this application takes the change difference of the bimodal class at the last step During the training process the minimum value Take the difference to obtain an indicator reflecting whether there is a change in category .
[0158] Step S36: Determine whether the predicted distribution result of the unlabeled image data meets the set conditions according to the spatial features and temporal features of the unlabeled image data.
[0159] In this embodiment, based on the obtained values of the spatial features and temporal features of the unlabeled image data, determine whether the historical prediction data set of the unlabeled image data meets the set conditions to determine whether it can be determined as pseudo-labeled image data.
[0160] Combined with the above embodiments, in one implementation manner, the embodiment of the present application further provides a pseudo-label enhanced training method using prediction inconsistent samples. In this pseudo-label enhanced training method using prediction inconsistent samples, step S36 may include: calculating the spatial features and temporal features of the unlabeled image data through a prediction inconsistent index algorithm to obtain the prediction inconsistent index value of the unlabeled image data;
[0161] The calculation expression of the prediction inconsistent index algorithm is:
[0162]
[0163] where and are weights;
[0164] Determine whether the predicted distribution result of the unlabeled image data meets the set conditions according to the relationship between the prediction inconsistent index value and the set threshold.
[0165] In this embodiment, the present application defines a prediction inconsistency metric for unlabeled image data. This prediction inconsistency metric is calculated based on the spatial and temporal features of the unlabeled image data. Whether the value of the prediction inconsistency metric based on the unlabeled image data is lower than a set threshold is used to determine whether the unlabeled image data meets the set conditions. Among them, the set threshold is preferably the mean value of the prediction inconsistency metric on the unlabeled image data set. This is only a preferred value, and it can also be other preset values. Specifically, after calculating the spatial and temporal features of the unlabeled image data through steps S31 to S36, the spatial and temporal features are brought into the prediction inconsistency metric algorithm for calculation to obtain the value of the prediction inconsistency metric of the unlabeled image data. Then, the value of the prediction inconsistency metric is compared with the set threshold. When the value of the prediction inconsistency metric does not exceed the set threshold, it is determined that the unlabeled image data meets the set conditions. At this time, the unlabeled image data (such as image data a) is determined as pseudo-labeled image data (i.e., image data a), and at the same time, the prediction category in the second stage of the unlabeled image data (i.e., image data a) is determined as the pseudo-label of the pseudo-labeled image data (i.e., image data a).
[0166] In this embodiment, the prediction inconsistency metric proposed by a pseudo-label enhancement training method using prediction inconsistent samples provided by the present application describes the characteristics of training dynamics with a prediction inconsistent pattern in space and time. Through this prediction inconsistency metric, this novel type of pseudo-label that conforms to the prediction inconsistent dynamics can be screened out. This is a new type of prediction label with a prediction inconsistent training dynamics pattern. The pseudo-label and the corresponding unlabeled image data can help the model learn the correct correlations while eliminating the incorrect correlations, complementing the deficiency of screening pseudo-labels by confidence scores.
[0167] It should be noted that for the method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequences, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present application.
[0168] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0169] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0170] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.
[0171] The above has introduced in detail a method for enhancing training using pseudo-labels of prediction-inconsistent samples provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A pseudo-label enhancement training method using prediction inconsistent samples, characterized in that, The method includes: Determining pseudo-labels of unlabeled image data through a baseline pseudo-label selection method to construct a baseline pseudo-label training set; Training a first model with a target training set composed of a true-label training set and the baseline pseudo-label training set, and using the first model after each training round to perform label prediction on the unlabeled image data to obtain corresponding prediction distribution results. All the prediction distribution results obtained by the unlabeled image data in all training rounds form a corresponding historical prediction data set; When the first model is trained successfully to obtain a corresponding second model, determining whether the historical prediction data set of the unlabeled image data meets a set condition, where the set condition is that the label predictions of the unlabeled image data during training are successively stable on two different category labels; Determining the unlabeled image data that meets the set condition as pseudo-label image data, and determining the category label on which the final label prediction is stable as the pseudo-label of the pseudo-label image data; Using the obtained large number of pseudo-label image data and the target training set to form an enhanced training set, and training an initially constructed model with the enhanced training set to obtain a trained and qualified label prediction model; Among them, training an initially constructed model with the enhanced training set to obtain a trained and qualified label prediction model includes: training the initially constructed model with the enhanced training set. When a trained and qualified label prediction model cannot be obtained after a preset number of trainings, determining the initially constructed model after the preset number of trainings as a third model; training the third model with the target training set, and using the third model after each training round to perform label prediction on new unlabeled image data to obtain corresponding prediction distribution results. All the prediction distribution results obtained by the new unlabeled image data in all training rounds form a corresponding historical prediction data set; when the third model is trained successfully, determining whether the historical prediction data set of the new unlabeled image data meets the set condition; determining the new unlabeled image data that meets the set condition as pseudo-label image data, and determining the category label on which the final label prediction is stable as the pseudo-label of the pseudo-label image data; using the obtained large number of new pseudo-label image data and the enhanced training set to form a new enhanced training set, and training the third model after the preset number of trainings with the new enhanced training set to obtain a trained and qualified label prediction model.
2. The pseudo-label enhancement training method using prediction inconsistent samples according to claim 1, wherein The method further includes: Testing the trained and qualified label prediction model with the labeled image data in the test set to evaluate the performance of the label prediction model on unseen image data; Determining whether to perform new training on the label prediction model according to the performance evaluation result.
3. A method for enhancing training using pseudo-labels of prediction inconsistent samples according to claim 1, characterized in that Determining pseudo-labels of unlabeled image data through a baseline pseudo-label selection method to construct a baseline pseudo-label training set includes: Performing label prediction on the unlabeled image data through a first model; Determining the predicted label category that meets the baseline pseudo-label selection method in the label prediction result as the pseudo-label of the unlabeled image data; Construct a baseline pseudo-label training set based on the determined unlabeled image data and their respective pseudo-labels.
4. A method for enhancing training using pseudo-labels of prediction inconsistent samples according to claim 3, characterized in that Before predicting the labels of the unlabeled image data through the first model, the method further includes: Pre-train the constructed initial model with the labeled image data in the true label training set; Determine that the initial model with qualified pre-training is the first model.
5. A method for enhancing training using pseudo-labels of prediction inconsistent samples according to claim 1, characterized in that Construct the initial model, including: Based on the selected type of neural network model, construct the basic structure of the model, and the basic structure includes: an input layer, a convolutional layer, a pooling layer, an activation function, and a fully connected layer; Add a softmax layer after the fully connected layer in the basic structure to convert the model output into a probability distribution; Set the loss function of the model to construct and obtain the final initial model.
6. A method for enhancing training using pseudo - labels of prediction - inconsistent samples according to claim 1, characterized in that, When the first model is trained qualified to obtain the corresponding second model, determine whether the historical prediction data set of the unlabeled image data meets the set conditions, including: When the first model is trained qualified to obtain the corresponding second model, calculate the historical prediction data set of the unlabeled image data through the prediction distribution mean algorithm to obtain the mean vector of the prediction distribution of the unlabeled image data; The calculation expression of the prediction distribution mean algorithm is: Among them, is the predicted distribution result obtained by the i-th unlabeled image data in the t-th training round, and the predicted distribution result records the probabilities of the i-th unlabeled image data belonging to various class labels in the t-th training round; represents the total number of training rounds; is the mean vector of the predicted distribution of the i-th unlabeled image data; Calculate the obtained mean vector of the prediction distribution through the target entropy algorithm to obtain the spatial features of the unlabeled image data; The calculation expression of the target entropy algorithm is: wherein is all other class labels except the class label corresponding to the maximum value in; is the mean of the predicted probabilities of the i-th unlabeled image data under the class label c; is the spatial feature of the i-th unlabeled image data; is the entropy after removing the maximum value; Calculate the historical prediction data set of the unlabeled image data through the change trend algorithm to obtain the change trend vector of the prediction distribution of the unlabeled image data; The calculation expression of the change trend algorithm is: Among them, and are consecutive training rounds of unlabeled image data, is the next training round; and respectively represent the prediction distribution results of the i-th unlabeled image data at training rounds and respectively, represents the change trend vector of the prediction distribution of the i-th unlabeled image data; Calculate the obtained change trend vector through the first algorithm to obtain the first operator value and the second operator value of the unlabeled image data; The calculation expression of the first algorithm is: Among them, is the first operator of the i-th unlabeled image data; is the second operator of the i-th unlabeled image data; is the final predicted category of the unlabeled image data during the training process, and this predicted category is defined as the predicted category in the second stage; is for excluding Among them, the predicted category corresponding to the largest change trend in the change trend vector of the predicted distribution of the unlabeled image data is defined as the predicted category in the first stage; is the change trend of the predicted probability of the i-th unlabeled image data on the predicted category in the first stage; is the change trend of the predicted probability of the i-th unlabeled image data on the predicted category in the second stage; is the change trend of the predicted distribution of the i-th unlabeled image data on other predicted categories except the predicted categories in the first stage and the second stage; Calculate the first operator value and the second operator value through the second algorithm to obtain the time features of the unlabeled image data; The calculation expression of the second algorithm is: where and are thresholds; is the change index of the i-th unlabeled image data; is the direction index of the i-th unlabeled image data; T is the last training epoch during the training process; is the prediction probability of the i-th unlabeled image data on the predicted class in the second stage in the t-th training epoch; is the prediction probability of the i-th unlabeled image data on the predicted class in the first stage in the t-th training epoch; is the time feature of the i-th unlabeled image data; and are weights; Determine whether the prediction distribution result of the unlabeled image data meets the set conditions according to the spatial features and time features of the unlabeled image data.
7. A method for enhancing training using pseudo-labels of predicted inconsistent samples according to claim 6, characterized in that Determine whether the prediction distribution result of the unlabeled image data meets the set conditions according to the spatial features and time features of the unlabeled image data, including: Calculate the spatial features and time features of the unlabeled image data through the prediction inconsistency index algorithm to obtain the prediction inconsistency index value of the unlabeled image data; The calculation expression of the prediction inconsistency index algorithm is: Among them, and are weights; Determine whether the prediction distribution result of the unlabeled image data meets the set conditions according to the relationship between the prediction inconsistency index value and the set threshold.
Citation Information
Patent Citations
Semi-supervised model training method, image recognition method and device
CN116935146A
Lotus phenotype identification method and device based on pseudo tag algorithm and MobileNetV2 network
CN117953281A