A data processing method, apparatus and device
By using the pseudo-label filtering of the unlabeled data set and multiple basic models, selecting suitable unlabeled data based on uncertainty for training, the problems of human resources and time consumption in machine learning model training are solved, and the model performance and efficiency are improved.
Patent Information
- Application Number
- CN202111523107.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-12-13
AI Technical Summary
In the prior art, the training of machine learning models requires a large amount of labeled data, which consumes a lot of human resources and time, and the labeling process consumes time and effort, affecting the user experience.
By obtaining the label-free data set and using the pseudo-labels output from multiple basic models, appropriate label-free data are selected based on uncertainty for training, the target model is obtained, reducing calibration operations for label-free data, and improving model performance and efficiency.
Reduce calibration operations on labelless data, save human resources, improve model performance and efficiency, reduce noise impact in pseudo-labels, and achieve robust and efficient semi-supervised learning.
Smart Images

Figure CN114298173B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a data processing method, apparatus, and device. Background Art
[0002] Machine learning is a way to achieve artificial intelligence. It is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. Machine learning is used to study how a computer simulates or implements human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning pays more attention to algorithm design, enabling a computer to automatically learn patterns from data and use the patterns to predict unknown data. Machine learning has been widely applied, such as in deep learning, data mining, computer vision, natural language processing, biometric recognition, search engines, medical diagnosis, speech recognition, and handwriting recognition.
[0003] In order to implement artificial intelligence processing using machine learning, a training data set can be constructed. The training data set includes a large number of labeled data (such as image data, that is, images with calibration boxes and calibration categories). Based on the training data set, a machine learning model can be trained, such as a machine learning model with object detection capabilities. The machine learning model can be used to perform object detection on the data to be detected. For example, the object box in the data to be detected can be detected, and the object category can be identified, such as vehicle category, animal category, electronic product category, etc.
[0004] In order to improve the performance of the machine learning model, a large number of labeled data need to be obtained. The more labeled data there are, the better the performance of the trained machine learning model. However, in order to obtain a large number of labeled data, a large amount of data needs to be labeled, which consumes a large amount of human resources and time. Summary of the Invention
[0005] This application provides a data processing method, and the method includes:
[0006] Obtain an unlabeled data set, where the unlabeled data set includes a plurality of unlabeled data; for each unlabeled data, the unlabeled data corresponds to a plurality of pseudo-labels, and the plurality of pseudo-labels are the pseudo-labels output by a plurality of basic models after the unlabeled data is input to the plurality of basic models;
[0007] For each base model, select the target unlabeled data corresponding to the base model from the unlabeled data set; wherein, for each unlabeled data in the unlabeled data set, based on the multiple pseudo-labels corresponding to the unlabeled data, determine the first uncertainty of the unlabeled data with respect to the base model and the second uncertainty of the unlabeled data with respect to the remaining base models other than the base model; based on the first uncertainty and the second uncertainty, determine whether the unlabeled data is the target unlabeled data corresponding to the base model or not the target unlabeled data corresponding to the base model.
[0008] Train the base model based on the target unlabeled data corresponding to the base model to obtain a trained target model; wherein, the target model is used to process application data.
[0009] This application provides a data processing device, the device includes:
[0010] An acquisition module, configured to acquire an unlabeled data set, the unlabeled data set includes a plurality of unlabeled data; for each unlabeled data, the unlabeled data corresponds to a plurality of pseudo-labels, and the plurality of pseudo-labels are the pseudo-labels output by the plurality of base models after inputting the unlabeled data to the plurality of base models.
[0011] A determination module, configured to, for each base model, select the target unlabeled data corresponding to the base model from the unlabeled data set; wherein, for each unlabeled data in the unlabeled data set, based on the multiple pseudo-labels corresponding to the unlabeled data, determine the first uncertainty of the unlabeled data with respect to the base model and the second uncertainty of the unlabeled data with respect to the remaining base models other than the base model; based on the first uncertainty and the second uncertainty, determine whether the unlabeled data is the target unlabeled data corresponding to the base model or not the target unlabeled data corresponding to the base model.
[0012] A training module, configured to train the base model based on the target unlabeled data corresponding to the base model to obtain a trained target model; the target model is used to process application data.
[0013] This application provides a data processing device, including: a processor and a machine-readable storage medium, the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is configured to execute the machine-executable instructions to implement the data processing method disclosed in the above examples of this application.
[0014] As can be seen from the above technical solutions, in the embodiments of the present application, the basic model can be trained based on unlabeled data and the pseudo-labels corresponding to the unlabeled data to obtain the target model. The pseudo-labels are the labels output by the basic model, rather than the labels manually marked by users. Thus, it is possible to avoid calibrating a large amount of unlabeled data, reduce the calibration operations of a large amount of data, save human resources, and reduce the calibration time. For each basic model, the target unlabeled data can be selected for the basic model based on the uncertainty corresponding to each unlabeled data in the unlabeled data set, so as to select appropriate and valuable unlabeled data for the basic model to participate in the training, enabling the performance of the basic model to be improved quickly and robustly. The relative measure of uncertainty is used for pseudo-label screening, and the pseudo-labels adapted to the basic model are specifically selected to participate in the training, reducing the probability that the basic model receives noise in the pseudo-labels, that is, the influence of the noise in the pseudo-labels is reduced, and the quality of the pseudo-labels is improved. Multiple basic models are used to provide pseudo-labels and cooperate in the learning and optimization process, enabling the knowledge of multiple basic models to flow and be shared efficiently. It is not strongly coupled with a single task and can be adapted to tasks such as detection, classification, and segmentation, ensuring generality. Under the semi-supervised learning of pseudo-labels, a more robust and efficient learning mode is adopted to obtain a model with better performance. The pseudo-labels are jointly provided and optimized by multiple basic models, and the influence of noise in the pseudo-labels of a single basic model is reduced. Overall, knowledge sharing is completed in a collaborative learning mode, reducing the number of unlabeled data during the training of a single model, and thus being beneficial to improving the performance and efficiency of semi-supervised learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings of the embodiments of the present application.
[0016] Figure 1 is a flowchart of a data processing method in an embodiment of the present application;
[0017] Figure 2 is a schematic structural diagram of a system in an embodiment of the present application;
[0018] Figure 3 is a flowchart of a data processing method in an embodiment of the present application;
[0019] Figure 4 is a schematic structural diagram of a data processing device in an embodiment of the present application;
[0020] Figure 5 is a hardware structure diagram of a data processing device in an embodiment of the present application. Detailed implementation manners
[0021] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and do not limit the present application. The singular forms "a", "the", and "said" used in the present application and the claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to any or all possible combinations including one or more of the associated listed items.
[0022] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, in addition, the word "if" used may be interpreted as "when" or "while" or "in response to determining".
[0023] In the embodiments of the present application, a data processing method is proposed, which can be applied to a data processing device. Refer to Figure 1 As shown, it is a schematic flowchart of the data processing method. The method may include:
[0024] Step 101, obtain an unlabeled data set, where the unlabeled data set may include multiple unlabeled data; for each unlabeled data, the unlabeled data corresponds to multiple pseudo-labels, and the multiple pseudo-labels are the pseudo-labels output by multiple basic models after the unlabeled data is input to the multiple basic models.
[0025] Exemplarily, for each unlabeled data in the unlabeled data set, the unlabeled data may be augmented A times to obtain A unlabeled data after data augmentation, where A may be a positive integer. For each unlabeled data after data augmentation, the unlabeled data after data augmentation may be input to multiple basic models, and the multiple basic models output the pseudo-labels corresponding to the unlabeled data, so that multiple pseudo-labels corresponding to the unlabeled data, that is, the pseudo-labels output by the multiple basic models, can be obtained.
[0026] Step 102: For each base model, select the target unlabeled data corresponding to the base model from the unlabeled dataset. Exemplarily, for each unlabeled data in the unlabeled dataset, based on the multiple pseudo-labels corresponding to the unlabeled data, determine the first uncertainty of the unlabeled data with respect to the base model and the second uncertainty of the unlabeled data with respect to the remaining base models other than the base model; based on the first uncertainty and the second uncertainty, determine whether the unlabeled data is the target unlabeled data corresponding to the base model or determine that the unlabeled data is not the target unlabeled data corresponding to the base model.
[0027] Exemplarily, based on the multiple pseudo-labels corresponding to the unlabeled data, determining the first uncertainty of the unlabeled data with respect to the base model and the second uncertainty of the unlabeled data with respect to the remaining base models other than the base model may include, but is not limited to: dividing the multiple pseudo-labels corresponding to the unlabeled data into a first pseudo-label set and a second pseudo-label set, where the pseudo-labels in the first pseudo-label set are the pseudo-labels output by the base model, and the pseudo-labels in the second pseudo-label set are the pseudo-labels output by the remaining base models other than the base model. Determine the first uncertainty based on the confidence levels corresponding to the pseudo-labels in the first pseudo-label set; determine the second uncertainty based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set.
[0028] Exemplarily, determining the first uncertainty based on the confidence levels corresponding to the pseudo-labels in the first pseudo-label set may include, but is not limited to: determining the entropy of the first average value based on the confidence levels corresponding to the pseudo-labels in the first pseudo-label set, and determining the average value of the first entropy based on the confidence levels corresponding to the pseudo-labels in the first pseudo-label set; determining the first uncertainty based on the entropy of the first average value and the average value of the first entropy.
[0029] Exemplarily, determining the second uncertainty based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set may include, but is not limited to: determining the entropy of the second average value based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set, and determining the average value of the second entropy based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set; determining the second uncertainty based on the entropy of the second average value and the average value of the second entropy.
[0030] Exemplarily, determining that the unlabeled data is the target unlabeled data corresponding to the base model or determining that the unlabeled data is not the target unlabeled data corresponding to the base model based on the first uncertainty and the second uncertainty may include, but is not limited to: determining the uncertainty difference of the unlabeled data with respect to the base model and the remaining base models based on the difference between the first uncertainty and the second uncertainty; if the uncertainty difference is greater than a first threshold (which can be configured according to experience), it may be determined that the unlabeled data is the target unlabeled data corresponding to the base model; or, if the uncertainty difference is not greater than the first threshold, it may be determined that the unlabeled data is not the target unlabeled data corresponding to the base model.
[0031] Exemplarily, determining that the unlabeled data is the target unlabeled data corresponding to the base model or determining that the unlabeled data is not the target unlabeled data corresponding to the base model based on the first uncertainty and the second uncertainty may include, but is not limited to: determining the uncertainty difference of the unlabeled data with respect to the base model and the remaining base models based on the difference between the first uncertainty and the second uncertainty; determining the average confidence based on the confidences corresponding to the pseudo-labels in the second pseudo-label set. On this basis, if the uncertainty difference is greater than a first threshold (which can be configured according to experience) and the average confidence is greater than a second threshold (configured according to experience), it is determined that the unlabeled data is the target unlabeled data corresponding to the base model; if the uncertainty difference is not greater than the first threshold, and / or, the average confidence is not greater than the second threshold, it is determined that the unlabeled data is not the target unlabeled data corresponding to the base model.
[0032] Step 103: Train the base model based on the target unlabeled data corresponding to the base model to obtain a trained target model; where the target model is used to process application data. That is to say, the target model can be deployed online, and the application data can be input into the target model, and the target model processes the application data to obtain a data processing result, that is, an artificial intelligence processing result.
[0033] Exemplarily, training the base model based on the target unlabeled data corresponding to the base model to obtain a trained target model may include, but is not limited to: generating a target pseudo-label corresponding to the target unlabeled data based on multiple pseudo-labels corresponding to the target unlabeled data corresponding to the base model; training the base model based on the target unlabeled data and the target pseudo-label to obtain a trained target model.
[0034] As can be seen from the above technical solutions, in the embodiments of the present application, the basic model can be trained based on unlabeled data and the pseudo-labels corresponding to the unlabeled data to obtain a target model. The pseudo-labels are the labels output by the basic model, rather than the labels manually marked by users. Therefore, it is possible to avoid labeling a large amount of unlabeled data, reduce the labeling operations of a large amount of data, save human resources, and reduce the labeling time. For each basic model, the target unlabeled data can be selected for the basic model based on the uncertainty corresponding to each unlabeled data in the unlabeled data set, so as to select appropriate and valuable unlabeled data for the basic model to participate in the training, and enable the performance of the basic model to be improved quickly and robustly. The relative measure of uncertainty is used for pseudo-label screening, and the pseudo-labels suitable for the basic model are specifically selected to participate in the training, reducing the probability of noise in the pseudo-labels received by the basic model, that is, the influence of noise in the pseudo-labels is reduced, and the quality of the pseudo-labels is improved. Multiple basic models are used to provide pseudo-labels and cooperate in the learning and optimization process, enabling the knowledge of multiple basic models to flow and be shared efficiently. It is not strongly coupled with a single task and can be adapted to tasks such as detection, classification, and segmentation, ensuring generality. Under the semi-supervised learning of pseudo-labels, a more robust and efficient learning mode is used to obtain a model with better performance. The pseudo-labels are jointly provided and optimized by multiple basic models. The influence of noise in the pseudo-labels of a single basic model is reduced, and overall, knowledge sharing is completed in a collaborative learning mode, reducing the number of unlabeled data in the training of a single model, which is conducive to improving the performance and efficiency of semi-supervised learning.
[0035] The technical solutions of the embodiments of the present application will be described below in conjunction with specific application scenarios.
[0036] Machine learning models (such as deep learning models and neural network models, etc.) have made great progress in fields such as computer vision and natural language processing. Various types of machine learning models emerge in an endless stream. However, the performance of machine learning models requires the support of a large amount of labeled data. The training process needs to perform gradient backpropagation on various labeled data to update parameters, and the generalization depends on the richness and extensiveness of the labeled data. The above process is called full-supervised learning or training. However, labeling a large amount of data requires consuming a large amount of human resources and a large amount of time. Fine-grained annotation of a large amount of data (such as target boxes in detection tasks, pixel points in segmentation tasks, etc.) is very time-consuming and laborious, and the user experience is relatively poor.
[0037] In real scenarios, unlabeled data (i.e., image data) is extremely easy to obtain. Therefore, how to make full use of a large amount of unlabeled data to train machine learning models has become a research hotspot. Among them, semi-supervised learning is a learning method that studies how to use a large amount of unlabeled data to improve the performance of machine learning models, and can use the information in the unlabeled data to train machine learning models to increase the generalization of machine learning models.
[0038] In an embodiment of the present application, a robust semi-supervised learning method is proposed, which can perform semi-supervised learning using pseudo-labels, utilize multi-model collaborative training, where multiple models jointly provide pseudo-labels, and screen the pseudo-labels through relative measurement of uncertainty, making it difficult for noise to enter the pseudo-labels. Thus, under semi-supervised learning with pseudo-labels, a more robust and efficient learning mode is adopted to obtain a machine learning model with better performance.
[0039] Exemplarily, semi-supervised learning refers to a learning method that can train a machine learning model using a certain amount of labeled data and a large amount of unlabeled data to improve the model performance. Pseudo-labels refer to the labels assigned to unlabeled data using non-artificial annotation methods, which are an approximation of the true labels and participate in the model training as the true labels of unlabeled data in semi-supervised learning. The quality of pseudo-labels determines the training effect of the model. Robustness means that pseudo-labels are jointly provided and optimized by multiple models, reducing the influence of noise in the pseudo-labels of a single model. For a machine learning model, being robust means that the model is less affected by noise. Efficiency means that for each model, the most suitable and valuable unlabeled data and the corresponding pseudo-labels are specifically selected for training to rapidly improve the model performance.
[0040] See Figure 2 As shown, it is a schematic diagram of the system structure of an embodiment of the present application. The system structure may include multiple base models (clients) and a pseudo-label management module (expert). In Figure 2 it, M base models are taken as an example, where M is a positive integer greater than 1. The M base models interact around the pseudo-label management module. The base models are used to provide pseudo-labels to the pseudo-label management module, and the pseudo-label management module is used to provide the most valuable unlabeled subset and optimized pseudo-labels to the base models, thereby performing learning optimization on the base models based on the unlabeled subset and optimized pseudo-labels. Information flow is achieved through the pseudo-labels of multiple base models, enabling each base model to obtain valuable pseudo-labels for itself by combining the knowledge of other base models, and at the same time participating in the process of providing appropriate pseudo-labels to other base models.
[0041] See Figure 2 As shown, the M base models perform pseudo-label interaction around the pseudo-label management module. For each base model, the base model can obtain pseudo-labels through forward prediction of the base model based on the data-augmented unlabeled data, provide the pseudo-labels to the pseudo-label management module, and obtain the unlabeled subset and optimized pseudo-labels suitable for itself from the pseudo-label management module for training.
[0042] The pseudo-label management module may include a pseudo-label pool (PLPool) sub-module and a selector (Selector) sub-module. The pseudo-label pool sub-module is used to integrate the pseudo-labels from each base model, and the selector sub-module is used to screen out the pseudo-labels suitable for the base model based on the relative measure of the uncertainty of the pseudo-labels.
[0043] Pseudo-label pool sub-module: The pseudo-label pool includes all unlabeled data and the pseudo-labels predicted under each base model. The pseudo-label pool sub-module is used to obtain the pseudo-labels from each base model and record the unlabeled data and pseudo-labels in the pseudo-label pool. For example, an unlabeled data (such as an unlabeled image) is augmented independently A times and then predicted by M base models, resulting in M*A pseudo-labels under different test conditions. The pseudo-label pool sub-module then collates these pseudo-labels and records the source base model of each pseudo-label. During the entire training process, the pseudo-label management module will continuously update the pseudo-label pool.
[0044] Selector sub-module: For each base model, the selector sub-module combines the relative measure of the uncertainty of the unlabeled data to screen out an unlabeled subset (including multiple unlabeled data) suitable for each base model to learn and the optimized pseudo-labels from the pseudo-label pool. It should be noted that unlabeled data is screened separately for each base model because the value of the same unlabeled data is different for each base model, and selecting the most valuable unlabeled data for each base model for training can improve the training efficiency.
[0045] Among them, the value of the unlabeled data for each base model can be understood from the following aspects. 1. Robustness value: For the unlabeled data x and the base model m, under various data augmentations, the prediction uncertainty of the base model m for the unlabeled data x is relatively large, but for other base models, the prediction results under various data augmentations are relatively consistent and the confidence levels are generally high. Then, the pseudo-label of the unlabeled data x is both highly correct and suitable for providing to the base model m for learning, and the unlabeled data x has a relatively large value for the base model m. 2. Efficiency value: Although the tasks of each baseline model are the same, the distributions of the training data may vary greatly. For example, baseline model 1 is good at rainy and snowy day scenes, and baseline model 2 is good at sunny day scenes. Then, the unlabeled data of the sunny day scene can bring relatively large benefits to baseline model 1, that is, the efficiency value is relatively high.
[0046] Since the value of the unlabeled data is different for different base models, therefore, each base model can correspond to an unlabeled subset (multiple unlabeled data) as training data, and this unlabeled subset has the greatest value for this base model. The base model can obtain the greatest gain by learning on this unlabeled subset.
[0047] In the above application scenario, a data processing method is proposed in an embodiment of the present application, which can be applied to a data processing device. Refer to Figure 3 As shown in
[0048] Step 301: Obtain M base models, a labeled dataset, and an unlabeled dataset.
[0049] Exemplarily, the base model is a model that needs to be trained. The base model can be a model for implementing an image classification task, a model for implementing an image detection task, or a model for implementing an image segmentation task. There is no limitation on the function of this base model. The base model can be a machine learning model, such as a deep learning model, a neural network model, etc. There is no limitation on the type of this base model.
[0050] Exemplarily, for the M base models, they can be base models with different structures or base models with the same structure but different parameters. There is no limitation on this. For example, all M base models are models for implementing an image classification task, and all M base models support C categories (such as category 1, category 2, and category 3, etc.). Another example is that all M base models are models for implementing an image detection task. Another example is that all M base models are models for implementing an image segmentation task.
[0051] Exemplarily, the labeled dataset is a set composed of multiple labeled data, that is, the labeled dataset includes multiple labeled data (such as image data), and the labeled data is data with calibration information.
[0052] Exemplarily, the unlabeled dataset is a set composed of multiple unlabeled data, that is, the unlabeled dataset can include multiple unlabeled data (such as image data). The unlabeled data does not have calibration information and cannot directly participate in model training. It needs to be calibrated before it can participate in model training.
[0053] Step 302: For each unlabeled data in the unlabeled dataset, perform A times of data augmentation on the unlabeled data to obtain A data-augmented unlabeled data, where A can be a positive integer.
[0054] Exemplarily, when performing data augmentation on the unlabeled data, methods such as spatial transformation and / or color transformation can be used to perform data augmentation on the unlabeled data to obtain data-augmented unlabeled data. Among them, the spatial transformation can include but is not limited to image scale transformation, and the color transformation can include but is not limited to at least one of the following: image brightness transformation, image saturation transformation, and image contrast transformation.
[0055] Exemplarily, when performing A data augmentations on the unlabeled data, different data augmentation methods are used, that is, the unlabeled data is augmented A times using different augmentation methods. For example, the method of the second data augmentation is different from that of the first data augmentation, the method of the third data augmentation is different from that of the first data augmentation and also different from that of the second data augmentation, and so on.
[0056] Step 303: For each unlabeled data after data augmentation, input the unlabeled data after data augmentation into M basic models, and the M basic models output pseudo-labels corresponding to the unlabeled data. Alternatively, the unlabeled data (i.e., not the unlabeled data after data augmentation) can also be directly input into M basic models, and the M basic models output pseudo-labels corresponding to the unlabeled data.
[0057] In summary, it can be seen that for each unlabeled data in the unlabeled dataset, the unlabeled data can correspond to multiple pseudo-labels. The multiple pseudo-labels are the pseudo-labels output by the M basic models after inputting the unlabeled data (such as the unlabeled data after data augmentation) into the M basic models.
[0058] For example, for each unlabeled data in the unlabeled dataset, for the convenience of description, hereinafter, taking the unlabeled data a as an example, the unlabeled data a is augmented 3 times to obtain the unlabeled data a1 after data augmentation, the unlabeled data a2 after data augmentation, and the unlabeled data a3 after data augmentation.
[0059] Input the unlabeled data a1 into the basic model 1. The basic model 1 processes the unlabeled data a1 to obtain the predicted label corresponding to the unlabeled data a1 and the confidence corresponding to the predicted label. Based on the predicted label and confidence corresponding to the unlabeled data a1, the pseudo-label and confidence corresponding to the unlabeled data a1 can be determined, and there is no limitation on this. For example, after inputting the unlabeled data a1 into the basic model 1, the basic model 1 outputs the pseudo-label a11-1 corresponding to the unlabeled data a1 and the confidence a11-2 corresponding to the pseudo-label a11-1.
[0060] Similarly, after inputting the unlabeled data a1 into the basic model 2, the basic model 2 can output the pseudo-label a12-1 corresponding to the unlabeled data a1 and the confidence a12-2 corresponding to the pseudo-label a12-1.
[0061] And so on. After inputting the unlabeled data a1 into the basic model M, the basic model M can output the pseudo-label a1M-1 corresponding to the unlabeled data a1 and the confidence a1M-2 corresponding to the pseudo-label a1M-1.
[0062] Input the unlabeled data a2 into the base model 1. The base model 1 processes the unlabeled data a2 to obtain the pseudo-label a21-1 and the confidence a21-2 corresponding to the pseudo-label a21-1. Similarly, after inputting the unlabeled data a2 into the base model 2, the base model 2 outputs the pseudo-label a22-1 and the confidence a22-2 corresponding to the pseudo-label a22-1. And so on. After inputting the unlabeled data a2 into the base model M, the base model M outputs the pseudo-label a2M-1 and the confidence a2M-2 corresponding to the pseudo-label a2M-1.
[0063] Input the unlabeled data a3 into the base model 1. The base model 1 processes the unlabeled data a3 to obtain the pseudo-label a31-1 and the confidence a31-2 corresponding to the pseudo-label a31-1. Similarly, after inputting the unlabeled data a3 into the base model 2, the base model 2 outputs the pseudo-label a32-1 and the confidence a32-2 corresponding to the pseudo-label a32-1. And so on. After inputting the unlabeled data a3 into the base model M, the base model M outputs the pseudo-label a3M-1 and the confidence a3M-2 corresponding to the pseudo-label a3M-1.
[0064] As can be seen from the above, for each unlabeled data in the unlabeled dataset, taking the unlabeled data a as an example, this unlabeled data a can correspond to multiple pseudo-labels, as shown in Table 1. Obviously, when performing A times (such as 3 times) of data augmentation on the unlabeled data a, the number of pseudo-labels is A*M.
[0065] Table 1
[0066]
[0067] Based on the multiple pseudo-labels corresponding to each unlabeled data in the unlabeled dataset, and the confidence corresponding to each pseudo-label, an unlabeled subset (including multiple unlabeled data) that matches the base model can be selected for each base model. For example, an unlabeled subset 1 that matches the base model 1 is selected for the base model 1,..., and an unlabeled subset M that matches the base model M is selected for the base model M.
[0068] For the convenience of description, in the subsequent embodiments, taking the selection of the unlabeled subset m for the base model m (the base model m is any one of all the base models) as an example, the process of selecting the unlabeled subset for other base models is similar to the process of selecting the unlabeled subset m, and will not be elaborated here.
[0069] Regarding the process of selecting the unlabeled subset m, the data processing method of this embodiment may further include:
[0070] Step 304: For the base model m, for each unlabeled data in the unlabeled dataset, based on the multiple pseudo-labels corresponding to the unlabeled data, determine the first uncertainty of the unlabeled data with respect to the base model m and the second uncertainty of the unlabeled data with respect to the remaining base models other than the base model m.
[0071] Exemplarily, the multiple pseudo-labels corresponding to the unlabeled data can be divided into a first pseudo-label set and a second pseudo-label set. The pseudo-labels in the first pseudo-label set are the pseudo-labels output by the base model m, and the pseudo-labels in the second pseudo-label set are the pseudo-labels output by the remaining base models other than the base model m. Then, determine the first uncertainty based on the confidence levels corresponding to the pseudo-labels in the first pseudo-label set, and determine the second uncertainty based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set.
[0072] For example, taking the unlabeled data a in the unlabeled dataset as an example, the processing process of other unlabeled data is similar to that of the unlabeled data a. The multiple pseudo-labels corresponding to the unlabeled data a can be divided into a first pseudo-label set and a second pseudo-label set. Assume that the base model m is the base model 1, then the remaining base models other than the base model m are the base models 2 - M. On this basis, as shown in Table 1, the first pseudo-label set includes the pseudo-labels output by the base model 1, such as the pseudo-labels a11-1, a21-1, and a31-1, and the second pseudo-label set includes the pseudo-labels output by the base models 2 - M, such as the pseudo-labels a12-1, a22-1, a32-1, a13-1, a23-1, a33-1, …, a1M-1, a2M-1, a3M-1.
[0073] Based on this, based on the confidence levels corresponding to the pseudo-labels in the first pseudo-label set, the first uncertainty can be determined. For example, based on the confidence level a11-2 corresponding to the pseudo-label a11-1, the confidence level a21-2 corresponding to the pseudo-label a21-1, and the confidence level a31-2 corresponding to the pseudo-label a31-1, determine the first uncertainty. Based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set, the second uncertainty can be determined. For example, based on the confidence level a12-2 corresponding to the pseudo-label a12-1, the confidence level a22-2 corresponding to the pseudo-label a22-1, the confidence level a32-2 corresponding to the pseudo-label a32-1, …, the confidence level a1M-2 corresponding to the pseudo-label a1M-1, the confidence level a2M-2 corresponding to the pseudo-label a2M-1, and the confidence level a3M-2 corresponding to the pseudo-label a3M-1, determine the second uncertainty.
[0074] In a possible implementation, the uncertainty corresponding to the unlabeled data a is used to represent the inconsistency of the prediction results (i.e., pseudo-labels) of M base models for the unlabeled data a. That is to say, the greater the uncertainty, the greater the inconsistency of the prediction results of the M base models for the unlabeled data a.
[0075] For example, the uncertainty can be represented by mutual information (MI), or other attributes can also be used to represent the uncertainty, which is not limited here. In this embodiment, mutual information is taken as an example to represent the uncertainty. On this basis, based on the confidence levels corresponding to each pseudo-label in the first pseudo-label set, the entropy of the first average value can be determined (that is, first calculate the average value of the confidence levels corresponding to each pseudo-label, and then calculate the entropy of this average value). Based on the confidence levels corresponding to each pseudo-label in the first pseudo-label set, the average value of the first entropy can be determined (that is, first calculate the entropy values of the confidence levels corresponding to each pseudo-label, and then calculate the average value of these entropy values). Then, the first uncertainty can be determined based on the entropy of the first average value and the average value of the first entropy.
[0076] Similarly, based on the confidence levels corresponding to each pseudo-label in the second pseudo-label set, the entropy of the second average value can be determined. Based on the confidence levels corresponding to each pseudo-label in the second pseudo-label set, the average value of the second entropy can be determined. Then, the second uncertainty can be determined based on the entropy of the second average value and the average value of the second entropy.
[0077] For example, a representation form of mutual information can be seen in the following formula:
[0078]
[0079] In the above formula, MI represents mutual information, that is, uncertainty, H represents entropy, p represents the confidence level corresponding to the pseudo-label (the value is distributed between 0 and 1), n represents the total number of pseudo-labels, and p i represents the confidence level corresponding to the i-th pseudo-label. Obviously, mutual information can be briefly described as the difference between the entropy of the average value and the average value of the entropy.
[0080] For the unlabeled data a, A times of data augmentation are performed (such as 3 times of data augmentation), and M base models are used for prediction. That is, the unlabeled data a corresponds to M * A pseudo-labels, and M * A pseudo-labels correspond to M * A confidence levels. Then n is M * A. Another representation form of the above mutual information can be seen in the following formula:
[0081]
[0082] In the above formula, MI represents mutual information, which is also uncertainty, H represents entropy, p represents the confidence corresponding to the pseudo-label, M*A represents the total number of pseudo-labels, i represents the i-th base model, that is, the 1-M-th base model, j represents the j-th data augmentation, that is, the 1-A-th data augmentation, p ij represents the confidence corresponding to the pseudo-label output by the i-th base model for the unlabeled data after the j-th data augmentation.
[0083] In summary, it can be seen that mutual information can be briefly described as the difference between the entropy of the average value and the average of the entropies. Intuitively, if the confidences for a certain object in each pseudo-label are the same, the mutual information is 0. If the confidences in each pseudo-label are very inconsistent and the entropy of each confidence is small enough (the prediction is confident enough), the mutual information is large. Therefore, mutual information can be understood as a measure of the "divergence" of pseudo-labels based on the output of multiple base models with multiple data augmentations. The greater the mutual information of the pseudo-labels, the greater the divergence of the pseudo-labels, indicating that each pseudo-label has a different "view" on the unlabeled data. Conversely, the smaller the mutual information of the pseudo-labels, the smaller the divergence of the pseudo-labels. At this time, there are two cases. The first is that each base model gives an ambiguous judgment (such as the confidence is neither high nor low, distributed around 0.5), and the second is that each base model reaches a confident and unified opinion.
[0084] Based on the representation form of mutual information, the first uncertainty of the unlabeled data for the base model m can be determined, and the second uncertainty of the unlabeled data for the remaining base models M / m (M / m represents the remaining base models except the base model m among all M base models) can be determined. For example, the first uncertainty MI(P m ) and the second uncertainty MI(P M / m ) can be determined through the following formula.
[0085]
[0086]
[0087] In the above formula, MI(P m ) represents the first uncertainty for the base model m, H represents entropy, A represents the number of pseudo-labels in the first pseudo-label set (that is, the pseudo-labels output by the base model m), represents the confidence corresponding to the i-th pseudo-label output by the base model m. To sum up, based on the confidences corresponding to each pseudo-label in the first pseudo-label set the average value of each confidence can be determined, and then the entropy of this average value (that is, the entropy of the first average value) can be calculated. Based on the confidences corresponding to each pseudo-label in the first pseudo-label set The confidence levels can be calculated The entropy value can be calculated, and then the average value of the entropy value (i.e., the average value of the first entropy) can be calculated. Based on the entropy of the first average value and the average value of the first entropy, the first uncertainty can be determined.
[0088] In the above formula, MI(P M / m ) represents the second uncertainty for the remaining base models except base model m among all M base models. H represents entropy, (M - 1)×A represents the number of pseudo-labels in the second pseudo-label set (i.e., the pseudo-labels output by the remaining base models except base model m among all M base models). i represents the i-th base model, i.e., the 1-(M - 1)-th base model. Here, the base models are the remaining base models except base model m among all M base models, that is, there are a total of (M - 1) base models. j represents the j-th data augmentation, i.e., the 1-A-th data augmentation, represents the confidence level corresponding to the pseudo-label output by the i-th base model (the remaining base models except base model m) for the unlabeled data after the j-th data augmentation. To sum up, based on the confidence levels corresponding to each pseudo-label in the second pseudo-label set The average value of the confidence levels can be determined The average value of the confidence levels can be calculated, and then the entropy of the average value (i.e., the entropy of the second average value) can be calculated. Based on the entropy of the second average value and the average value of the second entropy, the second uncertainty can be determined. The entropy values of the confidence levels can be calculated and then the average value of the entropy values (i.e., the average value of the second entropy) can be calculated. Based on the entropy of the second average value and the average value of the second entropy, the second uncertainty can be determined.
[0089] Step 305: For each unlabeled data in the unlabeled dataset, based on the difference between the first uncertainty and the second uncertainty corresponding to the unlabeled data, determine the uncertainty difference of the unlabeled data for base model m and the remaining base models (i.e., the remaining base models except base model m).
[0090] For example, based on the first uncertainty MI(P m ) and the second uncertainty MI(P M / m ), the uncertainty difference ΔMI(m, M) corresponding to the unlabeled data can be determined through the following formula: ΔMI(m, M) = MI(P m ) - MI(P M / m ). Of course, this formula is just an example, and the determination method of this uncertainty difference is not limited.
[0091] Step 306: Determine the average confidence level based on the confidence levels corresponding to each pseudo-label in the second pseudo-label set.
[0092] The pseudo-labels in the second pseudo-label set are the pseudo-labels output by the remaining base models other than the base model m. Based on the confidence levels corresponding to each pseudo-label in the second pseudo-label set, the average value of the confidence levels corresponding to these pseudo-labels can be calculated, and the average value of the confidence levels corresponding to these pseudo-labels is used as the average confidence level.
[0093] Step 307: For each unlabeled data in the unlabeled dataset, if the uncertainty difference corresponding to the unlabeled data is greater than the first threshold, and the average confidence level corresponding to the unlabeled data is greater than the second threshold, then it is determined that the unlabeled data is the target unlabeled data corresponding to the base model m. If the uncertainty difference corresponding to the unlabeled data is not greater than the first threshold, and / or the average confidence level corresponding to the unlabeled data is not greater than the second threshold, then it is determined that the unlabeled data is not the target unlabeled data corresponding to the base model m.
[0094] Exemplarily, if the uncertainty difference is larger, it indicates that the mutual information of a certain pseudo-label in the pseudo-label set of a single base model (i.e., the base model m) is larger, and the prediction results based on different data augmentations have a relatively large gap, while the disagreement of this pseudo-label in the pseudo-label sets of the remaining base models (i.e., the remaining base models other than the base model m among all M base models) is small. Obviously, the disagreement of a single base model is large, and it is not sure about this unlabeled data, while the disagreement of the remaining base models is small and the consistency is high. Then, this unlabeled data is "meaningful" for training a single base model as it can bring additional information. The above process corresponds to the relative measurement of uncertainty, and a single base model and the remaining base models form a relative relationship. As shown in Table 2, it shows the corresponding relationship between the certainty (uncertainty) and meaning of unlabeled data. When the certainty of a certain unlabeled data for the base model m is "uncertain", and the certainty of this unlabeled data for the remaining base models is "certain", then this unlabeled data is valuable unlabeled data and needs to participate in the training as sample data for the base model m.
[0095] Table 2
[0096]
[0097] Exemplarily, the value of the second uncertainty MI(P M / m ) is small. It may be that each base model gives an ambiguous and unsure judgment (the confidence level is neither high nor low, distributed around 0.5), that is, the pseudo-label is uncertain for all base models, and this pseudo-label should be ignored. Therefore, the average confidence level can also be used for screening, and only the unlabeled data with a relatively high average confidence level is retained as the target pseudo-label of the base model m.
[0098] In summary, when determining whether unlabeled data is the target unlabeled data corresponding to the base model m, the selection criterion can be: ΔMI(m, M) > σ1, and, avg(P M / m ) > σ2. That is to say, for each unlabeled data, if the uncertainty difference ΔMI(m, M) corresponding to the unlabeled data is greater than the first threshold σ1 (the threshold used to represent the uncertainty difference), and the average confidence avg(P M / m ) corresponding to the unlabeled data is greater than the second threshold σ2 (the threshold used to represent the average confidence), then the unlabeled data is the target unlabeled data corresponding to the base model m; otherwise, the unlabeled data is not the target unlabeled data corresponding to the base model m.
[0099] Exemplarily, if the unlabeled data is the target unlabeled data corresponding to the base model m, the unlabeled data can be added to the unlabeled subset m corresponding to the base model m; if the unlabeled data is not the target unlabeled data corresponding to the base model m, adding the unlabeled data to the unlabeled subset m corresponding to the base model m is prohibited. Obviously, after performing the above processing on each unlabeled data in the unlabeled dataset, the unlabeled subset m corresponding to the base model m can be obtained. The unlabeled subset m can include multiple unlabeled data, and these unlabeled data are all the target unlabeled data corresponding to the base model m.
[0100] Step 308: Generate a target pseudo-label corresponding to the target unlabeled data based on multiple pseudo-labels corresponding to the target unlabeled data corresponding to the base model m. Obviously, the target unlabeled data and the target pseudo-label are equivalent to labeled data with calibration information, that is, the target pseudo-label serves as the calibration information.
[0101] For example, after obtaining the unlabeled subset m corresponding to the base model m, the unlabeled subset m includes multiple target unlabeled data corresponding to the base model m. For each target unlabeled data, the target unlabeled data corresponds to multiple pseudo-labels. Based on the multiple pseudo-labels corresponding to the target unlabeled data, a target pseudo-label corresponding to the target unlabeled data can be generated. For example, the multiple pseudo-labels corresponding to the target unlabeled data are fused, and the fused pseudo-label is used as the target pseudo-label. Of course, pseudo-label fusion is only an example and is not limited thereto, as long as the target pseudo-label is generated based on multiple pseudo-labels.
[0102] For example, if M base models are used to implement a classification task, the multiple pseudo-labels corresponding to the target unlabeled data can be class labels. Based on the confidence levels corresponding to the multiple class labels (i.e., the confidence levels corresponding to the pseudo-labels), the class label with the highest confidence level is used as the target pseudo-label corresponding to the target unlabeled data.
[0103] If M basic models are used to implement the detection task, the multiple pseudo-labels corresponding to the target unlabeled data can be prediction boxes (such as rectangular prediction boxes, and the prediction box can be represented by coordinates). A fused prediction box can be generated based on the multiple prediction boxes. The fused prediction box includes multiple prediction boxes, that is, the fused prediction box covers the regions of multiple prediction boxes. The fused prediction box is used as the target pseudo-label corresponding to the target unlabeled data.
[0104] Of course, the above method is only an example for determining the target pseudo-label, and there is no limitation on this.
[0105] Step 309: Train the basic model m based on the target unlabeled data and the target pseudo-label corresponding to the basic model m to obtain the trained target model. For example, the unlabeled subset m includes multiple target unlabeled data, and each target unlabeled data corresponds to a target pseudo-label. The basic model m can be trained based on the multiple target unlabeled data and the corresponding target pseudo-labels to obtain the target model.
[0106] When training the basic model m, the basic model m can also be trained based on the labeled data set (see Step 301). That is, the basic model m is trained based on the labeled data set and the unlabeled subset m. There is no limitation on the training process of the basic model m. Among them, the labeled data in the labeled data set all have calibration information, and the target unlabeled data in the unlabeled subset m all have calibration information (i.e., the target pseudo-labels), so as to train based on the data with calibration information.
[0107] In summary, based on Steps 304 - 309, each basic model can be trained to obtain the trained model corresponding to the basic model. For example, when training the basic model 1, the trained model 1a corresponding to the basic model 1 is obtained; when training the basic model 2, the trained model 2a corresponding to the basic model 2 is obtained, and so on. When training the basic model M, the trained model Ma corresponding to the basic model M is obtained.
[0108] If the training end condition is satisfied, the trained model 1a is used as the target model corresponding to the basic model 1, the trained model 2a is used as the target model corresponding to the basic model 2, and so on. The trained model Ma is used as the target model corresponding to the basic model M, that is, the target models corresponding to each basic model are obtained.
[0109] If the training end condition is not satisfied, the trained model 1a is used as the basic model 1, the trained model 2a is used as the basic model 2, and so on. The trained model Ma is used as the basic model M, and return to Step 303 to repeat the execution to obtain the trained model 1b corresponding to the basic model 1, the trained model 2b corresponding to the basic model 2, and so on. The trained model Mb corresponding to the basic model M is obtained.
[0110] If the training end condition is satisfied, the trained model 1b is used as the target model corresponding to the base model 1, the trained model 2b is used as the target model corresponding to the base model 2, and so on. The trained model Mb is used as the target model corresponding to the base model M, that is, the target model corresponding to each base model is obtained.
[0111] If the training end condition is not satisfied, the trained model 1b is used as the base model 1, the trained model 2b is used as the base model 2, and so on. The trained model Mb is used as the base model M, and the above steps are repeated. This will not be elaborated here as long as the target model corresponding to each base model can be obtained.
[0112] In the above embodiment, if the number of model iterations reaches a preset number threshold (which can be configured according to experience), it can be determined that the training end condition is satisfied; otherwise, it can be determined that the training end condition is not satisfied. For another example, if the training duration of the model reaches a preset duration threshold (which can be configured according to experience), it can be determined that the training end condition is satisfied; otherwise, it can be determined that the training end condition is not satisfied. For another example, if the model performance reaches the expected index, it can be determined that the training end condition is satisfied; otherwise, it can be determined that the training end condition is not satisfied. Of course, the above are just a few examples, and the training end condition is not limited thereto.
[0113] Exemplarily, after obtaining the target models (i.e., the target model corresponding to the base model 1, the target model corresponding to the base model 2,..., the target model corresponding to the base model M), these target models can be output, that is, the target models are deployed online, and the target models process the application data. That is to say, the application data can be input to the target models, and the target models process the application data to obtain a data processing result, that is, an artificial intelligence processing result. For example, if the target model is used to implement a classification task, the target model processes the application data to obtain a classification result; if the target model is used to implement a detection task, the target model processes the application data to obtain a detection result.
[0114] As can be seen from the above technical solutions, in the embodiments of the present application, during the training process, each basic model can indirectly interact with other basic models. It not only accepts the unlabeled subset that matches this basic model to participate in the training of this basic model, but also conveys the knowledge of this basic model to other basic models in the form of pseudo-labels to participate in the training of other basic models, and overall completes the process of knowledge sharing in a collaborative learning mode. For each basic model, the relative measure of uncertainty is used to screen the pseudo-labels, and the pseudo-labels suitable for each basic model are specifically selected for its training. Although there is noise in the pseudo-labels, since the noise often cannot reach a high prediction consistency in each amplification of each basic model, therefore, the noise can be naturally filtered to reduce the noise in the pseudo-labels received by each basic model. When optimizing the pseudo-labels, the uncertainty is used to screen the pseudo-labels, and for each basic model, a suitable subset of pseudo-labels is provided for training. It has a high filtering ability for the noise in the pseudo-labels, improves the quality of the pseudo-labels, and at the same time reduces the number of unlabeled data during the training of a single basic model, which is conducive to improving the performance and efficiency of semi-supervised learning. It can use multiple basic models to provide pseudo-labels, and the process of collaborative learning and optimization enables the knowledge of multiple basic models to be efficiently circulated and shared. It is not strongly coupled with a single task and can be adapted to tasks such as detection, classification, and segmentation, ensuring generality. It emphasizes the robustness and efficiency of semi-supervised learning, comprehensively utilizes the knowledge of multiple basic models, and selects suitable unlabeled data and corresponding pseudo-labels for each basic model. The value of the pseudo-labels is measured by uncertainty, and a subset of pseudo-labels suitable for the learning of each basic model is selected, while also ensuring a low noise ratio. Through the collaborative provision and screening of pseudo-labels by multiple basic models for learning, the role of knowledge transfer is achieved. It is a semi-supervised learning technology applicable to multiple fields and is applicable to various applications such as image classification and object detection. Using the relative measure of uncertainty for the value of pseudo-labels, making full use of the knowledge of multiple basic models to select the most suitable unlabeled data and corresponding pseudo-labels for the learning of each basic model, and training in a more efficient form, the performance and generalization ability of the basic model will be improved.
[0115] Based on the same application concept as the above method, in the embodiments of the present application, a data processing device is proposed. Refer to Figure 4 As shown in the figure, which is a schematic structural diagram of the data processing device, the device may include:
[0116] An acquisition module 41, configured to acquire an unlabeled data set, where the unlabeled data set includes multiple unlabeled data; for each unlabeled data, the unlabeled data corresponds to multiple pseudo-labels, and the multiple pseudo-labels are the pseudo-labels output by the multiple basic models after inputting the unlabeled data into the multiple basic models;
[0117] A determination module 42, configured to select, for each base model, target unlabeled data corresponding to the base model from the unlabeled dataset; wherein, for each unlabeled data in the unlabeled dataset, based on a plurality of pseudo-labels corresponding to the unlabeled data, determine a first uncertainty of the unlabeled data with respect to the base model and a second uncertainty of the unlabeled data with respect to the remaining base models other than the base model; and determine whether the unlabeled data is the target unlabeled data corresponding to the base model or not based on the first uncertainty and the second uncertainty.
[0118] A training module 43, configured to train the base model based on the target unlabeled data corresponding to the base model to obtain a trained target model; the target model is used to process application data.
[0119] Exemplarily, the obtaining module 41 is further configured to, for each unlabeled data in the unlabeled dataset, perform data augmentation on the unlabeled data A times to obtain A unlabeled data after data augmentation, where A is a positive integer; for each unlabeled data after data augmentation, input the unlabeled data after data augmentation into a plurality of base models, and the plurality of base models output pseudo-labels corresponding to the unlabeled data.
[0120] In a possible implementation manner, when the determination module 42 determines the first uncertainty of the unlabeled data with respect to the base model and the second uncertainty of the unlabeled data with respect to the remaining base models other than the base model based on the plurality of pseudo-labels corresponding to the unlabeled data, it is specifically configured to: divide the plurality of pseudo-labels corresponding to the unlabeled data into a first pseudo-label set and a second pseudo-label set; wherein, the pseudo-labels in the first pseudo-label set are the pseudo-labels output by the base model, and the pseudo-labels in the second pseudo-label set are the pseudo-labels output by the remaining base models other than the base model; determine the first uncertainty based on the confidence levels corresponding to the pseudo-labels in the first pseudo-label set; and determine the second uncertainty based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set.
[0121] In a possible implementation manner, when the determining module 42 determines the first uncertainty based on the confidence levels corresponding to the respective pseudo-labels in the first pseudo-label set, it is specifically configured to: determine the entropy of the first average value based on the confidence levels corresponding to the respective pseudo-labels in the first pseudo-label set, determine the average value of the first entropy based on the confidence levels corresponding to the respective pseudo-labels in the first pseudo-label set, and determine the first uncertainty based on the entropy of the first average value and the average value of the first entropy; wherein, when the determining module determines the second uncertainty based on the confidence levels corresponding to the respective pseudo-labels in the second pseudo-label set, it is specifically configured to: determine the entropy of the second average value based on the confidence levels corresponding to the respective pseudo-labels in the second pseudo-label set, determine the average value of the second entropy based on the confidence levels corresponding to the respective pseudo-labels in the second pseudo-label set, and determine the second uncertainty based on the entropy of the second average value and the average value of the second entropy.
[0122] In a possible implementation manner, when the determining module 42 determines whether the unlabeled data is the target unlabeled data corresponding to the base model or not based on the first uncertainty and the second uncertainty, it is specifically configured to: determine the uncertainty difference of the unlabeled data with respect to the base model and the remaining base models based on the difference between the first uncertainty and the second uncertainty; if the uncertainty difference is greater than a first threshold, determine that the unlabeled data is the target unlabeled data corresponding to the base model; or, if the uncertainty difference is not greater than the first threshold, determine that the unlabeled data is not the target unlabeled data corresponding to the base model.
[0123] In a possible implementation manner, when the determining module 42 determines whether the unlabeled data is the target unlabeled data corresponding to the base model or not based on the first uncertainty and the second uncertainty, it is specifically configured to: determine the uncertainty difference of the unlabeled data with respect to the base model and the remaining base models based on the difference between the first uncertainty and the second uncertainty; determine the average confidence level based on the confidence levels corresponding to the respective pseudo-labels in the second pseudo-label set; if the uncertainty difference is greater than a first threshold and the average confidence level is greater than a second threshold, determine that the unlabeled data is the target unlabeled data corresponding to the base model; if the uncertainty difference is not greater than the first threshold, and / or, the average confidence level is not greater than the second threshold, determine that the unlabeled data is not the target unlabeled data corresponding to the base model.
[0124] In a possible implementation, when the training module 43 trains the base model based on the target unlabeled data corresponding to the base model to obtain a trained target model, it is specifically configured to: generate a target pseudo-label corresponding to the target unlabeled data based on a plurality of pseudo-labels corresponding to the target unlabeled data corresponding to the base model; and train the base model based on the target unlabeled data and the target pseudo-label to obtain the trained target model.
[0125] Based on the same application concept as the above method, an embodiment of the present application proposes a data processing device. Refer to Figure 5 As shown, the data processing device may include: a processor 51 and a machine-readable storage medium 52, where the machine-readable storage medium 52 stores machine-executable instructions that can be executed by the processor 51; the processor 51 is configured to execute the machine-executable instructions to implement the data processing method disclosed in the above examples of the present application. For example, the processor 51 is configured to execute the machine-executable instructions to implement the following steps:
[0126] Obtain an unlabeled data set, where the unlabeled data set includes a plurality of unlabeled data; for each unlabeled data, the unlabeled data corresponds to a plurality of pseudo-labels, and the plurality of pseudo-labels are pseudo-labels output by a plurality of base models after the unlabeled data is input to the plurality of base models;
[0127] For each base model, select the target unlabeled data corresponding to the base model from the unlabeled data set; wherein, for each unlabeled data in the unlabeled data set, based on the plurality of pseudo-labels corresponding to the unlabeled data, determine a first uncertainty of the unlabeled data with respect to the base model and a second uncertainty of the unlabeled data with respect to the remaining base models other than the base model; determine whether the unlabeled data is the target unlabeled data corresponding to the base model or not based on the first uncertainty and the second uncertainty;
[0128] Train the base model based on the target unlabeled data corresponding to the base model to obtain a trained target model; wherein, the target model is used to perform data processing on application data.
[0129] Based on the same application concept as the above method, an embodiment of the present application further provides a machine-readable storage medium, where a number of computer instructions are stored on the machine-readable storage medium, and when the computer instructions are executed by a processor, the data processing method disclosed in the above examples of the present application can be implemented.
[0130] Among them, the above-mentioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, and so on. For example, the machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.
[0131] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, laptop computer, cellular phone, camera phone, smart phone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or a combination of any several of these devices.
[0132] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0133] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0134] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0135] Moreover, these computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 or a plurality of processes and / or blocks Figure 1 or the functions specified in a plurality of blocks.
[0136] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process Figure 1 or a plurality of processes and / or blocks Figure 1 or the functions specified in a plurality of blocks.
[0137] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A data processing method, characterized in that, The method includes: Obtaining an unlabeled image dataset, where the unlabeled image dataset includes multiple unlabeled image data; for each unlabeled image data, the unlabeled image data corresponds to multiple pseudo-labels, and the multiple pseudo-labels are obtained by performing data augmentation A times on the unlabeled image data to obtain A data-augmented unlabeled image data, and after inputting each data-augmented unlabeled image data into multiple base models, the pseudo-labels output by the multiple base models; where A is a positive integer greater than 1; For each base model, selecting the target unlabeled image data corresponding to the base model from the unlabeled image dataset; where for each unlabeled image data in the unlabeled image dataset, based on the multiple pseudo-labels corresponding to the unlabeled image data, determining the first uncertainty of the unlabeled image data with respect to the base model and the second uncertainty of the unlabeled image data with respect to the remaining base models other than the base model; based on the first uncertainty and the second uncertainty, determining whether the unlabeled image data is the target unlabeled image data corresponding to the base model or not the target unlabeled image data corresponding to the base model; where determining the entropy of the first average value based on the confidence levels corresponding to the pseudo-labels in the first pseudo-label set, and determining the average value of the first entropy based on the confidence levels corresponding to the pseudo-labels in the first pseudo-label set; determining the first uncertainty based on the entropy of the first average value and the average value of the first entropy; determining the second uncertainty based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set; where the pseudo-labels in the first pseudo-label set are the pseudo-labels corresponding to the unlabeled image data output by the base model, and the pseudo-labels in the second pseudo-label set are the pseudo-labels corresponding to the unlabeled image data output by the remaining base models other than the base model; Training the base model based on the target unlabeled image data corresponding to the base model to obtain a trained target model; where the target model is used to perform data processing on application image data.
2. The method according to claim 1, wherein The determining the second uncertainty based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set includes: determining the entropy of the second average value based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set, and determining the average value of the second entropy based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set; determining the second uncertainty based on the entropy of the second average value and the average value of the second entropy.
3. The method according to claim 1, wherein The determining whether the unlabeled image data is the target unlabeled image data corresponding to the base model or not the target unlabeled image data corresponding to the base model based on the first uncertainty and the second uncertainty includes: Based on the difference between the first uncertainty and the second uncertainty, determining the uncertainty difference of the unlabeled image data with respect to the base model and the remaining base models; If the uncertainty difference is greater than the first threshold, determine that the unlabeled image data is the target unlabeled image data corresponding to the base model; or, if the uncertainty difference is not greater than the first threshold, determine that the unlabeled image data is not the target unlabeled image data corresponding to the base model.
4. The method according to claim 1, characterized in that Determining that the unlabeled image data is the target unlabeled image data corresponding to the base model, or not the target unlabeled image data corresponding to the base model, based on the first uncertainty and the second uncertainty, includes: Based on the difference between the first uncertainty and the second uncertainty, determine the uncertainty difference of the unlabeled image data with respect to the base model and the remaining base models; Determine the average confidence based on the confidence corresponding to each pseudo-label in the second pseudo-label set; If the uncertainty difference is greater than the first threshold and the average confidence is greater than the second threshold, determine that the unlabeled image data is the target unlabeled image data corresponding to the base model; If the uncertainty difference is not greater than the first threshold, and / or the average confidence is not greater than the second threshold, determine that the unlabeled image data is not the target unlabeled image data corresponding to the base model.
5. The method according to claim 1, wherein Training the base model based on the target unlabeled image data corresponding to the base model to obtain a trained target model includes: Generate a target pseudo-label corresponding to the target unlabeled image data based on multiple pseudo-labels corresponding to the target unlabeled image data corresponding to the base model; train the base model based on the target unlabeled image data and the target pseudo-label to obtain the trained target model.
6. A data processing device, characterized in that, The apparatus includes: An acquisition module, configured to acquire an unlabeled image data set, the unlabeled image data set including multiple unlabeled image data; for each unlabeled image data, the unlabeled image data corresponds to multiple pseudo-labels, and the multiple pseudo-labels are obtained by performing A times of data augmentation on the unlabeled image data to obtain A data-augmented unlabeled image data, and after inputting each data-augmented unlabeled image data into multiple base models, the pseudo-labels output by the multiple base models; where A is a positive integer greater than 1; A determination module, configured to select, for each base model, target unlabeled image data corresponding to the base model from the unlabeled image dataset; wherein, for each unlabeled image data in the unlabeled image dataset, based on a plurality of pseudo-labels corresponding to the unlabeled image data, determine a first uncertainty of the unlabeled image data with respect to the base model and a second uncertainty of the unlabeled image data with respect to the remaining base models other than the base model; determine whether the unlabeled image data is the target unlabeled image data corresponding to the base model or not based on the first uncertainty and the second uncertainty; wherein, determine the entropy of the first average value based on the confidence levels corresponding to the pseudo-labels in the first pseudo-label set, and determine the average value of the first entropy based on the confidence levels corresponding to the pseudo-labels in the first pseudo-label set; determine the first uncertainty based on the entropy of the first average value and the average value of the first entropy; determine the second uncertainty based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set; wherein, the pseudo-labels in the first pseudo-label set are the pseudo-labels corresponding to the unlabeled image data output by the base model, and the pseudo-labels in the second pseudo-label set are the pseudo-labels corresponding to the unlabeled image data output by the remaining base models other than the base model. A training module, configured to train the base model based on the target unlabeled image data corresponding to the base model to obtain a trained target model; wherein, the target model is used to process application image data.
7. The device according to claim 6, characterized in that Wherein, When the determination module determines the second uncertainty based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set, it is specifically configured to: determine the entropy of the second average value based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set, determine the average value of the second entropy based on the confidence levels corresponding to the pseudo-labels in the second pseudo-label set, and determine the second uncertainty based on the entropy of the second average value and the average value of the second entropy; Wherein, when the determination module determines whether the unlabeled image data is the target unlabeled image data corresponding to the base model or not based on the first uncertainty and the second uncertainty, it is specifically configured to: determine the uncertainty difference of the unlabeled image data with respect to the base model and the remaining base models based on the difference between the first uncertainty and the second uncertainty; if the uncertainty difference is greater than a first threshold, determine that the unlabeled image data is the target unlabeled image data corresponding to the base model; or, if the uncertainty difference is not greater than the first threshold, determine that the unlabeled image data is not the target unlabeled image data corresponding to the base model; Wherein, when the determining module determines whether the unlabeled image data is the target unlabeled image data corresponding to the base model or not based on the first uncertainty and the second uncertainty, it is specifically used for: determining the uncertainty difference of the unlabeled image data for the base model and the remaining base models based on the difference between the first uncertainty and the second uncertainty; determining the average confidence based on the confidence corresponding to each pseudo-label in the second pseudo-label set; if the uncertainty difference is greater than a first threshold and the average confidence is greater than a second threshold, determining that the unlabeled image data is the target unlabeled image data corresponding to the base model; if the uncertainty difference is not greater than the first threshold, and / or, the average confidence is not greater than the second threshold, determining that the unlabeled image data is not the target unlabeled image data corresponding to the base model; Wherein, when the training module trains the base model based on the target unlabeled image data corresponding to the base model to obtain the trained target model, it is specifically used for: generating the target pseudo-label corresponding to the target unlabeled image data based on a plurality of pseudo-labels corresponding to the target unlabeled image data corresponding to the base model; training the base model based on the target unlabeled image data and the target pseudo-label to obtain the trained target model.
8. A data processing device, characterized in that, Including: a processor and a machine-readable storage medium, the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method steps according to any one of claims 1-5.
Citation Information
Patent Citations
Model training method and device, quality determination method and device, electronic equipment and storage medium
CN112149733A
Method To Decide A Labeling Priority To A Data
US20210027143A1