A Semi-Supervised Active Learning Method Based on Dual Classifiers
Through the dual classifier iterative update and pseudo-label automatic annotation method, the problem of insufficient utilization of unlabeled image information in existing active learning is solved, the classification accuracy and robustness of the model are improved, and the annotation cost is reduced.
Patent Information
- Application Number
- CN202310613548.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-05-29
AI Technical Summary
The existing active learning methods only select images that are difficult to classify for labeling in each learning step, and ignore unlabeled image information, resulting in a reduced robustness of the model.
Using a semi-supervised active learning method based on dual classifiers, the first and second picture classifiers are iteratively updated, and the JS divergence of the predicted probability distribution of unlabeled pictures and the double classifier information volume are selected, and combined with pseudo-label automatic annotation, improving the accuracy of picture selection.
It improves the classification accuracy and robustness of the model, reduces the annotation cost, enhances the model's utilization of unlabeled image information, and improves the representativeness of the image dataset.
Smart Images

Figure CN116704282B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine learning, and particularly relates to a semi-supervised active learning method based on a dual classifier. Background Art
[0002] In the field of computer vision, active learning is a machine learning method for alleviating the shortage of labeled images, aiming to achieve the expected performance of the target model with as few labeled images as possible, thereby significantly reducing the labeling cost of images. The key step is to design an image query strategy. By using the active learning method, the system can iteratively select unlabeled images that are difficult for the model to classify according to a specific query strategy and send them to the query user, asking the query user to return the true labels of these images.
[0003] Current active learning methods usually only select some images that are difficult to classify in each learning step and require the query user to label them. Although this process reduces image labeling, it often ignores the information of unlabeled images and reduces the robustness of the model. Summary of the Invention
[0004] The purpose of the present invention is to provide a semi-supervised active learning method based on a dual classifier, which combines a dual classifier with a semi-supervised learning method to make full use of the information of unlabeled images to improve the robustness of the model.
[0005] The present invention adopts the following technical solutions: A semi-supervised active learning method based on a dual classifier, comprising the following steps:
[0006] Obtain an image set to be labeled, and divide the image set to be labeled into a first image set and a second image set; the number of the first image set is less than that of the second image set;
[0007] Iteratively update the first image set and the second image set through a first image classifier and a second image classifier until the number of the first image set reaches an image number threshold, and complete the classification of the image set to be labeled;
[0008] Among them, iterative update is performed using the Jensen-Shannon (JS) divergence of the predicted probability distribution of the images in the second image set.
[0009] Further, iterative update using the JS divergence of the predicted probability distribution of the images in the second image set and the information amount of the dual classifier includes:
[0010] Extract the feature data of each image in the second image set through a feature extractor;
[0011] Input the feature data into the trained first image classifier to obtain a corresponding first predicted probability distribution, and input the feature data into the trained second image classifier to obtain a corresponding second predicted probability distribution;
[0012] Calculate the Jensen-Shannon divergence according to the first predicted probability distribution and the second predicted probability distribution;
[0013] Select the corresponding image according to the Jensen-Shannon divergence and add it to the first image set, and delete the image from the second image set.
[0014] Furthermore, calculating the Jensen-Shannon divergence according to the first predicted probability distribution and the second predicted probability distribution includes:
[0015]
[0016] Among them, JSD(p(x)||q(x)) represents the Jensen-Shannon divergence between the first predicted probability distribution and the second predicted probability distribution, p(x) represents the first predicted probability distribution, q(x) represents the second predicted probability distribution, and x represents the image in the second image set.
[0017] Furthermore, selecting the corresponding image according to the Jensen-Shannon divergence and adding it to the first image set and deleting the image from the second image set includes:
[0018] Select b images from the second image set in descending order of the Jensen-Shannon divergence and add them to the first image set, and delete the corresponding images from the second image set.
[0019] Furthermore, each iteration update also includes:
[0020] Calculate the double-classifier information amount according to the first predicted probability distribution and the second predicted probability distribution;
[0021] Select m images from the second image set in ascending order of the double-classifier information amount and add them to the first image set to obtain the first image training set; among them, the label corresponding to each of the m images is determined according to the first predicted probability distribution;
[0022] Train the first image classifier with the first image training set, and train the second image classifier with a subset of the first image training set.
[0023] Furthermore, calculating the double-classifier information amount according to the first predicted probability distribution and the second predicted probability distribution includes:
[0024] InfoTC = entropy (k) (x) + entropy (k′) (x),
[0025] Among them, InfoTC represents the double-classifier information amount, and entropy (k) (x) represents the classifier information amount of the first image classifier for the image x, and entropy (k′)(x) represents the classifier information amount of the second image classifier for image x.
[0026] Furthermore, the calculation methods of the classifier information amounts of the first image classifier and the second image classifier for image x are as follows:
[0027]
[0028]
[0029] where entropy (k) (·) represents the classifier information amount of the first image classifier, C represents the total number of labels of the image, represents the probability that the predicted label of image x by the first image classifier with parameter θ is c, entropy (k′) (·) represents the classifier information amount of the second image classifier, represents the probability that the predicted label of image x by the second image classifier with parameter θ′ is c.
[0030] Furthermore, the loss function of the first image classifier is:
[0031]
[0032] where, represents the loss function of the first image classifier, represents the expectation on the first image training set D, c represents the predicted label of image x determined according to the first prediction probability distribution, y represents the true label of image x, is an indicator function, which outputs 1 if a is true and 0 otherwise, p c (x) represents the probability of classifying the predicted label of image x into c, D represents the first image training set;
[0033] The loss function of the second image classifier is:
[0034]
[0035] where, represents the expectation on the subset S of the first image training set, represents the loss function of the second image classifier, q c (x) represents the probability of classifying the predicted label of image x into c, S represents the subset of the first image training set.
[0036] Furthermore, after dividing the image set to be labeled into the first image set and the second image set, it further includes:
[0037] Receiving the true label of each image in the first image set.
[0038] Another technical solution of the present invention: A semi-supervised active learning method based on a dual classifier, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above method is implemented.
[0039] The beneficial effects of the present invention are as follows: Through the dual classifier mode, the information of unlabeled pictures can be introduced during the iterative update process. At the same time, by combining the JS divergence of the predicted probability distributions of the dual classifiers to select pictures, the accuracy of picture selection can be greatly improved, so that the pictures in the first picture set are more representative, and thus the classification accuracy of the model trained with the picture data set after active learning is higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic diagram of a semi-supervised active learning method based on a dual classifier according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The present invention will be described in detail below with reference to the drawings and specific embodiments.
[0042] The present invention discloses a semi-supervised active learning method based on a dual classifier, including the following steps: obtaining a picture set to be labeled, and dividing the picture set to be labeled into a first picture set and a second picture set; the number of the first picture set is less than that of the second picture set; performing iterative update on the first picture set and the second picture set through a first picture classifier and a second picture classifier until the number of the first picture set reaches a picture number threshold, and completing the classification of the picture set to be labeled; wherein, the JS divergence of the predicted probability distributions of the pictures in the second picture set is used for iterative update.
[0043] Through the dual classifier mode, the information of unlabeled pictures can be introduced during the iterative update process. At the same time, by combining the JS divergence of the predicted probability distributions of the dual classifiers to select pictures, the accuracy of picture selection can be greatly improved, so that the pictures in the first picture set are more representative, and thus the classification accuracy of the model trained with the picture data set after active learning is higher.
[0044] Semi-supervised learning (SSL) is a key issue in the fields of pattern recognition and machine learning, and it is a learning method that combines supervised learning and unsupervised learning. Semi-supervised learning uses a large number of unlabeled pictures and a small number of labeled pictures at the same time for pattern recognition work. When using semi-supervised learning, it will require as few people as possible to do the work, and at the same time, it can bring relatively high accuracy. Therefore, semi-supervised learning is attracting more and more attention.
[0045] The method of the present invention automatically labels pseudo - labels for unlabeled pictures using self - training semi - supervised learning methods, and uses the pictures with pseudo - labels to improve the classification performance of the model. At the same time, pictures with the largest prediction differences between two classifiers are selected and sent to the user to return the true labels, which can effectively query the pictures that are difficult for the model to classify, thus reducing the labeling cost. Combining the active learning method and the self - training semi - supervised learning method can achieve complementary advantages and further reduce the labeling cost of samples without reducing the classification performance.
[0046] The present invention proposes to jointly train two classifiers with all labeled pictures and a subset of labeled pictures. The former is the main classifier (i.e., the first picture classifier), which is not only used to measure the difference between the two classification networks, but also used to predict and evaluate the final picture classification task. The secondary classifier (i.e., the second picture classifier) is trained from a subset of labeled pictures. The classifier trained in this way can not only retain the important information of the main classifier, but also make a certain distinction from the main classifier.
[0047] In this method, first, the picture set to be labeled is divided into a first picture set and a second picture set. The first picture set is randomly selected, and the number of pictures in it is small (less than the picture number threshold) and much smaller than the number of pictures in the second picture set. Then these pictures are sent to the querier for manual labeling. After the labeling is completed, the true label of each picture in the first picture set is received as the initial picture set for the iterative update of the network.
[0048] Specifically, for the dual - classifier model, F(·) represents the feature extractor, C(·) represents the first picture classifier, represents the first deep neural network, and C′(·) is the second picture classifier.
[0049] The iterative update using the Jensen - Shannon (JS) divergence of the predicted probability distribution of pictures in the second picture set and the information quantity of the dual - classifier includes: extracting the feature data of each picture in the second picture set through the feature extractor; inputting the feature data into the trained first picture classifier to obtain the corresponding first predicted probability distribution, and inputting the feature data into the trained second picture classifier to obtain the corresponding second predicted probability distribution; calculating the JS divergence according to the first predicted probability distribution and the second predicted probability distribution; selecting the corresponding pictures according to the JS divergence and adding them to the first picture set and deleting these pictures from the second picture set.
[0050] In the process of iterative update, first, the labeled picture set D(x, y) (i.e., the above - mentioned initial picture set) is input into the feature extractor, where x represents the picture and y represents the corresponding true label. Then, a predicted probability distribution is output through the main classifier C(·)
[0051] For example, for the input image x, if the image set to be recognized contains 4 categories, then p(x) = {0.8, 0.1, 0.05, 0.05}. The first image classifier C(·) is trained using the cross-entropy loss, as shown in the following equation:
[0052]
[0053] where represents the loss function of the first image classifier, represents the expectation on the first image training set D, c represents the predicted label of the image x determined according to the first predicted probability distribution, y represents the true label of the image x, is the indicator function, which outputs 1 if a is true and 0 otherwise, p c (x) represents the probability of classifying the predicted label of the image x into c, and D represents the first image training set;
[0054] Next, a subset S(x, y) is constructed from the first image set. This subset can be randomly selected, and the second image classifier is trained using this subset. During the training process, the second image classifier will obtain the predicted probability distribution q(x) for each image, that is The loss function of the second image classifier is:
[0055]
[0056] where represents the expectation on the subset S of the first image training set, represents the loss function of the second image classifier, q c (x) represents the probability of classifying the predicted label of the image x into c, and S represents the subset of the first image training set.
[0057] During the network training process, before the second image classifier C′(·) passes the output, its output is separated from F(·) so as not to affect the first image classifier of the task. Although the sub-classifiers are still trained using the typical cross-entropy loss, their main goal is not the correct classification task, but to learn the decision boundaries for active learning and semi-supervised learning sample queries. Specifically, the training process of the dual-classifier is shown in Table 1.
[0058] Table 1 Dual-classifier training process
[0059]
[0060]
[0061] In active learning, in each round of active learning training, some unlabeled images that are most difficult for the model to classify are searched for. After manual annotation, they are added to the labeled image pool for relevant training of the model. Based on the active learning strategy of dual classifiers, two similar but different classifiers are trained. If the prediction results of the two classification networks for a certain image are inconsistent, then this image is very likely to be the most difficult to classify. Using the active learning strategy to query out these most difficult to classify images can reduce the annotation cost.
[0062] Based on the above assumptions, the present invention proposes an index called the classifier prediction difference degree, denoted as DoPD (Degree of Prediction Difference), to represent the prediction difference between the first image classifier and the second image classifier. Given an image x, the prediction result of the first image classifier for it is p(x), and the prediction result of the second image classifier for it is q(x). The calculation of the classifier prediction difference degree is as follows:
[0063] DoPD = JSD(p(x)||q(x)),
[0064] Furthermore, calculating the JS divergence according to the first prediction probability distribution and the second prediction probability distribution includes:
[0065]
[0066] Among them, JSD(p(x)||q(x)) represents the JS divergence between the first prediction probability distribution and the second prediction probability distribution, p(x) represents the first prediction probability distribution, q(x) represents the second prediction probability distribution, and x represents the image in the second image set.
[0067] In addition, selecting the corresponding images according to the JS divergence and adding them to the first image set and deleting the images in the second image set includes: selecting b images in the second image set in descending order of the JS divergence and adding them to the first image set and deleting the corresponding images in the second image set.
[0068] Specifically, sorting the calculated classifier prediction difference degrees, and the selected b images are as shown in the formula:
[0069]
[0070] Among them, It represents the set of selected pictures. These pictures are deleted from the second picture set, sent to the querier at the same time, and the true labels of these pictures are received from the querier to form a new first picture set. However, the information of unlabeled pictures is not utilized at present. Therefore, some unlabeled pictures also need to be selected from the second picture set and added to the new first picture set to form the final first picture training set for training the first picture classifier and the second picture classifier.
[0071] Specifically, each iteration update also includes: calculating the information quantity of the double classifier according to the first prediction probability distribution and the second prediction probability distribution; selecting m pictures from the second picture set in ascending order of the information quantity of the double classifier and adding them to the first picture set to obtain the first picture training set; where the label corresponding to each of the m pictures is determined according to the first prediction probability distribution; training the first picture classifier through the first picture training set, and training the second picture classifier through a subset of the first picture training set.
[0072] Most pictures that the model can highly confirm are called highly confirmed pictures. The predicted labels of such pictures are very likely to be the true labels. For example, the predicted probability distribution of picture x1 is p(x1) = {0.8, 0.1, 0.05, 0.05}, and the confidence level of predicting it as class 1 is 0.8. The predicted probability distribution of picture x2 is p(x2) = {0.4, 0.1, 0.3, 0.2}, and the confidence level of predicting it as class 1 is 0.4. Relatively speaking, picture x1 is a picture that the model can highly confirm, while x2 is not.
[0073] Automatically select highly confirmed pictures from unlabeled pictures based on the semi-supervised learning strategy of the double classifier and automatically assign pseudo-labels to them through the first picture classifier, without manual cost. That is to say, if the model predicts a high confidence level for a certain picture on both the first picture classifier and the second picture classifier, the information quantity contained in such pictures is relatively small. Using the semi-supervised learning strategy to select these pictures that the model is already relatively certain about and jointly form the first training picture set with the labeled pictures to train the first picture classifier and the second picture classifier can reduce the labeling cost and also achieve the technical effect of utilizing the information of unlabeled pictures.
[0074] Specifically, the present invention proposes an index called the information quantity of the double classifier, denoted as InfoTC (Information of Two Classifiers). The smaller the information quantity of the double classifier, the greater the degree of certainty of the model for the picture. Given picture x, the entropy value of the prediction result of the first picture classifier for x is entropy (k) (x), and the entropy value of the prediction result of the second picture classifier for x is entropy (k′)If (x), then calculating the double-classifier information quantity according to the first prediction probability distribution and the second prediction probability distribution includes:
[0075] InfoTC = entropy (k) (x) + entropy (k′) (x),
[0076] where InfoTC represents the double-classifier information quantity, and entropy (k) (x) represents the classifier information quantity of the first image classifier for image x, and entropy (k′) (x) represents the classifier information quantity of the second image classifier for image x.
[0077] More specifically, the calculation methods of the classifier information quantities of the first image classifier and the second image classifier for image x are as follows:
[0078]
[0079]
[0080] where entropy (k) (·) represents the classifier information quantity of the first image classifier, C represents the total number of labels of the image, represents the probability that the predicted label of image x by the first image classifier with parameter θ is c, and entropy (k′) (·) represents the classifier information quantity of the second image classifier, represents the probability that the predicted label of image x by the second image classifier with parameter θ′ is c.
[0081] Then, for m images with pseudo-labels, select the image with the smallest double-classifier information quantity. The formula is as follows:
[0082]
[0083] where, is a set of m images with relatively small double-classifier information quantities.
[0084] After selecting the images, assign one-hot labels to the selected images, that is, use the largest category predicted by the machine learning model (the first main image classifier) as the hard pseudo-label y of the image i , and train the classifier model together with the images annotated by the query user. After each round of training, erase the pseudo-label information of these images and put these images back into the second image set again to avoid reusing these images in subsequent training.
[0085] In summary, the present invention rationally combines the advantages of semi-supervised learning and active learning based on self-training using a dual-classifier strategy, and can maximize the image classification performance with the minimum labor cost.
[0086] To verify the technical effects of the present invention, the following simulation experiments were also carried out. The method 2CoSAL proposed by the present invention was compared with passive learning Random based on the random sampling method, and classic comparison methods VAAL, Learning Loss, CDAL, and CEAL commonly used in active learning in recent years.
[0087] (1) Experiments on CIFAR-10.
[0088] For the CIFAR-10 dataset, the initial sample pool for the experiment was set to 4000 images (i.e., labeled images). During training, the active learning model was used to sample 2000 samples from the remaining unlabeled sample pool each time. The experimental results are shown in Table 2. When the number of labeled image samples reached 6000, this method and most other active learning methods had a certain improvement compared to the random method. However, compared with other active learning methods, the improvement of this method was the largest, with an accuracy of 85.13%. The accuracy of the relatively good comparison method CEAL was 85.01%, the accuracy of CDAL was 84.15%, and the random method (i.e., passive learning) was 83.35%. VAAL had an accuracy of 82.89% which was lower than the random method. Subsequently, this method still significantly exceeded other methods, and the gap gradually widened. In the last round, this method reached an accuracy of 91.83%, far exceeding the accuracy of 87.63% of the random method. Although the accuracies of other active learning methods also had a large improvement compared to the random method, except for CEAL, most of these methods only had an accuracy of about 89%. Although the CEAL method had a similar accuracy to this method in the early stage, in the last few rounds, its accuracy was about 1% lower than this method. The experiments on CIFAR-10 demonstrated the superiority of this method.
[0089] Table 2 Experimental results in CIFAR-10
[0090]
[0091] (2) Experiments on CIFAR-100.
[0092] For the CIFAR-100 dataset, the initial sample pool for the experiment was set to 4,000 images. During training, the active learning model was used to sample 2,000 samples from the remaining unlabeled sample pool each time. The experimental results of each method are shown in Table 3. When the number of labeled image samples reached 6,000, this method, the Learning Loss method, and the CEAL method all had a certain improvement compared to the random method. However, this method had a greater improvement, with its accuracy reaching 47.15%. The best-performing comparative method, CEAL, had an accuracy of 46.29%, and the random method had an accuracy of 43.80%. VAAL had an accuracy of 40.29%, which was lower than the random method. Subsequently, this method still exceeded the Learning Loss method and the CEAL method. In the last round, this method reached an accuracy of 67.67%, far exceeding the accuracy of the random method, which was 62.07%. The accuracies of CEAL, Learning Loss, and VAAL reached 67.59%, 65.35%, and 65.42% respectively. It is worth noting that the VAAL method was lower than or equal to the random method before selecting 12,000 samples and only had a significant improvement compared to the random method after 14,000 samples.
[0093] Table 3 Experimental results in CIFAR-100
[0094]
[0095] The present invention also discloses a semi-supervised active learning method based on a dual classifier, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above method is implemented.
[0096] Another embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above various method embodiments can be implemented.
[0097] Another embodiment of the present invention provides a computer program product. When the computer program product runs on a data storage device, the data storage device can be made to execute the steps in the above various method embodiments.
[0098] When the integrated unit module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to a storage device, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0099] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0100] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0101] In the embodiments provided by the present invention, it should be understood that the disclosed device / equipment and method can be implemented in other ways. For example, the device / equipment embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0102] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0103] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present application.
Claims
1. A semi-supervised active learning method based on a dual classifier, characterized in that Including the following steps: Obtain the picture set to be labeled, and divide the picture set to be labeled into a first picture set and a second picture set; the number of the first picture set is less than that of the second picture set; Iteratively update the first picture set and the second picture set through a first picture classifier and a second picture classifier until the number of the first picture set reaches a picture number threshold, and complete the classification of the picture set to be labeled; Wherein, the JS divergence of the predicted probability distribution of the pictures in the second picture set is used to iteratively update the first picture set and the second picture set; Using the JS divergence of the predicted probability distribution of the pictures in the second picture set and the double-classifier information quantity for iterative update includes: Extract the feature data of each picture in the second picture set through a feature extractor; Input the feature data into the trained first picture classifier to obtain the corresponding first predicted probability distribution, and input the feature data into the trained second picture classifier to obtain the corresponding second predicted probability distribution; Calculate the JS divergence according to the first predicted probability distribution and the second predicted probability distribution; Select the corresponding pictures according to the JS divergence and add them to the first picture set and delete the pictures in the second picture set; Select the corresponding pictures according to the JS divergence and add them to the first picture set and delete the pictures in the second picture set includes: Select b pictures from the second picture set in the order of the JS divergence from large to small, add them to the first picture set, and delete the corresponding pictures in the second picture set; Each time of iterative update further includes: Calculate the double-classifier information quantity according to the first predicted probability distribution and the second predicted probability distribution; Select m pictures from the second picture set in the order of the double-classifier information quantity from small to large and add them to the first picture set to obtain a first picture training set; wherein, the label corresponding to each picture in the m pictures is determined according to the first predicted probability distribution; Train the first picture classifier through the first picture training set, and train the second picture classifier through a subset of the first picture training set.
2. The semi-supervised active learning method based on a dual classifier according to claim 1, wherein Calculating the JS divergence according to the first predicted probability distribution and the second predicted probability distribution includes: Wherein, JSD(p(x)q(x)) represents the JS divergence between the first predicted probability distribution and the second predicted probability distribution, p(x) represents the first predicted probability distribution, q(x) represents the second predicted probability distribution, and x represents the pictures in the second picture set.
3. The semi-supervised active learning method based on a dual classifier according to claim 2, characterized in that Calculating the double-classifier information quantity according to the first predicted probability distribution and the second predicted probability distribution includes: InfoTC = entropy (k) (x) + entropy (k′) (x), Among them, InfoTC represents the information quantity of the double classifier, and entropy (k) (x) represents the classifier information quantity of the first image classifier for the image x, and entropy (k′) (x) represents the classifier information quantity of the second image classifier for the image x.
4. The semi-supervised active learning method based on a dual classifier according to claim 3, characterized in that The method for calculating the classifier information quantity of the first picture classifier and the second picture classifier for the picture x is: Among them, entropy (k) (·) represents the classifier information amount of the first image classifier, C represents the total number of labels of the image, represents the probability that the predicted label of the image x by the first image classifier with parameter θ is c, entropy (k′) (·) represents the classifier information amount of the second image classifier, represents the probability that the predicted label of the image x by the second image classifier with parameter θ′ is c.
5. The semi-supervised active learning method based on a dual classifier according to claim 4, characterized in that, The loss function of the first picture classifier is: Among them, represents the loss function of the first image classifier, represents the expectation on the first image training set D, c represents the predicted label of the image x determined according to the first predicted probability distribution, y represents the true label of the image x, is an indicator function that outputs 1 if a is true and 0 otherwise, p c (x) represents the probability that the predicted label of the image x is classified as c, and D represents the first image training set; The loss function of the second picture classifier is: Among them, represents the expectation on the subset S of the first image training set, represents the loss function of the second image classifier, q c (x) represents the probability that the predicted label of the image x is classified as c, and S represents the subset of the first image training set.
6. The semi-supervised active learning method based on a dual classifier according to claim 5, wherein, After dividing the picture set to be labeled into a first picture set and a second picture set, it further includes: Receive the true label of each picture in the first picture set.
7. A semi-supervised active learning device based on a dual classifier, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Multiclass image classification method based on active learning and semi-supervised learning
CN101853400A
Semi-supervised learning-based extraterrestrial picture segmentation method and device
CN115205309A