A Semi-Supervised Active Learning Method Based on Semantic Consistency

Semantic-aware cropping and loss-based training enhance semi-supervised learning by reducing annotation costs and improving classification precision through targeted feature identification and iterative model refinement.

CN115759243BActive Publication Date: 2025-07-15NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211454359.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2025-07-15
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

In the prior art, image samples are marked with high cost and low classification accuracy, and random cutting methods are prone to missing foreground objects, resulting in the problem of incorrectly cropping samples.

Method used

Semi-supervised active learning method with semantic consistency is adopted to perform semantic perceptual clipping of the original image, combining semi-supervised learning and active learning, and by calculating semi-supervised losses and active learning losses, the classification weight of the image classification network is updated to avoid incorrectly cropping samples.

Benefits of technology

It reduces the cost of image sample annotation, improves image classification accuracy, reduces the error-clipping sample rate, improves detection efficiency and classifier accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115759243B_ABST
    Figure CN115759243B_ABST
Patent Text Reader

Abstract

The present invention discloses a semi-supervised active learning method based on semantic consistency. First, semantic-aware cropping is used to localize the sample images, which solves the problem that random cropping may lead to false positives, thereby causing the model to misinterpret the cropped content and improving the classification accuracy of the model. Secondly, a semi-supervised learning is used to train the image classification network. Finally, through active learning, the labeled sample set is expanded, and the total loss of the model is determined by combining the semi-supervised loss and the active learning loss. The classification weights of the total loss are iteratively updated until the preset classification accuracy is satisfied, achieving the technical effects of reducing the cost of adding labels and improving the classification accuracy of the classification network model during the image classification process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a semi-supervised active learning method based on semantic consistency. Background Art

[0002] Active learning is a special case of machine learning, which aims to maximize the performance gain of the model while annotating as few samples as possible. Active learning can use a sample selection strategy to screen out some of the most difficult-to-classify samples, that is, the most representative samples, and hand them over to humans for confirmation, review, and annotation. Then, the data labeled by humans is used to train the model to further improve the model's performance.

[0003] The query method of uncertainty sampling adopted by active learning is to extract the sample data that is difficult to distinguish in the model and provide it to business experts or annotators for annotation, so as to achieve the ability to improve the algorithm effect at a relatively fast speed.

[0004] Semi-supervised learning (SSL) is a key issue in the fields of pattern recognition and machine learning, and it is a learning method that combines supervised learning and unsupervised learning. Semi-supervised learning uses a large amount of unlabeled data and a small amount of labeled data at the same time for pattern recognition work. When using semi-supervised learning, it will require as few people as possible to do the work, and at the same time, it can bring relatively high accuracy. Therefore, semi-supervised learning is attracting more and more attention.

[0005] Before annotation, random cropping is usually used to crop the sample images, which is prone to false positives, resulting in the problem of easily missing foreground objects and causing incorrect cropped samples. On the one hand, it increases the cost of manual annotation, and on the other hand, there is a technical problem of low accuracy in image classification.

[0006] Therefore, a new semi-supervised active learning method needs to be proposed to solve the technical problems existing in the above-mentioned prior art. Summary of the Invention

[0007] The purpose of the present invention is to provide a semi-supervised active learning method based on semantic consistency to solve the problems of how to reduce the cost of image sample annotation and how to improve the accuracy of image classification.

[0008] The present invention adopts the following technical solutions:

[0009] Embodiment 1 of the present invention provides a semi-supervised active learning method based on semantic consistency, including:

[0010] Perform semantic-aware cropping on each original image to obtain the corresponding cropped image after cropping. The cropped image contains the image features to be labeled;

[0011] Pass the original image and the cropped image through semi-supervised learning to output the semi-supervised learning loss;

[0012] Pass the cropped image corresponding to each original image through active learning to output the active learning loss;

[0013] Perform weighting on the active learning loss and the semi-supervised learning loss to obtain the total loss;

[0014] Update the classification weights of the image classification network according to the total loss through backpropagation.

[0015] Optionally, perform semantic-aware cropping on each original image to obtain the corresponding cropped image after cropping. The cropped image contains the image features to be labeled, including:

[0016] Take the original image as the input of the cropping network, and calculate the image features of the cropped image and the position information of the image feature points;

[0017] Normalize the image features and output the corresponding heatmap;

[0018] Judge whether the brightness of the heatmap is greater than the preset activation point threshold. When it is greater than the preset threshold, output the cropped image containing the height, width, and the position information of the image feature points.

[0019] Optionally, pass the original image and the cropped image through semi-supervised learning to output the semi-supervised learning loss, including:

[0020] Perform softmax prediction on the original image and the cropped image, and respectively output the first prediction result of the original image and the first prediction result of the cropped image;

[0021] Perform KL divergence calculation on the first prediction result of the original image and the first prediction result of the cropped image, and output the semi-supervised learning loss.

[0022] Optionally, pass the cropped image corresponding to each original image through active learning to output the active learning loss, including:

[0023] Perform softmax prediction on the original image and the cropped image, and respectively output the second prediction result of the original image and the second prediction result of the cropped image;

[0024] Perform cross-entropy calculation on the second prediction result of the original image and the second prediction result of the cropped image, and output the active learning loss.

[0025] Optionally, after calculating the cross entropy between the second prediction result of the original image and the second prediction result of the cropped image, the following steps are further included:

[0026] Sort the original images according to the calculated cross entropy, and select a preset number of original images with the largest cross entropy for annotation;

[0027] Append the annotated images to the initial label sample set to form a new label sample set.

[0028] Optionally, the initial label sample set is used to train an image classification network to determine the initial value of the classification weight.

[0029] Embodiment 2 of the present invention provides a semi-supervised active learning device based on semantic consistency, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a semi-supervised active learning method based on semantic consistency as described in any one of the above.

[0030] The beneficial effects of the present invention are as follows: By performing semantic-aware cropping on the original images, cropped images containing key image features are located, and compared with the random cropping method, foreground objects will not be missed to cause incorrect cropped samples; and the semi-supervised loss and the active learning loss are calculated in sequence, and the classifier is trained based on the total loss, which improves the accuracy of semi-supervised learning and active learning, as well as the cost of small sample annotation. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a flowchart of a semi-supervised active learning method based on semantic consistency provided in Embodiment 1 of the present invention;

[0032] Figure 2 It is a schematic diagram of a semi-supervised active learning training process provided in Embodiment 1 of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0033] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0034] Combined with Figure 1 Embodiment 1 of the present invention provides a semi-supervised active learning method based on semantic consistency, including:

[0035] Step S101: Perform semantic-aware cropping on each original image to obtain a corresponding cropped image, and the cropped image contains image features to be annotated;

[0036] Optionally, performing semantic-aware cropping on each original image to obtain a corresponding cropped image, and the cropped image contains image features to be annotated includes:

[0037] The original image is used as the input of the cropping network to calculate the image features of the cropped image and the location information of the image feature points;

[0038] Normalize the image features and output the corresponding heat map;

[0039] Determine whether the brightness of the heat map is greater than the preset activation point threshold. When it is greater than the preset threshold, output a cropped image containing the height, width, and location information of the image feature points.

[0040] In one embodiment, in combination with Table 1, the machine learning model can capture the approximate location information of the object to guide the selection of cropping.

[0041] Table 1

[0042]

[0043]

[0044] In the present invention, the main parameters of the cropping network are pre-set, including: cropping scale s, cropping rate r, response threshold k, adjustment parameter β, etc. The original image is used as the input of the cropping network, and the feature map in the learning process is used to obtain the object frame, as shown in formula (1):

[0045]

[0046] Where M represents the normalization of image features, k is the preset activation point threshold, which can be determined according to the actual situation. It is used to indicate that the function outputs a box when it is judged that the brightness of the heat map is greater than the preset activation point threshold. L is a function for calculating rectangles, and finally returns a positioning box B. After obtaining the positioning box B, considering that the positioning box B is not completely accurate, and in order to cover the background information around some objects, an image feature point is selected in the box, and the position information of the image feature point is calculated, for example: the pixel coordinates (x, y) of any point in the positioning box B; in a special case, the center of the crop is limited to the box, the coordinates of the crop center are (1, 1), and (x, y) is taken as (1, 1). The "center suppression sampling" method is used to reduce the probability of cropping concentrated in the center of the picture, thereby increasing the variance of the sampling, so as to better obtain local semantic features. Specifically, we use Beta distribution as the probability distribution of sampling, and control the distribution to be U-shaped with low in the middle and high around. As shown in formula (2), the final cropping formula is recorded as:

[0047] (x; y; h; w) = Crop (s, r, B) (2)

[0048] It should be noted that only in the case of semantic-aware cropping can the heatmap be detected in the cropped image. The heatmap contains the main image features, while random detection generally cannot directly locate the position of the heatmap. Therefore, it may lead to the situation that some parts of the cropped image contain or do not contain the main image features. For this reason, the image cropping method provided by the present invention can well avoid cropping incorrect samples, reduce the error sample rate, and improve the accuracy and detection efficiency of sample annotation.

[0049] Step S102: Perform semi-supervised learning on the original image and the cropped image, and output the semi-supervised learning loss;

[0050] Optionally, performing semi-supervised learning on the original image and the cropped image to output the semi-supervised learning loss includes:

[0051] Perform softmax prediction on the original image and the cropped image, and respectively output the first prediction result of the original image and the first prediction result of the cropped image;

[0052] Perform KL divergence calculation on the first prediction result of the original image and the first prediction result of the cropped image, and output the semi-supervised learning loss.

[0053] It should be noted that the first prediction result is the prediction result obtained by performing softmax prediction on the original image and the cropped image during the semi-supervised learning process.

[0054] In one embodiment, based on the smoothness assumption and the clustering assumption, data points with different labels are separated in the low-density region, and similar data points have similar outputs. Then, if an actual perturbation is added to an unlabeled original image, its prediction result will not change significantly, that is, the output is consistent. Using samples that match the perturbation and consistency rules to train the model can improve the robustness of the model. Described using KL divergence in semi-supervised learning, the semi-supervised learning loss is obtained. Based on the original image and the cropped image obtained in step S101 as the input of the semi-supervised learning network, after softmax prediction based on the consistency classification loss, the prediction results of the original image and the cropped image are output. The prediction results include the classification type of the image and the corresponding information content. The key point of the consistency classification loss lies in the similarity of the two output distributions. Therefore, in order to avoid the appearance of false positive samples caused by random cropping, semantic-aware cropping is used instead of random cropping to achieve the best effect.

[0055] Assume that the classifier prediction result of the original image is I, and the cropped image obtained by using semantic-aware cropping is I'. The consistency-regularized semi-supervised loss is as follows:

[0056]

[0057] Where, is the semi - supervised loss, f cls (I) is the first prediction result of the original image, f cls (I′) is the first prediction result of the cropped image.

[0058] Step S103: Pass the cropped image corresponding to each original image through active learning and output the active learning loss;

[0059] Optionally, passing the cropped image corresponding to each original image through active learning and outputting the active learning loss includes:

[0060] Perform softmax prediction on the original image and the cropped image, and respectively output the second prediction result of the original image and the second prediction result of the cropped image;

[0061] Perform cross - entropy calculation on the second prediction result of the original image and the second prediction result of the cropped image, and output the active learning loss.

[0062] In one embodiment, in active learning, select the samples that are most inconsistent with the model prediction output as the samples with the largest amount of information and submit them for manual annotation. The most uncertain samples, because the features of these inconsistent samples are very likely to indicate that the machine learning model cannot recognize the local semantic features of the picture. To describe the inconsistency, similar to semi - supervised learning, cross - entropy is used for description. Similarly, since random cropping may lead to the appearance of false - positive samples, semantic - aware cropping is used to replace random cropping to prevent the selection of false - positive samples by active learning, so as to achieve the best effect.

[0063] Suppose the classification prediction output result of the original image is I, and the cropped image obtained by its semantic - aware cropping is I′. The calculation method of the consistency active learning loss is shown in formula (4):

[0064] A sort = cross - entropy(f 分类预测 (I), f 分类预测 (I′)), (4)

[0065] where A sort is the active learning loss, f 分类预测 (I) is the second prediction result of the original image, f 分类预测 (I′) is the second prediction result of the cropped image.

[0066] Optionally, after performing cross - entropy calculation on the second prediction result of the original image and the second prediction result of the cropped image, it further includes:

[0067] Sort the original images according to the calculated cross - entropy, and select a preset number of original images with the largest cross - entropy for annotation;

[0068] Append the labeled images to the initial label sample set to form a new label sample set.

[0069] Sort the calculated active learning losses in descending order according to cross-entropy, and select the preset number of original images with the largest cross-entropies for annotation. The screening method is described by the following formula (5):

[0070]

[0071] where x is the preset number, is the sorting result.

[0072] Form a sample from the x original images with the largest cross-entropies selected, and submit them to humans for annotation. Then add the labeled images to the initial labeled label sample set for training a new model to save the annotation cost.

[0073] Step S104: Obtain the total loss by weighting the active learning loss and the semi-supervised learning loss;

[0074] Step S105: Update the classification weights of the image classification network by backpropagation according to the total loss.

[0075] In one embodiment, the active learning loss and the semi-supervised learning loss are weighted to obtain the total loss, which is used as the loss function of the next image classification network. Update the classification weights of the classification network by backpropagation according to the total loss. For example, before the update, the classification weight of the classification network for animal recognition is 0.5 for cats and 0.5 for dogs. After the update, the classification weight of cats is 0.6 and the classification weight of dogs is 0.4 to improve the classification accuracy of the actual classifier for cats and dogs.

[0076] Optionally, the initial label sample set is used to train the image classification network to determine the initial values of the classification weights.

[0077] It should be noted that before performing semi-supervised active learning, first use semi-supervised learning based on the initial label samples to train the classifier, and determine the initial values of the classification weights of the classification network based on the semi-supervised learning loss.

[0078] Combined with Figure 2 The following is a specific training process of the classification network of the present invention:

[0079] Semi-supervised active learning training is divided into two stages. In the first stage, an initial model is trained using an initial labeled sample set. The second stage is divided into two branches: active learning and semi-supervised learning. In the actual training process, the two branches of active learning and semi-supervised learning are carried out in sequence. In the active learning branch, an active learning strategy is used to calculate the information content of the unlabeled sample set, and unlabeled samples with rich information content are selected for expert annotation. The newly labeled samples are appended to the labeled sample set, and the active learning loss is obtained through the image classification network. In the semi-supervised learning branch, a semi-supervised learning strategy is used to calculate the semi-supervised learning loss of each batch of unlabeled samples. The active learning loss and the semi-supervised learning loss are superimposed to obtain the total model loss, and the image classification network is updated. The iteration is repeated until the image classification network reaches the set classification accuracy threshold or meets the performance requirements.

[0080] In the specific implementation process, it needs to be divided into three stages. In the first stage, an initial model is trained using an initial labeled sample set. In the second stage, an active learning strategy is used to further expand the labeled sample set. In the third stage, a semi-supervised learning strategy is added on the basis of the second stage to calculate the consistency loss and participate in backpropagation. First is the semi-supervised learning stage. A semi-supervised loss is added using consistency regularization. Here, the semi-supervised loss is described by the KL divergence difference of the semantic-aware cropping consistency output. The specific loss is as described in formula (1), and it is used for two-stage backpropagation together with the active learning loss to train the network. Then, the samples queried by the active learning strategy as described in formula (5) are submitted to humans for annotation. The above process is repeated until the expected accuracy or the expected budget annotation cost is reached.

[0081] Based on the above semi-supervised active learning method based on semantic consistency, the classification effects compared with the random cropping method on Cifar-10 are shown in Table II as follows:

[0082] Table II

[0083] It can be seen from Table II that as the number of samples increases, the detection accuracy of the semi-supervised active learning method based on semantic consistency provided by the present invention is higher than that of the random cropping classification method.

[0084]

[0085] The beneficial effect of the present invention is that it organically combines semi-supervised learning and active learning, and uses a novel semantic-aware cropping strategy, which can maximize the image classification performance with the minimum human cost.

[0086] Embodiment 2 of the present invention provides a semi-supervised active learning device based on semantic consistency, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a semi-supervised active learning method based on semantic consistency as described in any one of the above.

[0087] It should be noted that the semi-supervised active learning method that can be implemented by the device in this embodiment is consistent with the implementation manner of the above method embodiment, and will not be elaborated here.

Claims

1. A semi-supervised active learning method based on semantic consistency, characterized in that including performing semantic-aware cropping on each original image to obtain a corresponding cropped image, where the cropped image contains image features to be annotated; performing semi-supervised learning on the original image and the cropped image to output a semi-supervised learning loss; performing active learning on each original image corresponding to the cropped image to output an active learning loss; weighting the active learning loss and the semi-supervised learning loss to obtain a total loss; backpropagating according to the total loss to update the classification weights of the image classification network; performing semantic-aware cropping on each original image to obtain a corresponding cropped image, where the cropped image contains image features to be annotated includes: Take the original image as the input of the cropping network and calculate the image features of the cropped image. The parameters of the cropping network include: cropping scale s , cropping rate r , preset activation point threshold k , height of the cropped image , width of the cropped image ; Normalize the image features and output the corresponding heat map M ; Determine whether the brightness of the heat map is greater than a preset activation point threshold. When it is greater than the preset activation point threshold, obtain the positioning box , L is a function for calculating a rectangle; selecting an image feature point within the bounding box, calculating the position information of the image feature point, and outputting a cropped image including the height, width, and the position information of the image feature point.

2. The semi-supervised active learning method based on semantic consistency according to claim 1, characterized in that, performing semi-supervised learning on the original image and the cropped image to output a semi-supervised learning loss includes: performing softmax prediction on the original image and the cropped image, and respectively outputting a first prediction result of the original image and a first prediction result of the cropped image; calculating the KL divergence between the first prediction result of the original image and the first prediction result of the cropped image, and outputting the semi-supervised learning loss.

3. A semi-supervised active learning method based on semantic consistency according to claim 1, characterized in that, performing active learning on each original image corresponding to the cropped image to output an active learning loss includes: performing softmax prediction on the original image and the cropped image, and respectively outputting a second prediction result of the original image and a second prediction result of the cropped image; calculating the cross entropy between the second prediction result of the original image and the second prediction result of the cropped image, and outputting the active learning loss.

4. A semi-supervised active learning method based on semantic consistency according to claim 3, characterized in that, after calculating the cross entropy between the second prediction result of the original image and the second prediction result of the cropped image, it further includes: sorting the original images according to the calculated cross entropy, and screening out a preset number of original images with the largest cross entropy for annotation; appending the annotated images to the initial label sample set to form a new label sample set.

5. The semi-supervised active learning method based on semantic consistency according to claim 4, characterized in that The initial label sample set is used to train the image classification network to determine the initial value of the classification weights.

6. A semi-supervised active learning device based on semantic consistency, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, it implements a semi-supervised active learning method according to any one of claims 1-5.