Training Method for a General Semi-Supervised Object Detection Framework
By proposing a training method of a general semi-supervised object detection framework in the semi-supervised object detection algorithm, the generality problem and pseudo-label imbalance between different detection models are solved, and more efficient semi-supervised learning performance and better detection accuracy are achieved.
Patent Information
- Application Number
- CN202310549422.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-05-16
AI Technical Summary
The existing semi-supervised object detection algorithms are difficult to generalize among different detection models, and the problem of quality and quantity imbalance of pseudo-labels limits the semi-supervised learning performance of the model.
A training method of a general semi-supervised object detection framework is proposed, by initializing the teacher model and student model, using pseudo-labels and label images for supervised learning, and updating the teacher model through exponential sliding average. At the same time, the lower limit of the number of pseudo labels is set, the noise pseudo labels are filtered, the accuracy and recall rate of pseudo labels are balanced, and the imbalance problem of pseudo labels is alleviated through mixed interpolation and splicing data enhancement technology.
On the premise of ensuring the universality of the framework, the semi-supervised detection performance of most mainstream detection models is improved, the cost of manual labeling is reduced, and the recognition accuracy of objects in the picture is improved.
Smart Images

Figure CN116563634B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to a training method for a general semi-supervised object detection framework. Background Art
[0002] Object detection is an important task in computer vision. Currently, most object detection tasks are based on deep learning algorithms, and the performance depends on the amount of labeled data in model training, such as label information like the position and category of objects. However, the high cost of manual annotation makes it necessary for relevant personnel to balance model performance and annotation cost when conducting experimental research. Therefore, semi-supervised learning algorithms have emerged, aiming to utilize unlabeled data to further improve the object detection performance of the model.
[0003] An existing semi-supervised learning framework for object detection - STAC initializes two detection models with the same structure as the teacher model and the student model respectively. The teacher model is trained in a supervised learning manner on labeled data, and then the teacher model is used to predict pseudo-labels for weakly data-augmented unlabeled images. Through confidence threshold filtering and geometric mapping, the pseudo-labels and the corresponding strongly data-augmented unlabeled images are combined into pseudo-label images. Finally, the student model is trained in a supervised learning manner on the labeled images and the pseudo-label images. Due to the limited scale of the labeled data, the performance of the teacher model limits the quality of the pseudo-labels. Therefore, subsequent work uses the exponential moving average method to update the parameters of the teacher model with the parameters of the student model, so that the teacher model improves as the performance of the student model improves, which is called the detection average teacher framework. At the same time, due to the complexity of the detection task, the differences between detection models in different paradigms are very large. Subsequent semi-supervised object detection algorithms usually make specific improvements based on the detection average teacher framework in combination with specific detection models.
[0004] The prior art usually makes targeted improvements based on the characteristics shown by specific detection models during semi-supervised learning. Although these methods improve the semi-supervised learning performance of the corresponding detection models, it is difficult to be used for detection models in other paradigms. For example:
[0005] (1) Representative semi-supervised object detection algorithms based on Faster R-CNN include Unbiased Teacher, Soft Teacher, and PseCo. Unbiased Teacher changes the classification loss of Faster R-CNN from cross-entropy loss to Focal Loss to alleviate the class imbalance problem of pseudo-labels. However, many single-stage detection models already use Focal Loss as the classification loss; Soft Teacher randomly jitters the pseudo-labeled bounding boxes as region proposals and re-enters them into the R-CNN network, and filters out more accurate pseudo-labels by calculating the consistency of the network output. However, single-stage detection models do not have the R-CNN structure; PseCo enforces the model to learn scale consistency on pseudo-label data through misaligned alignment of the feature pyramid. However, the feature pyramid is not a necessary component for DETR series models.
[0006] (2) The representative semi-supervised object detection algorithm Dense Teacher based on FCOS combines the characteristics of FCOS dense prediction and designs a method of dense pseudo-labels to train the student model, improving the semi-supervised learning performance of single-stage detection models. However, two-stage detection models and DETR series models adopt a sparse prediction paradigm and obviously cannot use dense pseudo-labels for semi-supervised learning.
[0007] (3) The representative semi-supervised object detection algorithm Consistent Teacher based on RetinaNet, considering that the positive and negative sample assignment strategy of RetinaNet is not robust to noisy pseudo-labels, designs an adaptive anchor matching strategy to improve the consistency of positive and negative samples. However, some advanced single-stage detection models and DETR series models have unique designed positive and negative sample assignment strategies, and directly replacing the strategy will have an uncertain impact on the detection performance of the model.
[0008] (4) The representative semi-supervised object detection algorithm Omni-DETR based on Deformable DETR fully explores different paradigms of label guidance for various types based on the detection average teacher framework, but lacks research on the detection average teacher framework itself, resulting in the semi-supervised learning performance of the model not being fully exploited.
[0009] In summary, making targeted improvements to the characteristics exhibited by object detection models during semi-supervised learning will reduce the application scope of the method. Summary of the Invention
[0010] To solve at least some of the above problems in the prior art, the present invention provides a training method for a general semi-supervised object detection framework, including:
[0011] Initialize two detection models as the teacher model and the student model respectively, where the structures and parameters of the teacher model and the student model are the same;
[0012] The teacher model predicts pseudo-labels for the weakly data-augmented unlabeled images. After threshold filtering and geometric mapping, the pseudo-labels and the corresponding strongly data-augmented unlabeled images form pseudo-label images; and
[0013] The student model is trained in a supervised learning manner with the labeled images and the pseudo-label images as inputs. During the training process, the parameters of the student model are optimized by gradient descent, and the updated student model is integrated into the teacher model by exponential moving average.
[0014] Furthermore, set a lower limit on the number of pseudo-labels for each pseudo-label image.
[0015] Furthermore, set that each pseudo-label image includes at least one pseudo-label. If the number of pseudo-labels for a pseudo-label image in a batch is less than 1, the learning weight of the pseudo-label image in this batch is set to zero.
[0016] Furthermore, filter out noisy pseudo-labels by setting a pseudo-label threshold.
[0017] Furthermore, when using the cross-entropy loss function as the classification loss function, the pseudo-label threshold is 0.7.
[0018] Furthermore, when using the Focal Loss function as the classification loss function, the pseudo-label threshold is 0.4.
[0019] Furthermore, when training the student model with the labeled images and the pseudo-label images as inputs, set the ratio and corresponding weights of the labeled images and the unlabeled images in each batch according to the scales of the labeled dataset and the unlabeled dataset to balance the supervision signals of the labels and the pseudo-labels.
[0020] Furthermore, interpolate the pseudo-label images in the current batch and the pseudo-label images in the adjacent batches cached, and take the union of the corresponding pseudo-labels.
[0021] Furthermore, randomly select 4 images from the pseudo-label images in the adjacent batches cached for downsampled mosaic splicing, and take the union of the corresponding pseudo-labels.
[0022] The present invention also provides a computer-readable storage medium, on which a computer program is stored. The computer program, when executed by a processor, executes the steps according to the above method.
[0023] The present invention has at least the following beneficial effects: (1) A training method for a general semi-supervised object detection framework disclosed by the present invention fully exploits the potential performance of the general semi-supervised object detection framework by filtering pseudo-label empty graphs, balancing the accuracy and recall rate of pseudo-labels, and balancing the learning weights of labels and pseudo-labels. On the premise of ensuring the generality of the framework, competitive semi-supervised detection performance is achieved on most mainstream detection models. This method designs a general semi-supervised object detection framework by fully exploiting the potential performance and extended performance of the detection average teacher framework according to the characteristics of the semi-supervised object detection task itself, rather than optimizing a specific module or strategy of the model according to the characteristics shown by a specific detection model during semi-supervised learning. (2) In addition, the problem of the imbalance in the number of positive and negative samples of pseudo-labels is alleviated through pseudo-label interpolation data augmentation, and the problem of the scale imbalance of positive samples of pseudo-labels is alleviated through pseudo-label splicing data augmentation, further improving the detection accuracy of mainstream models of different paradigms on the basis of the general semi-supervised object detection framework. The general semi-supervised object detection framework constructed by the method of the present invention is applied to the field of object detection in computer vision, which can improve the recognition accuracy of objects in pictures and reduce the human annotation cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] To further clarify the above and other advantages and features of the embodiments of the present invention, a more specific description of the embodiments of the present invention will be presented with reference to the accompanying drawings. It can be understood that these drawings only depict typical embodiments of the present invention and thus will not be considered as limiting its scope.
[0025] Figure 1 FIG. shows a schematic diagram of the training process of a general semi-supervised object detection framework according to an embodiment of the present invention; and
[0026] Figure 2 FIG. shows a schematic diagram of the training process of a general semi-supervised object detection framework based on hybrid pseudo-labels according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] It should be noted that the components in the drawings may be exaggerated for illustration purposes and are not necessarily to scale.
[0028] In the present invention, each embodiment is only intended to illustrate the solution of the present invention and should not be construed as restrictive.
[0029] In the present invention, unless otherwise specified, the quantifiers "a" and "one" do not exclude the scenario of multiple elements.
[0030] It should also be noted here that, in the embodiments of the present invention, for the sake of clarity and simplicity, only a part of the components or assemblies may be shown. However, those of ordinary skill in the art can understand that, under the teaching of the present invention, the required components or assemblies can be added according to the specific scenario requirements.
[0031] It should also be noted here that within the scope of the present invention, the terms such as "same", "equal", "equal to" do not mean that the two values are absolutely equal, but allow a certain reasonable error. That is to say, these terms also cover "substantially the same", "substantially equal", "substantially equal to".
[0032] It should also be noted here that in the description of the present invention, the orientation or positional relationships indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than explicitly or implicitly indicating that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as explicitly or implicitly indicating relative importance.
[0033] In addition, the numbering of the steps of each method of the present invention does not limit the execution order of the method steps. Unless otherwise specified, the method steps can be executed in different orders.
[0034] Figure 1 The schematic diagram of the training process of a general semi-supervised object detection framework according to an embodiment of the present invention is shown.
[0035] As Figure 1 shown, a training method for a general semi-supervised object detection framework includes:
[0036] Step 1, initialize two detection models, one as the teacher model and the other as the student model, and their structures and parameters are exactly the same.
[0037] Step 2, the teacher model predicts pseudo-labels for the weakly data-augmented unlabeled images, performs threshold filtering and geometric mapping on the pseudo-labels, and forms pseudo-label images by combining the pseudo-labels with the corresponding strongly data-augmented unlabeled images.
[0038] Step 3, train the student model in a supervised learning manner with the labeled images and pseudo-label images as inputs. During the training process, optimize the parameters of the student model through gradient descent, and integrate the updated student model into the teacher model by means of exponential moving average.
[0039] Since the detection models of different paradigms have different robustness to noise pseudo-labels, the semi-supervised learning performance of some detection models is limited or even cannot converge in the detection of the mean teacher framework. Therefore, three key technologies are proposed to ensure the general effectiveness of the above general semi-supervised object detection framework:
[0040] 1. Filter empty pseudo-label images. At the initial stage of model training, the overall confidence of pseudo-labels is low, and the pseudo-labels of many images will be completely filtered out, resulting in a large number of empty pseudo-label images, seriously deviating from the true distribution. By setting the lower limit of the number of pseudo-labels (pseudo-annotated boxes) for each pseudo-label image, the stability of model training is ensured. It is set that each pseudo-label image includes at least one pseudo-label. If the number of pseudo-labels of a pseudo-label image in a batch is less than 1, the learning weight of the pseudo-label image in this batch is set to zero.
[0041] 2. Balance the accuracy and recall rate of pseudo-labels. The number of detection models is much larger than the number of classification loss functions. Therefore, different detection models may use the same classification loss function. The distribution of classification confidence predicted by different detection models using the same classification loss function is similar. By setting a pseudo-label threshold for a loss function, it is applicable to multiple detection models, ensuring the general effectiveness of the above general semi-supervised object detection framework. By setting a reasonable pseudo-label threshold, most noise pseudo-labels are filtered out, while ensuring the quantity and quality of pseudo-labels, and balancing the accuracy and recall rate of pseudo-labels. In an embodiment of the present invention, when using the cross-entropy loss function as the classification loss function, the pseudo-label threshold is 0.7. In another embodiment of the present invention, when using the Focal Loss function as the classification loss function, the pseudo-label threshold is 0.4.
[0042] 3. Balance the learning weights of labels and pseudo-labels. When training the student model with labeled images and pseudo-label images as inputs, the ratio and corresponding weights of labeled images and unlabeled images in each batch are set in combination with the scales of the labeled dataset and the unlabeled dataset to balance the supervision signals of labels and pseudo-labels.
[0043] Figure 2 FIG. shows a schematic diagram of the training process of a general semi-supervised object detection framework based on hybrid pseudo-labels according to an embodiment of the present invention.
[0044] The inventors found that the key to restricting the performance of semi-supervised object detection is two imbalance problems: the imbalance problem of the number of positive and negative samples of pseudo-labels and the imbalance problem of the scales of positive samples of pseudo-labels. Therefore, a hybrid pseudo-label method is proposed. As Figure 2As shown in the figure, the hybrid pseudo-label method is used in the above general semi-supervised object detection framework. The pseudo-label images and labeled images processed by the hybrid pseudo-label method are used as the input for training the student model, which can further improve the performance of the general semi-supervised object detection framework. Specifically, it includes:
[0045] 1. Hybrid interpolation data augmentation: The number of negative samples in the pseudo-labels is much larger than that of positive samples. Therefore, it is necessary to increase the number of positive samples to solve the problem of imbalance in the number of positive and negative samples in the pseudo-labels. Interpolate the pseudo-label images of the current batch and the neighboring batches cached, and take the union of the corresponding pseudo-labels. Each pseudo-label corresponds to multiple positive samples, and the number of pseudo-labels is doubled on average, alleviating the problem of imbalance in the number of positive and negative samples.
[0046] 2. Hybrid mosaic data augmentation: Randomly select 4 images from the pseudo-label images of the neighboring batches cached for downsampled mosaic splicing, and take the union of the corresponding pseudo-labels. Through the downsampling operation, the number of pseudo-labels of small and medium-sized objects is supplemented, alleviating the problem of scale imbalance of positive samples.
[0047] The general semi-supervised object detection framework obtained by the technical solution of the present invention can be used in the field of object detection in computer vision to achieve the following technical effects: improving the recognition accuracy of objects in images, reducing the manual annotation cost, achieving better detection effects with less labeled data, and further improving the detection effect of the model using more unlabeled data. The principle is to solve the problem that all pseudo-labels will be filtered out at the initial stage of model training, resulting in a large number of empty pseudo-label images, by filtering empty pseudo-label images, select an appropriate pseudo-label threshold, while ensuring the recall rate and accuracy, and combine the scale of the labeled dataset and the unlabeled dataset to set the ratio and corresponding weights of the labeled data and unlabeled data in each batch, balance the supervision signals of the labels and pseudo-labels, and improve the recognition accuracy when the general semi-supervised object detection framework is used for object detection. Then, the number of positive samples is increased by hybrid interpolation data augmentation to solve the problem of imbalance in the number of positive and negative samples in the pseudo-labels, and the number of pseudo-labels of small and medium-sized objects is supplemented by hybrid mosaic data augmentation, alleviating the problem of scale imbalance of positive samples, and further improving the performance of the general semi-supervised object detection framework.
[0048] In order to verify the generality of the general semi-supervised object detection framework constructed by the above method, on the COCO dataset, a representative 10% semi-supervised data division of COCO (using 10% of the training set as labeled data and the remaining data as unlabeled data) is adopted to widely verify most mainstream detection models. The verification results are shown in Table 1. The general semi-supervised object detection framework based on hybrid pseudo-labels has an average improvement of about 9 mAP compared to the supervised baseline on 11 mainstream paradigm detection models.
[0049] Table 1 Comparison of verification results using the COCO dataset divided with 10% semi-supervised data
[0050]
[0051] In addition, semi-supervised training was carried out on the full COCO data (using the training set as labeled data and simultaneously using the unlabeled dataset of COCO) to verify the gain of semi-supervised learning over supervised learning for different detection models using different backbone networks (ResNet50, Swin-L). The verification results are shown in Table 2. The general semi-supervised object detection framework based on hybrid pseudo-labels improves the upper limit of the accuracy of detection models in different paradigms and is superior to previous semi-supervised object detection algorithms (Soft Teacher, PseCo, and Dense Teacher) optimized for specific models. The general semi-supervised object detection framework based on hybrid pseudo-labels is also applicable to the cutting-edge detection model DINO, and there is still an obvious gain when using the stronger backbone network Swin-L. Compared with Soft Teacher based on the HTC model, the architecture established by the method of the present invention is the first semi-supervised object detection algorithm that does not use additional labeled data and has an accuracy exceeding 60 mAP.
[0052] Table 2 Comparison of verification results using the full COCO data
[0053]
[0054] Finally, ablation experiments were carried out on the pseudo-label interpolation data augmentation and pseudo-label splicing data augmentation using the representative Faster R-CNN model to prove the role of each improvement. The verification results are shown in Table 3. The structure of the general semi-supervised object detection framework is simpler than that of Soft Teacher, has better generality (Soft Teacher is only applicable to Faster R-CNN), and higher accuracy. Both pseudo-label interpolation data augmentation and pseudo-label splicing data augmentation can improve the performance of the detection average teacher, and among them, pseudo-label splicing data augmentation more significantly improves the detection ability of small and medium-sized objects of the model.
[0055] Table 3 Comparison of verification results of the baseline model Soft Teacher, the general semi-supervised object detection framework, and the general semi-supervised object detection framework plus pseudo-label interpolation data augmentation and pseudo-label splicing data augmentation
[0056]
[0057] In addition, the embodiments can be provided as a computer program product that may include one or more machine-readable media having machine-executable instructions stored thereon, which when executed by one or more machines such as a computer, a computer network, or other electronic devices, can cause the one or more machines to perform operations in accordance with the embodiments of the present invention. The machine-readable media may include, but are not limited to, floppy disks, optical disks, CD-ROMs (Compact Disc Read-Only Memories), and magneto-optical disks, ROMs (Read-Only Memories), RAMs (Random Access Memories), EPROMs (Erasable Programmable Read-Only Memories), EEPROMs (Electrically Erasable Programmable Read-Only Memories), magnetic or optical cards, flash memories, or other types of media / machine-readable media suitable for storing machine-executable instructions.
[0058] In addition, the embodiments can be downloaded as a computer program product, wherein the program can be transmitted from a remote computer (e.g., a server) to a requesting computer (e.g., a client) by one or more data signals implemented and / or modulated by a carrier wave or other propagation medium via a communication link (e.g., a modem and / or a network connection). Thus, the machine-readable media used herein may include such a carrier wave, but this is not necessary.
[0059] Although some embodiments of the present invention have been described in this application document, those skilled in the art can understand that these embodiments are merely shown as examples. Those skilled in the art can conceive of numerous variations, alternatives, and improvements under the teachings of the present invention without departing from the scope of the present invention. The appended claims are intended to define the scope of the present invention and thereby cover the methods and structures within the scope of these claims themselves and their equivalents.
Claims
1. A training method for a general semi-supervised object detection framework, characterized in that Including: Initialize two detection models as the teacher model and the student model respectively, where the structures and parameters of the teacher model and the student model are the same; The teacher model predicts pseudo-labels for the weakly data-augmented unlabeled images, performs threshold filtering and geometric mapping on the pseudo-labels, and combines the pseudo-labels with the corresponding strongly data-augmented unlabeled images to form pseudo-label images; And Train the student model in a supervised learning manner with the labeled images and the pseudo-label images as inputs. During the training process, optimize the parameters of the student model through gradient descent, and integrate the updated student model into the teacher model through exponential moving average; Set the lower limit of the number of pseudo-labels for each pseudo-label image: Set that each pseudo-label image includes at least one pseudo-label. If the number of pseudo-labels of a pseudo-label image in a batch is less than 1, the learning weight of the pseudo-label image in this batch is set to zero; When training the student model with the labeled images and the pseudo-label images as inputs, set the ratio and corresponding weights of the labeled images and the unlabeled images in each batch in combination with the scales of the labeled dataset and the unlabeled dataset to balance the supervision signals of the labels and the pseudo-labels; Interpolate the pseudo-label images in the current batch and the cached neighboring batches of pseudo-label images, and take the union of the corresponding pseudo-labels; Randomly select 4 images from the cached neighboring batches of pseudo-label images for downsampled mosaic stitching, and take the union of the corresponding pseudo-labels.
2. The training method of the general semi-supervised object detection framework according to claim 1, characterized in that Filter out the noisy pseudo-labels by setting the pseudo-label threshold.
3. The training method of the general semi-supervised object detection framework according to claim 2, characterized in that When using the cross-entropy loss function as the classification loss function, the pseudo-label threshold is 0.
7.
4. The training method of the general semi-supervised object detection framework according to claim 2, characterized in that When using the Focal Loss function as the classification loss function, the pseudo-label threshold is 0.
4.
5. A computer-readable storage medium, on which a computer program is stored, and the computer program, when executed by a processor, executes the steps of the method according to any one of claims 1-4.
Citation Information
Patent Citations
Semi-supervised learning-based image segmentation model training method and segmentation method
CN114255237A
Semi-supervised learning power equipment target detection model training method, detection method and device
CN116091858A