Self-training weak supervision object detection method with uncertainty perception

The self-training framework with precise sampling and weighted localization refinement addresses uncertainties in weakly supervised object detection, enhancing model adaptability and accuracy.

CN120318496APending Publication Date: 2025-07-15HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510469987.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing weakly supervised object detection methods are unable to adapt to different scenes and image changes due to pseudo-label uncertainty and bounding box regression uncertainty.

Method used

Adopting an uncertainty-aware self-training framework, by constructing teacher modules and student modules, generating pseudo-labels using accurate positive and negative sampling strategies, introducing soft pseudo-labels and weighted positioning refine branches, optimizing the pseudo-labels and bounding box regression uncertainty of the model, designing loss functions to maximize likelihood losses, and balancing classification and positioning tasks.

Benefits of technology

It effectively reduces the uncertainty of pseudo-label and bounding box regression, improves the model's adaptability to different scenes and image changes, and improves the accuracy and generalization ability of object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318496A_ABST
    Figure CN120318496A_ABST
Patent Text Reader

Abstract

The invention discloses a self-training weak supervision object detection method based on uncertainty perception, solves the problem that the detection accuracy is greatly influenced by the uncertainty of an existing weak supervision object detection method, and belongs to the field of computer vision and deep learning. The method comprises the following steps: constructing a training framework which comprises a pre-training classifier, a full connection layer and a self-training module; an image enters a pre-training classifier, and a full connection layer generates a group of candidate frame feature vectors and inputs the candidate frame feature vectors to a self-training module; the self-training module comprises parallel K + 1 layers of the teacher module and the student module, and the K + 1 layer is a weighted positioning refinement branch; during training, a first-layer pseudo label of the student module is obtained by adopting a positive and negative sampling strategy according to a prediction result of the teacher module; in K + 1 layers of the student module, a false label of the next layer is obtained by adopting a positive and negative sampling strategy according to the prediction result of the previous layer; and in the reasoning process, finally outputting a prediction result from the (K + 1) th layer. According to the method, the influence of uncertainty is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an uncertainty-aware self-training weakly supervised object detection method, belonging to the fields of computer vision and deep learning. Background Art

[0002] Object detection is a classical problem in computer vision and has been significantly developed in recent years with the emergence of large-scale accurately annotated datasets. However, obtaining accurately labeled bounding box data is both time-consuming and expensive, and sometimes even infeasible. These problems have promoted the research on weakly supervised object detection, which only uses image-level annotations to train deep neural networks. Weakly supervised object detection models need to be trained on limited weakly labeled data and at the same time have good generalization ability to accurately detect targets on unseen test data. Weakly supervised object detection methods usually regard the task as a classification task or a latent variable learning task, and by designing appropriate classification loss functions or complex loss functions, train the detector to find the positive candidate boxes as real instances among numerous candidate boxes. However, due to the limitations of weakly labeled data and the risk of overfitting in the model training process, the model may perform well on the training data, but its performance drops on the test data and cannot well adapt to different scenarios and image variations, thus leading to the problem of converging to a local optimal solution in the training stage and highlighting the most discriminative regions of the object.

[0003] Based on the above, the existing weakly supervised object detection methods highly rely on the quality of the pseudo-labels generated during the model. The limitations of its model can be summarized as the uncertainty of pseudo-labels. During the training process, one is the uncertainty of pseudo-labels, and the other is the uncertainty of bounding box regression caused by the imbalance between the classification task and the localization task. It is precisely these uncertainties that greatly affect the accuracy of the detection model. Summary of the Invention

[0004] Aiming at the problem that the uncertainty of the existing weakly supervised object detection methods greatly affects the detection accuracy, the present invention provides an uncertainty-aware self-training weakly supervised object detection method.

[0005] An uncertainty-aware self-training weakly supervised object detection method of the present invention includes:

[0006] Construct a training framework, including a pre-trained classifier, a fully connected layer, and a self-training module;

[0007] According to the input image, obtain multiple candidate boxes, input the input image and the candidate boxes into the pre-trained classifier to obtain candidate box feature blocks, and the candidate box feature blocks generate a set of candidate box feature vectors through two fully connected layers;

[0008] The self-training module includes a teacher module and a student module, and the student module includes K + 1 parallel layers;

[0009] During training, each candidate box feature vector is simultaneously input into the K+1 layer of both the teacher module and the student module. The label of the teacher module is the image-level label, and the pseudo-label of the first layer of the student module is obtained from the prediction result of the teacher module; in the K+1 layer of the student module, the pseudo-label of the subsequent layer is obtained from the prediction result of the previous layer.

[0010] During the inference process, the final output is the prediction result from the K+1 layer.

[0011] Preferably, the pseudo-label is a soft pseudo-label;

[0012] The method for obtaining the soft pseudo-label:

[0013] The instance-level pseudo-label of the k-th layer of the student module is generated by combining the prediction result of the teacher module and the prediction result of the k-1 layer of the student module with the positive and negative sampling method k = 1, 2,..., K+1; in the positive and negative sampling method, the threshold for negative sample selection is set such that the IOU overlap degree is in the range of [0.1, 0.5). Samples with an IOU overlap degree lower than 0.1 are discarded, and the others are positive samples;

[0014] Soft pseudo-label is:

[0015]

[0016] where is the loss weight. When k = 1, is equal to the classification confidence score of the candidate box with the highest score in the prediction result of the teacher module When k > 1, is equal to the classification confidence score of the candidate box with the highest score in the k-1 layer of the prediction result of the student module

[0017] Preferably, the K+1 layer is a weighted localization refinement branch, including a classification branch and a localization branch. Both the classification branch and the localization branch are fully connected layers and softmax layers connected in sequence, and the classification branch and the localization branch output classification information and position information respectively.

[0018] Preferably, the loss function Loss box is:

[0019]

[0020] where σ1 is the observation noise parameter of the localization branch, Loss cls represents the multi-label cross-entropy loss of the classification branch, σ2 is the observation noise parameter of the classification branch, and Loss loc represents the loss of the localization branch;

[0021] Loss cls 、Loss loc is as follows:

[0022]

[0023] Among them, X c,r and B c,r are the classification confidence score and bounding box of the candidate box r predicted by the (K + 1)-th layer for the category c respectively. L smooth-L1 represents the smooth L1 loss for localization, and λ is the weight used to balance the classification and localization losses of the (K + 1)-th layer; R represents the set of candidate boxes, and |R| represents the total number of candidate boxes. is the soft pseudo-label of the (K + 1)-th layer corresponding pseudo-box.

[0024] Preferably, the teacher module includes two branches. Each branch is a fully connected layer and a softmax layer connected in sequence. The fully connected layers in the two branches do not share parameters, and the outputs of the softmax layers in the two branches are obtained through element-wise multiplication operation to get the prediction result of the teacher module.

[0025] Preferably, each of the 1st layer to the K-th layer includes a fully connected layer and a softmax layer connected in sequence.

[0026] Preferably, the loss function Loss basic of the training framework is as follows:

[0027] Loss basic = Loss w + Loss r + Loss box

[0028] The loss function Loss w of the teacher module is as follows:

[0029]

[0030] Among them, c is the category, C is the set of image categories, y c = 1 or y c = 0 indicates whether the category c exists, and τ(c) represents the prediction score of the category c;

[0031] The loss function of the 1st layer to the K-th layer of the student module is as follows:

[0032]

[0033] Among them, represents the classification confidence score corresponding to the candidate box r predicted by the k-th layer;

[0034] Total loss function Loss of the first K layers of the student module r is:

[0035]

[0036] where γ r is used to balance the weights of the first layer to the Kth layer of the student module.

[0037] Preferably, the pre-trained classifier is an ImageNet classifier.

[0038] Advantages of the present invention: The present invention provides an uncertainty-aware training framework to reduce the impact of these uncertainties. Specifically, an accurate positive and negative sampling strategy is adopted to generate pseudo-labels to solve the problem of pseudo-label uncertainty. The IOU overlap threshold for negative sample selection is set in the range of [0.1, 0.5). At the same time, soft pseudo-labels in the range of (0, 1) are used to supervise the student module instead of hard labels. Then, the last layer of the student module is designed based on Bayesian uncertainty modeling, and weighted localization refines the last layer to overcome the uncertainty of bounding box output. In the localization task, the error between the prediction and the true target follows a Gaussian distribution. For the classification task, the classification error follows a Boltzmann distribution. The bounding box branch includes a joint task of classification and localization, and the weighted total loss function is determined by maximizing the log-likelihood of these two errors. The present invention solves the problem of prominent discriminant regions in weakly supervised object detection methods, overcomes the pseudo-label uncertainty and bounding box regression uncertainty in the self-training framework, enables the model to adapt to different scenarios and image variations, breaks through the limitation of the lack of regression ability in weakly supervised object detectors, and promotes the implementation of object detection technology in artificial intelligence deep learning to a certain extent.

[0039] In the present invention, the VOC2007 and VOC2012 datasets are used to train the model. VOC2007 contains 9963 images and there are 20 classes of objects to be detected. There are 5011 training sets in the dataset to train the model proposed by the present invention and 4952 test images to evaluate the accuracy of the model. VOC2012 contains 22531 images and there are 20 classes of objects to be detected. There are 11540 training sets in the dataset to train the model proposed by the present invention and 10991 test images to evaluate the accuracy of the model. The experimental results follow the evaluation metrics given by the official, that is, the accuracy mAP of the model is evaluated on the test set, and the correct localization performance CorLoc of the model is evaluated on the training set. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 Shows the heatmap evolution of the target object at different iteration times;

[0041] Figure 2 This is the overall architecture diagram of the present invention;

[0042] Figure 3 This is the comparison of the qualitative detection results between the method of the present invention and the baseline method. Specific Embodiments

[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the scope of protection of the present invention.

[0044] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0045] Next, the present invention will be further described in conjunction with the accompanying drawings and specific embodiments, but it is not limited to the present invention.

[0046] In view of the problem that the existing weakly supervised object detection method is prone to converge to a local optimal solution, this embodiment analyzes the noise sources in the training and testing processes, including pseudo-label uncertainty and bounding box regression uncertainty. An uncertainty-aware self-training framework for an online teacher-student block is designed to eliminate the uncertainty of weakly supervised object detection during the training process. Based on self-paced learning, this embodiment focuses on gradually optimizing all image pseudo-labels, enabling the network model to gradually obtain more accurate localization. The uncertainty-aware self-training weakly supervised object detection method in this embodiment includes:

[0047] Construct a training framework, including a pre-trained classifier, a fully connected layer, and a self-training module;

[0048] According to the input image, obtain multiple candidate boxes, input the input image and the candidate boxes into the pre-trained classifier to obtain candidate box feature blocks, and the candidate box feature blocks generate a group of candidate box feature vectors through two fully connected layers;

[0049] The self-training module includes a teacher module and a student module, and the student module includes K + 1 parallel layers;

[0050] During training, each candidate box feature vector is simultaneously input into the K + 1 layers of the teacher module and the student module. The label of the teacher module is the image-level label, and the pseudo-label of the first layer of the student module is obtained from the prediction result of the teacher module; the pseudo-label of the subsequent layer in the K + 1 layers of the student module is obtained from the prediction result of the previous layer;

[0051] During the inference process, the final output comes from the prediction result of the K + 1 layer.

[0052] Given an image I and its image-level label y = [y1, y2, …, y c ∈ R C×1 , where C represents the number of image categories, and y c = 1 or y c = 0 indicates whether the category c exists. Candidate boxes are generated through the selective search method where |R| is the total number of candidate boxes. The overall architecture of the method proposed in this embodiment is as Figure 2 shown. In this embodiment, weakly supervised object detection is regarded as a self-training learning task. Two types of uncertainties generated during the training process are analyzed, and an attempt is made to optimize the model from the perspective of uncertainty. Pseudo-labels are generated by combining the output of the teacher module with the positive and negative sampling method and are used to guide the online training of the student module. The student module has K + 1 layers, and the pseudo-labels generated by the previous layer of the student module can be used to supervise the next layer of the student module. During the self-training process, the pseudo-labels are continuously updated and optimized in the step-by-step iteration of the student module, thereby improving the representation ability of the entire module.

[0053] Construct a teacher module. The teacher module includes two branches, each branch is a fully connected layer and a softmax layer connected in sequence. The fully connected layers in the two branches do not share parameters, and the outputs of the softmax layers in the two branches are obtained through element-wise multiplication operation to get the prediction result of the teacher module. Specifically:

[0054] Given an image I and the corresponding candidate box P, each candidate box feature is learned from a pre-trained classifier, and the pre-trained classifier can adopt the ImageNet classifier, including a feature extraction layer, a pooling layer, and then project the learned features to two different score matrices through two fully connected layers that do not share parameters, namely x cls and x det ∈ R C×∣R∣ , which are used for category c and candidate box r in the image respectively. The two matrices are normalized in different directions through the softmax layer, that is, the category direction and the candidate box direction, and are interpreted as performing classification and localization:

[0055]

[0056] Then, the score of each candidate box r for the specified category c is generated through the element-wise multiplication operation x R = σ(x cls ) ⊙ σ(x det ). Finally, the predicted score of category c is obtained by summing over all candidate boxes The image-level loss function Loss w of the teacher module is defined as follows:

[0057]

[0058] Positive and negative sampling strategy of this embodiment: Introduce K + 1 online branches to form a student module to refine the scores of candidate boxes. The instance-level pseudo-label of each candidate box is generated by combining the scores of candidate boxes in the previous step with the positive and negative sampling method. Specifically, first select the candidate box with the highest score for each category; then, regard the candidate boxes that highly overlap with the candidate box with the highest score in space as positive instances, and select the candidate boxes with low overlap as negative instances. Note that in the past methods, the IOU overlap degree of 0.5 was used as the threshold for distinguishing positive and negative samples. To eliminate this uncertainty, this embodiment sets the threshold for negative sample selection to an IOU overlap degree in the range of [0.1, 0.5), which means that all samples with scores falling within this range are classified as negative samples, and samples with an IOU overlap degree lower than 0.1 are discarded. At the same time, use soft pseudo-labels to supervise each layer of the student module in the range of (0, 1), rather than hard labels. Soft pseudo-labels are defined as follows:

[0059]

[0060] where is the loss weight. When k = 1, is equal to the classification confidence score of the candidate box with the highest score in the prediction result of the teacher module When k > 1, is equal to the classification confidence score of the candidate box with the highest score in the (k - 1)-th layer of the prediction result of the student module

[0061] In this embodiment, from the first layer to the K-th layer, it all includes a fully connected layer and a softmax layer connected in sequence. The loss function of the first layer to the K-th layer of the student module is defined as follows:

[0062]

[0063] represents the classification confidence score corresponding to the candidate box r predicted by the k-th layer;

[0064] The total loss function of the first K layers of the student module is defined as follows:

[0065]

[0066] where λ r is used to balance the weights of the first to K-th layers of the student module.

[0067] In addition, a weighted localization refinement branch is added at the end of the student module to further improve the accuracy of the student module.

[0068] Weighted Localization Refinement Branch: To reduce the gap between the classification and localization tasks, a weighted localization refinement branch is introduced as the last layer of the student module and the output layer of the self-training framework. The introduced weighted localization refinement branch performs two tasks simultaneously, namely the classification task and the localization task. Each task has accidental uncertainties in task perception due to limitations in data and model understanding. Since weakly supervised object detection lacks instance-level annotations, it is necessary to balance the classification and localization tasks. Therefore, based on Bayesian deep learning, the uncertainty of the task is analyzed to define the weighted sum of the losses, so that the losses of each task have a similar scale. The weighted localization refinement branch consists of two streams, a classification branch and a localization branch, and is supervised by the pseudo-ground truth of the K-th layer of the student module. Specifically, the pseudo-ground truth refers to the pseudo-labels i ∈P p and the corresponding candidate boxes from the positive samples p generated in the previous step. Its formula is as follows:

[0069]

[0070] rectangle=(top,left,bottom,right0,category=PNS(score0

[0071] where, represents the positive sample corresponding to a rectangular candidate box, represents the coordinates of the corresponding candidate box of this positive sample, represents the pseudo-label category of the corresponding candidate box of this positive sample; category represents the category, rectangle represents the coordinates of the candidate box, w i represents the number of positive samples, w j represents the number of image categories, |P p | represents the total number of positive samples, C represents the total number of image categories. "PNS" represents the positive and negative sampling strategy, and the score score is the confidence score of the candidate box. The classification branch and the localization branch are supervised by the pseudo-labels and the corresponding candidate boxes respectively. The loss of the weighted localization refinement branch can be expressed as:

[0072]

[0073] where, Loss cls represents the multi-label cross-entropy loss of the classification branch, and Loss loc represents the smooth L1 loss of the localization branch. X c,r and B c,rThey are respectively the classification confidence score and the bounding box of the candidate box r predicted by the (K + 1)-th layer for the category c. L smooth-L1 represents the smooth L1 loss for localization, λ is the weight used to balance the classification and localization losses of the (K + 1)-th layer; R represents the set of candidate boxes. Then, the weighted loss function of the bounding box branch is analyzed, and the uncertainties in the classification and localization tasks are used to weight the loss function of the bounding box branch. For the localization task, the error between the model prediction and the true target follows a Gaussian distribution, and the Gaussian likelihood can be defined as:

[0074]

[0075] N represents the normal distribution, and f W (x) is the sufficient statistic, and σ1 is the observation noise parameter of the localization branch;

[0076] The corresponding log-likelihood can be written as:

[0077]

[0078] where y loc is the ideal output of the localization branch. For the classification task, the classification error follows a Boltzmann distribution, and the likelihood probability of the class label is defined as:

[0079]

[0080] After taking the logarithm and expanding, it can be further expressed as:

[0081]

[0082] where y cls is the ideal output of the classification branch, σ2 is the observation noise parameter of the classification branch, is the c-th element of the vector f W (x), log∑ c′ exp() represents the normalization process, and c′ is used to iterate over all categories. For the joint task of classification and localization, the optimization objective is to maximize the log-likelihood of the model, that is, to minimize the negative log-likelihood, as follows:

[0083]

[0084] where L(W, σ1, σ2) represents the loss function of the joint task of classification and localization, W represents the weight matrix, L1(W) represents the loss function of the classification task in the joint task, and L2(W) represents the loss function of the localization task in the joint task;

[0085] Therefore, the complete weighted localization refinement branch loss can be expressed as:

[0086]

[0087] Among them, σ1 and σ2 are regarded as the learnable parameters of the two branches in the (K + 1)-th layer. Finally, by combining the above three loss functions Loss w 、Loss r and Loss box to train the entire uncertainty-aware framework as follows:

[0088] Loss basic = Loss w + Loss r + Loss box

[0089] Figure 1 shows the heatmaps of the target object at different iteration times. Obviously, as the number of iterations increases, the network model gradually obtains more accurate localization. Self-paced learning, as a robust learning strategy, can also alleviate the uncertainty caused by noisy samples in some computer tasks. The self-paced curriculum learning strategy distinguishes the knowledge of confident instances from simple images in the early learning stage, and learns more ambiguous instances from more complex images in the later learning stage. In this embodiment, instead of solving the uncertainty problem of weakly supervised object detection by first learning simple samples and then learning difficult samples, it focuses on gradually optimizing the pseudo-labels of all images.

[0090] Prepare training samples. The present invention selects the VOC2007 and VOC2012 datasets to verify the effectiveness of the present invention. VOC2007 contains 9963 images and there are 20 types of objects to be detected. There are 5011 training sets in the dataset to train the model proposed by the present invention and 4952 test images to evaluate the accuracy of the model. VOC2012 contains 22531 images and there are 20 types of objects to be detected. There are 11540 training sets in the dataset to train the model proposed by the present invention and 10991 test images to evaluate the accuracy of the model. The experimental results follow the evaluation metrics given officially, that is, evaluate the accuracy mAP of the model on the test set and evaluate the correct localization performance CorLoc of the model on the training set.

[0091] The present invention uses the VGG16 classification network pre-trained on the ImageNet dataset as the backbone network of the framework. About 2000 candidate boxes are generated on each image by the selective search method. In the basic multi-instance detector, the refinement branch is trained 3 times. In the uncertainty-aware self-training framework, the student module has 4 branches, including a weighted localization refinement branch. For the first three branches of the student module, the present invention sets the weight coefficient of each student layer to λ r= [3, 1, 1]. During the training process, the entire model was trained for 75k iterations, with the learning rate changing from 0.001 in the first 30k iterations to 0.0001 in the last 45k iterations. In the present invention, the mini-batch size was set to 2, the momentum was set to 0.9, and the weight decay was set to 0.0005. For training and test data augmentation, the shortest side of the image was adjusted to one of six scales {480, 576, 688, 864, 1000, and 1200}, the longest side was set to not exceed = 2000, and each image was horizontally flipped. The present invention uses the deep learning framework PyTorch and NVIDIA RTX 3090 GPU to train the entire model.

[0092] Experiments have proven that the present invention alleviates the problem of prominent discriminative regions in the weak-supervised object detection process, overcomes the uncertainties generated by the model in the training framework, including pseudo-label uncertainties, and the uncertainties caused by classification and localization imbalances. This method enables the model to learn more accurate and sufficient target feature representations and achieve better experimental results.

[0093] The effectiveness of the method proposed in the present invention was trained and tested on the general VOC2007 and VOC2012 datasets respectively to verify the effectiveness of the method proposed in the present invention. In addition, the visualization results on the VOC2007 and VOC2012 test sets were also shown. Ablation experiments were conducted simultaneously to explore the impact of each component in the weak-supervised framework on the quantitative and qualitative results.

[0094] Table 1 Effectiveness analysis of the uncertainty self-training framework

[0095] Method mAP (%) Baseline model 47.4 Baseline model + Positive and negative sample sampling strategy 49.9 Baseline model + Positive and negative sample sampling strategy + Unweighted localization refinement branch 53.5 Baseline model + Positive and negative sample sampling strategy + Weighted localization refinement branch 54.3

[0096] Table 2 Ablation study of the negative sample selection threshold

[0097] Threshold mAP (%) 0 47.4 0.1 49.9 0.2 47.9

[0098] Table 3 Accuracy comparison on the VOC2007 test set

[0099]

[0100]

[0101] Table 4 Correct localization performance comparison on the VOC2012 training set

[0102]

[0103] Table 5 Correct localization performance comparison on the VOC2007 and VOC2012 training sets

[0104]

[0105]

[0106] Impact of Uncertainty-Aware Self-Training: Table 1 reports the ablation results of the uncertainty-aware self-training framework. The first row is the re-trained OICR (i.e., the baseline). The second row is the self-training structure with an exact positive and negative sampling strategy for optimizing the representation of the student module. Compared with the baseline, the result is improved by 2.5%, which indicates that the self-training framework effectively eliminates the impact of pseudo-label uncertainty. The third row is the result using an unweighted localization refinement branch, whose loss consists of the losses of two fully-connected layers. The fourth row is the result of the weighted localization refinement branch, which is optimized by the weighted sum of the classification and localization losses. The weighted localization refinement branch has better performance, with the mAP increasing from 53.5% to 54.3%, which further proves the effectiveness of introducing dynamic balance parameters. The weighted localization refinement branch is used in all subsequent experiments unless otherwise specified. Compared with the first row and the fourth row, the entire uncertainty-aware self-training framework improves the mAP performance by 6.9%, which indicates that the uncertainty of the model is a significant bottleneck in weakly-supervised object detection, and the model is proven to be effective in reducing network uncertainty. In addition, in all experiments, soft labels are used to supervise each layer of the student module.

[0107] Impact of Negative Sample Selection Threshold: Table 2 reports the impact of the negative sample selection threshold. The first column represents the lower bound of the IoU overlap range for negative sample selection. For example, "0.2" means the threshold is set in the range of [0.2, 0.5). The first row is the re-trained OICR, with negative samples selected in the range of [0, 0.5). The second row indicates that the threshold is set in the range of [0.1, 0.5). The method of the present invention improves the mAP by 2.5% compared with the baseline OICR (47.4% vs. 49.9%). In addition, compared with the second row and the third row, the model significantly improves the detection accuracy from 47.9% to 49.9%. Therefore, the IoU overlap range for negative sample selection is set to [0.1, 0.5).

[0108] Comparison with other baseline models and visualization results: Experiments were conducted on the PASCAL VOC0712 dataset, and the method proposed in the present invention was compared with other baseline methods for weakly supervised object detection. The results are shown in Tables 3, 4, and 5. Table 3 reports the detection performance of various methods on the VOC2007 test set. In terms of the accuracy mAP metric, the method of the present invention is 14.0%, 7.9%, 7.5%, and 1.3% higher than OICR, MELM, CBASH, and SLV respectively, and achieves the highest performance (55.2%). Compared with the self-paced method that can also alleviate the influence of uncertainty in weakly supervised object detection, the method of the present invention is 9.3% and 7.6% higher respectively. Obviously, the single model without any post-processing not only outperforms the single model but also greatly exceeds the ensemble model or the supervised re-training model. In addition, the self-training module proposed in the present invention performs well on both rigid and non-rigid objects. Especially for "bird", "chair", and "person", the accuracies of 61.3%, 34.6%, and 27.9% are obtained respectively. Table 4 shows the performance of different methods on the VOC2012 test dataset. The proposed single model is about 2% higher than other advanced weakly supervised detection methods and is only 0.1% lower than the method of Ren et al. The improvement obtained on the VOC2007 test dataset further verifies the effectiveness of the proposed method. Table 5 shows the localization accuracy results obtained on the VOC2007 and 2012 training validation datasets. The method achieves an average localization accuracy of 70.9% on the VOC2007 training validation dataset, far leading other advanced methods. For example, the method of the present invention exceeds methods such as OICR, ZigZag, C-SPCL, and CBASH by 10.3%, 9.7%, 8.5%, and 5.1% respectively in terms of accuracy mAP. Compared with the SLV method, this method is only 0.1% lower. On the VOC2012 training validation dataset, this method leads other advanced baseline models, and the localization performance Corloc reaches 70.9%.

[0109] The method proposed in the present invention has significant advantages over other methods, which benefits from the proposed uncertainty-aware module for improving the feature distribution and reducing the noise generated by the model. Compared with other methods of the self-training framework, the advantage lies in not only analyzing the uncertainty of the model but also taking corresponding measures to eliminate them. Some detection results are as Figure 3 shown. Obviously, the model is more inclined to focus on larger and more accurate object regions than the baseline model. Compared with the baseline results, more accurate bounding boxes are generated, which depends on the ability of the uncertainty-aware self-training learning framework to eliminate uncertainty.

[0110] Although the present invention has been described herein with reference to particular embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Accordingly, it should be understood that numerous modifications may be made to the exemplary embodiments, and other arrangements may be devised, without departing from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the different dependent claims and the features described herein may be combined in ways different from those described in the original claims. It should also be understood that the features described in connection with separate embodiments may be used in other described embodiments.

Claims

1. An uncertainty-aware self-training weakly supervised object detection method, characterized in that, Including: Construct a training framework, including a pre-trained classifier, a fully connected layer, and a self-training module; According to the input image, obtain multiple candidate boxes, input the input image and the candidate boxes into the pre-trained classifier to obtain candidate box feature blocks, and the candidate box feature blocks generate a group of candidate box feature vectors through two fully connected layers; The self-training module includes a teacher module and a student module, and the student module includes K + 1 parallel layers; During training, each candidate box feature vector is simultaneously input into the K + 1 layers of the teacher module and the student module. The label of the teacher module is the image-level label, and the pseudo-label of the first layer of the student module is obtained from the prediction result of the teacher module; in the K + 1 layers of the student module, the pseudo-label of the next layer is obtained from the prediction result of the previous layer; During the inference process, the final output is the prediction result from the K + 1 layer.

2. The uncertainty-aware self-training weak supervision object detection method according to claim 1, wherein The pseudo-label is a soft pseudo-label; The method for obtaining the soft pseudo-label: Generate the instance-level pseudo-labels of the k-th layer of the student module by combining the prediction results of the teacher module and the prediction results of the (k-1)-th layer of the student module with the positive and negative sampling method In the positive and negative sampling method, the threshold for negative sample selection is set such that the IOU overlap is in the range of [0.1, 0.5). Samples with an IOU overlap lower than 0.1 are discarded, and the others are positive samples; Soft pseudo-label is as follows: Among them, is the loss weight. When k = 1, is equal to the classification confidence score of the candidate box with the highest score in the prediction result of the teacher module When k > 1, is equal to the classification confidence score of the candidate box with the highest score in the (k - 1)-th layer of the prediction result of the student module 3. The uncertainty-aware self-training weak supervision object detection method according to claim 2, wherein The K + 1 layer is a weighted localization refinement branch, including a classification branch and a localization branch. Both the classification branch and the localization branch are fully connected layers and softmax layers connected in sequence, and the classification branch and the localization branch output classification information and position information respectively.

4. The uncertainty-aware self-training weakly supervised object detection method according to claim 3, wherein The loss function Loss of the (K + 1)-th layer box is as follows: Among them, σ1 is the observation noise parameter of the positioning branch, and Loss cls represents the multi-label cross-entropy loss of the classification branch, σ2 is the observation noise parameter of the classification branch, and Loss loc represents the loss of the positioning branch; Loss cls 、Loss loc is as follows: Among them, X c,r and B c,r are the classification confidence score and the bounding box of the candidate box r predicted by the (K + 1)-th layer for the category c, respectively. L smooth-L1 represents the smooth L1 loss for localization, and λ is the weight used to balance the classification and localization losses of the (K + 1)-th layer; R represents the set of candidate boxes, |R| represents the total number of candidate boxes, is the soft pseudo-label of the (K + 1)-th layer corresponding pseudo-box.

5. The uncertainty-aware self-training weakly supervised object detection method according to claim 1, characterized in that The teacher module includes two branches, each branch is a fully connected layer and a softmax layer connected in sequence. The fully connected layers in the two branches do not share parameters, and the outputs of the softmax layers in the two branches are obtained through an element-wise multiplication operation to get the prediction result of the teacher module.

6. The uncertainty-aware self-training weakly supervised object detection method according to claim 1, wherein Each of the first layer to the K layer includes a fully connected layer and a softmax layer connected in sequence.

7. The uncertainty-aware self-training weakly supervised object detection method according to claim 4, wherein The loss function Loss of the training framework basic is as follows: Loss basic = Loss w + Loss r + Loss box Loss function of the teacher module w is as follows: where c is the category, C is the set of image categories, and y c = 1 or y c = 0 indicates whether the category c exists, and τ(c) represents the predicted score of the category c; Loss function from the first layer to the Kth layer of the student module is as follows: Among them, represents the classification confidence score corresponding to the candidate box r predicted by the k-th layer; Total loss function Loss of the first K layers of the student module r is as follows: Among them, λ r is used to balance the weights from the first layer to the Kth layer of the student module.

8. The uncertainty-aware self-training weakly-supervised object detection method according to claim 1, characterized in that, The pre-trained classifier is an ImageNet classifier.

9. An uncertainty-aware self-training weakly supervised object detection device, comprising a storage device, a processor, and a computer program stored in the storage device and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the uncertainty-aware self-training weakly supervised object detection method as described in any one of claims 1 to 8.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the uncertainty-aware self-training weakly supervised object detection as described in any one of claims 1 to 8.