Active learning initialization method guided by visual language large model and weak supervision object detector
Through the combined method of visual language big model and weakly supervised object detector, the problem of high marking cost in the initialization phase and low recall of detection results in the active learning object detection technology is solved, and efficient active learning initialization and training of fully supervised object detectors are achieved.
Patent Information
- Application Number
- CN202411904099.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-06
AI Technical Summary
The existing active learning object detection technology is expensive to label in the initialization stage, and the detection result recall rate of visual language big models is low.
The active learning initialization method guided by visual language big model and weakly supervised object detector is used to inquire about the number of categories to be detected through class-related prompts, iteratively detect spatial coordinates, and the weakly supervised detector is trained in combination with candidate areas and image-level labels. The fusion detection results are used as pseudo-truth values for training of fully supervised object detectors.
The annotation cost in the initialization stage is reduced, the detection capability of the visual language big model is improved, the annotation cost and detection performance is balanced, high-quality false truth generation is achieved, and the fully supervised object detector is further trained.
Smart Images

Figure CN119942178A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an active learning initialization method guided by a large visual language model and a weakly supervised object detector, and belongs to the technical field of initialization of object detectors in computer vision. Background Art
[0002] Active learning object detection technology selects valuable images from unlabeled images for instance-level annotation in the active learning stage, balancing the annotation cost and detection performance. Active learning object detection technology mainly includes two stages: the initialization stage and the active learning stage; among them, the active learning stage depends to a large extent on the initialization of the object detector (initialization stage). In order to ensure that valuable images are mined in the active learning stage, the existing initialization stage randomly selects some images (5%, 6%, 12%) for instance annotation, and then uses the instance-annotated images to train the existing fully supervised detector, so that the object detector has preliminary detection capabilities. The limitations of existing active learning object detection technology include: 1) The annotation cost of the initialization process is too high, and it is highly dependent on instance-level labels. 2) It has a low recall rate.
[0003] Inspired by the large-scale visual language model that uses class-related cues designed with image-level labels to locate objects, the output of the large visual language model provides an opportunity to initialize active learning as pseudo-truth. After experimental exploration, the visual language model locates objects with high precision and low recall, and the weakly supervised object detector locates objects based on thousands of candidate regions with high recall. Figure 1 As shown in Figure 1, the localization results of three existing large visual language models and weakly supervised object detectors are visualized. It can be seen that the existing large visual language models directly applied to the active learning initialization process have a low recall rate. Summary of the invention
[0004] In view of the problems of high labeling cost in the initialization stage of existing active learning object detection technology and low recall rate of detection results of large visual language models, the present invention provides an active learning initialization method guided by a large visual language model and a weakly supervised object detector.
[0005] The present invention provides an active learning initialization method guided by a large visual language model and a weakly supervised object detector, comprising:
[0006] The visual language model is used to obtain the number of categories to be detected in the input image through class-related prompts, and then the spatial coordinate detection result is obtained based on the number of categories to be detected; the detection process is iterated repeatedly to obtain the final spatial coordinate detection result;
[0007] At the same time, a candidate region generation algorithm is used to obtain the candidate region of the input image; the image-level label of the input image and the candidate region are then used to train a weakly supervised object detector, and the object positioning detection result of the category to be detected is obtained;
[0008] The final spatial coordinate detection results and object positioning detection results are fused to obtain pseudo-true values, which are used to train the fully supervised object detector in the active learning initialization phase.
[0009] According to the active learning initialization method guided by a large visual language model and a weakly supervised object detector of the present invention, the large visual language model is a MiniGPT-v2 model.
[0010] According to the active learning initialization method guided by the visual language large model and the weakly supervised object detector of the present invention, the visual question answering identifier [vqa] based on the visual language large model is combined with the class-related prompts designed by the image-level label of the input image, and the number of categories to be detected in the input image is queried to obtain the number of categories to be detected.
[0011] According to the active learning initialization method guided by the visual language large model and the weakly supervised object detector of the present invention, the visual language large model outputs the spatial coordinate detection results of the categories to be detected based on the number of categories to be detected and the object detection identifier [detection].
[0012] According to the active learning initialization method guided by the visual language large model and the weakly supervised object detector of the present invention, the number of categories to be detected obtained by the mth iteration detection is n m for:
[0013] n m =M vl (I,T1)
[0014] Where M vl represents the visual language model, I represents the input image, T1 represents the input text, including the visual question answering identifier [vqa] and the class-related prompts designed by the image-level label of the input image.
[0015] According to the active learning initialization method guided by the visual language large model and the weakly supervised object detector of the present invention, the spatial coordinate detection result obtained by the mth iteration detection is for:
[0016]
[0017] Where T2 is the class-related hint designed for the image-level label of the input image, and M is the maximum number of iterations.
[0018] According to the active learning initialization method guided by the large visual language model and the weakly supervised object detector of the present invention, the spatial coordinate detection result obtained in each next iterative detection is calculated with the spatial coordinate detection result obtained in the adjacent previous iterative detection. When the intersection-and-union ratio is lower than 0.5, the newly generated spatial coordinate detection result is retained and the current spatial coordinate detection result is updated, otherwise the newly generated spatial coordinate detection result is ignored; until the final spatial coordinate detection result is obtained.
[0019] According to the active learning initialization method guided by the visual language large model and the weakly supervised object detector of the present invention, the object positioning detection result of the to-be-detected category output by the weakly supervised object detector is expressed as
[0020]
[0021] Where M w represents a weakly supervised object detector and R represents a candidate region.
[0022] According to the active learning initialization method guided by the visual language large model and the weakly supervised object detector of the present invention, the final spatial coordinate detection result and the object positioning detection result are fused, and the obtained pseudo-true value is expressed as y:
[0023]
[0024] Where η1 is the weak supervision result retention coefficient;
[0025] The weak supervision result retention coefficient η1 is selected based on the confidence of the weak supervision object detector output and the intersection and union ratio of the final spatial coordinate detection result output by the visual language large model and the object positioning detection result output by the weak supervision object detector.
[0026] According to the active learning initialization method guided by the visual language large model and the weakly supervised object detector of the present invention, the calculation method of the weakly supervised result retention coefficient η1 is:
[0027]
[0028] Where C(·) represents the confidence of the output of the weakly supervised object detector, α is the confidence threshold, IoU represents the intersection over union ratio, and β represents the intersection over union ratio threshold.
[0029] Beneficial effects of the present invention: The method of the present invention designs a multi-step reasoning process for the visual language model, including a category number perception step and an iteration step, so as to improve the detection ability of the visual language large model. The visual language large model and the weakly supervised object detector are used to guide the initialization of active learning. The visual language large model based on the multi-step reasoning process locates objects through class-related prompts designed by image-level labels. At the same time, the present invention uses image-level labels to train weakly supervised detectors and locate objects. The detection results of the visual language large model and the weakly supervised object detector are fused as pseudo-true values, and a fully supervised object detector can be further trained.
[0030] The method of the present invention solves the problem of high labeling cost in the initialization stage of active learning object detection technology, strengthens the reasoning ability of large visual language models, breaks through the limitation of traditional active learning initialization methods that require instance-level labels, and balances the labeling cost and detection performance in the active learning initialization process.
[0031] The present invention relates to the research on initializing a fully supervised object detector when annotations are scarce during the active learning initialization process, which, to a certain extent, promotes the implementation of object detection technology based on artificial intelligence deep learning and conforms to the development trend of contemporary intelligent manufacturing. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is a visual comparison of the positioning results of three existing large visual language models and weakly supervised object detectors;
[0033] Figure 2 It is a schematic diagram of the multi-step reasoning process of the visual language large model based on the chain of thought. In the figure, How many motorcycles are there in the image means how many motorcycles are there in the image.
[0034] Figure 3 Schematic diagram of the active learning initialization process guided by the visual language large model and weakly supervised object detector;
[0035] Figure 4 It is a visualization result comparison diagram of the pseudo-true value and the true value obtained by the method of the present invention. DETAILED DESCRIPTION
[0036] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0037] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0038] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.
[0039] Specific implementation method 1. Combination Figure 2 and Figure 3 As shown, the present invention provides an active learning initialization method guided by a large visual language model and a weakly supervised object detector, comprising:
[0040] The visual language model is used to obtain the number of categories to be detected in the input image through class-related prompts, and then the spatial coordinate detection result is obtained based on the number of categories to be detected; the detection process is iterated repeatedly to obtain the final spatial coordinate detection result;
[0041] At the same time, a candidate region generation algorithm is used to obtain the candidate region of the input image; the image-level label of the input image and the candidate region are then used to train a weakly supervised object detector, and the object positioning detection result of the category to be detected is obtained;
[0042] The final spatial coordinate detection results and object positioning detection results are fused to obtain pseudo-true values, which are used to train the fully supervised object detector in the active learning initialization phase.
[0043] This implementation utilizes the high precision of the large visual language model and the high recall of the weakly supervised object detector, and both types of models can be driven by image-level labels (the large visual language model uses class-related cues designed by image-level labels to locate objects, and the weakly supervised object detector has the ability to locate objects after training with image-level labels), thereby achieving active learning initialization without instance-level labeling of images.
[0044] Combination Figure 2 As shown, this embodiment enhances the detection capability of the visual language model based on the multi-step reasoning process of the thought chain. The detection process of the visual language model based on the multi-step reasoning includes a number perception step and an iteration step, with the input being the image and text, and the output being the coordinate information of the bounding box. Figure 2 The black fonts represent the prompt templates, the red fonts represent the image-level labels, the blue fonts represent the number of categories output after the number perception step, and the green fonts represent the task identifiers specific to the visual language model. The active learning initialization process guided by the enhanced visual language model and the weakly supervised object detector is shown in Figure 2. Figure 3The detection results of the enhanced visual language large model and the detection results of the weakly supervised object detector are fused into pseudo-true values, and a fully supervised detector is further initialized. Since the entire initialization process always follows the weakly supervised paradigm, that is, only image-level labels are accessed, the need for instance annotations in the active learning initialization process is reduced.
[0045] As an example, the visual language model is the MiniGPT-v2 model.
[0046] In this implementation, the visual question answering identifier [vqa] based on the visual language large model is combined with the class-related prompts designed based on the image-level labels of the input images to inquire about the number of categories to be detected in the input image and obtain the number of categories to be detected.
[0047] The visual language model outputs the spatial coordinate detection results of the categories to be detected based on the number of categories to be detected and the object detection identifier [detection].
[0048] Combination Figure 2 As shown, the multi-step reasoning process designed in this embodiment imitates the common sense of human beings to detect objects. People roughly count the number of objects to be detected, and then find the objects to be detected based on the counted number. Figure 1 As shown in Figure 2, considering the reasoning ability of the visual language large model, GPT4V has the disadvantage of bounding box offset, so this implementation selects MiniGPT-v2 as the visual language large model M vl. MiniGPT-v2 includes a visual feature extraction network, a linear mapping layer, and a large language model. Among them, the EVA model serves as a visual feature extraction network, and the visual feature extraction network EVA is frozen to ensure the generalization ability of the EVA model. In addition, the linear mapping layer obtains visual tags from the frozen visual feature extractor and projects them into the large language model space. However, in order to solve the problem of reduced training and testing efficiency caused by very long sequence inputs generated by high-resolution images, 4 adjacent visual tags are connected in the visual embedding space and projected into an embedding in the same feature space of the large language model, further reducing the number of visual input tags by 4 times. Through the above process, MiniGPT-v2 can process high-resolution images more efficiently. Then, the open source LLaMA2 is used as the language feature extraction network of MiniGPT-v2. In order to ensure the robustness of the large language model, the visual language tasks are performed directly relying on the LLaMA-2 language tags. The input end of this embodiment includes visual input and text input, and the output end is the coordinate information of the object detection box, for example: " <35> <45> <65> <70> Based on the visual question answering identifier "[vqa]", MiniGPT-v2 is asked the number of the current category for each category. The class-related prompt designed using the image-level label is "[vqa]How many people are in the image?", and the number of objects to be detected is recorded as n. Then, the spatial position of the object to be detected is located by combining the number of current classes n, the image-level label and the object detection identifier [detection] obtained above, for example, "[detection]three persons", and the detection result is recorded as Next, iteratively ask the number of objects to be detected n m and locate the object to be detected Where m represents the number of iterations. Finally, the detection results output by MiniGPT-v2 after multi-step reasoning are Compared with y, it suppresses the problem of missed detection and improves the recall rate of object detection, further improving the performance of object detection.
[0049] In addition, for the weakly supervised detector M w , relying on the candidate region generation network (such as selective search algorithm, multi-scale combination grouping algorithm) to obtain the candidate region R, and output the detection result through the candidate region R and the input image I Detection results of MiniGPT-v2 based on multi-step reasoning And the detection results of weakly supervised object detectors After fusion, it is used as the pseudo-true value of the active learning initialization stage to initialize a fully supervised object detector.
[0050] This implementation constructs a simple and effective bounding box fusion strategy to obtain the pseudo-true value of the initial stage of active learning. For all the detection results of MiniGPT-v2, the detection results that exceed the confidence threshold α and whose intersection-over-union ratio with the detection results of MiniGPT-v2 is lower than the threshold β are selected. Based on the bounding box fusion process, the selection of the detection results of MiniGPT-v2 ensures that the pseudo-true value contains the complete object area, and the selection of the weakly supervised detector ensures that the pseudo-true value contains the complete object area or most of the complete object area, which is used as the pseudo-true value of the active learning initialization process.
[0051] This implementation proposes an initialization method for active learning object detection technology from a novel perspective, which improves the need for instance-level annotation in traditional active learning initialization methods. Through a multi-step reasoning process based on mind linking, the problem of missed detection of the visual language large model is suppressed, the recall rate of the visual language large model is improved, and the detection ability of the visual language large model is strengthened. On this basis, an active learning initialization method guided by the visual language large model and a weakly supervised object detector is designed, and the detection results of the visual language large model and the weakly supervised object detector are integrated to obtain high-quality pseudo-true values. A fully supervised object detector is further trained in the active learning initialization to solve the problem of high annotation cost of the active learning initialization method.
[0052] M based on the visual language model vl The class-related prompts designed by the visual question answering identifier [vqa] and image-level labels of (MiniGPT-v2) ask the number of categories to be detected in the current image, for example: [vqa] How many people are in the image? . The above process can be expressed as:
[0053] n=M vl (I,T1);
[0054] n represents the answer of MiniGPT-v2 of the visual language model. Based on the number of objects to be detected in the current image and the object detection identifier [detection], for example [detection] three persons, MiniGPT-v2 outputs the spatial coordinates of the objects to be detected. The above process can be expressed as:
[0055]
[0056] in Represents the detection results output by MiniGPT-v2, and T2 is the designed class-related prompt.
[0057] Then, in the iteration step, the number of categories is iteratively queried and MiniGPT-v2 iteratively outputs the position of the object to be detected to improve the recall rate of the visual language model. Since the output process of the visual language model is based on a temperature random sampling process, the output results are diversified, which leads to different outputs of MiniGPT-v2 under the premise of inputting the same image and text.
[0058] Furthermore, the number of categories to be detected obtained by the mth iteration is n m for:
[0059] n m =M vl (I,T1)
[0060] Where M vl represents the visual language model, I represents the input image, T1 represents the input text, including the visual question answering identifier [vqa] and the class-related prompts designed by the image-level label of the input image.
[0061] The spatial coordinate detection results obtained by the mth iteration detection for:
[0062]
[0063] Where T2 is the class-related hint designed for the image-level label of the input image, and M is the maximum number of iterations.
[0064] n m is the answer of MiniGPT-v2 to the number of objects to be detected in the current mth iteration. Represents the detection result output by MiniGPT-v2 at the mth iteration step.
[0065] After iteratively querying the position of the object to be detected, the spatial coordinate detection result obtained in each next iterative detection is calculated by intersection and union ratio with the spatial coordinate detection result obtained in the previous iterative detection. When the intersection and union ratio is lower than 0.5, the newly generated spatial coordinate detection result is retained and the current spatial coordinate detection result is updated, otherwise the newly generated spatial coordinate detection result is ignored; until the final spatial coordinate detection result is obtained.
[0066] Based on the above process, the enhanced MiniGPT-v2 detection results were obtained. In order to further ensure the recall rate of the detection results, a weakly supervised object detector that only requires image-level label training was introduced.
[0067] Furthermore, the object location detection result of the to-be-detected category output by the weakly supervised object detector is expressed as
[0068] Where M wrepresents a weakly supervised object detector and R represents a candidate region.
[0069] With thousands of candidate regions as input, weakly supervised detection results It has a high recall rate. To this end, the detection results based on the enhanced visual language model And the detection results of weakly supervised object detectors with high recall Fusion is performed to obtain the pseudo-true value. A fully supervised detector is trained in the active learning initialization phase. This implementation adopts a simple and effective boundary fusion method. Since MiniGPT-v2 outputs a compact bounding box, all detection results of MiniGPT-v2 are selected. In addition, since the weakly supervised object detector has a low precision and a high recall rate, for the detection results of the weakly supervised object detector, the detection results with a confidence higher than the threshold α and an intersection-over-union ratio with the detection results of MiniGPT-v2 less than the threshold β are selected.
[0070] The final spatial coordinate detection results and object positioning detection results are fused, and the obtained pseudo-true value is expressed as y:
[0071]
[0072] Where η1 is the weak supervision result retention coefficient;
[0073] The weak supervision result retention coefficient η1 is selected based on the confidence of the weak supervision object detector output and the intersection and union ratio of the final spatial coordinate detection result output by the visual language large model and the object positioning detection result output by the weak supervision object detector.
[0074] The calculation method of weak supervision result retention coefficient η1 is:
[0075]
[0076] Where C(·) represents the confidence of the output of the weakly supervised object detector, α is the confidence threshold, IoU represents the intersection over union ratio, and β represents the intersection over union ratio threshold.
[0077] Verification experiment:
[0078] Prepare training samples. VOC2007 / 2012 and MS COCO 2017 datasets are selected to verify the effectiveness of the present invention. The model proposed in the present invention is further trained using a training set of 5011 images in the VOC2007 dataset and a training set of 11540 images in the VOC2012 dataset, and the accuracy of the model is evaluated using 4952 test images in the VOC2007 test set. In addition, COCO2017 is used to train and verify the model. The training set includes 118,000 images, the verification set includes 5,000 images, and there are 80 categories of objects to be detected. Not all images in the VOC2007 / 2012 and COCO2017 training sets are used to train the model. The first 5,000 images in the VOC dataset and the COCO dataset are selected to obtain pseudo-true values, further realizing the active learning initialization process. The experimental results follow the official evaluation indicators. The mAP accuracy of the model is evaluated on the test set of the VOC dataset, and the AP, AP50, and AP75 of the model are evaluated on the validation set of the COCO dataset.
[0079] MiniGPT-v2 is used as the large model of visual language, OD-WSCL is used as the baseline weakly supervised object detector, and RetinaNet is used as the fully supervised detector. The first 5000 images and corresponding image-level labels are selected from the VOC dataset and the COCO dataset to generate pseudo-truth values, and the fully supervised detector RetinaNet is further initialized. 0.1 is selected as the temperature hyperparameter of MiniGPT-v2 to obtain robust and stable detection results, and the iteration step M=2 is set to improve the recall rate of MiniGPT-v2. In addition, the confidence threshold α is set to 0.95 and the intersection-over-union threshold β is set to 0.5 to fuse the detection results of MiniGPT-v2 and OD-WSCL as pseudo-truth values. For fair comparison, the batch size, optimizer and learning rate used in the active learning initialization process of the present invention are consistent with the baseline model. Specifically, the RetinaNet detector is trained for 26 epochs, and the learning rate decays to 0.1 times the original at the 20th epoch. The RetinaNet detector uses the SGD optimizer, and the initial learning rate is set to 0.01. The OD-WSCL detector is trained for 30K iterations using the SGD optimizer with an initial learning rate of 0.01. The code for this method is based on PyTorch and trained using 8 RTX3090 GPUs.
[0080] Experiments have shown that the method of the present invention reduces the need for instance-level annotation in the initialization phase of traditional active learning object detection technology, improves the detection capability of large visual language models, mines pseudo-true values that contain complete object areas, and improves the detection performance of the initialization process. The effectiveness of the method proposed in the present invention was verified by training on the general VOC2007 / 2012 dataset and testing on the test set of VOC2007. The effectiveness of the method proposed in the present invention was verified by training on the general VOC2007 / 2012 dataset and testing on the test set of VOC2007. Finally, the pseudo-true values and true values were displayed on the VOC dataset.
[0081] 1) Analysis of the effectiveness of the multi-step reasoning process MiniGPT-v2 combined with weakly supervised object detectors for active learning initialization on the VOC2007 test set. As shown in Table 1, it can be observed that the detection results based on the original MiniGPT-v2 output are used as pseudo-truths to train the RetinaNet detector, and a mAP of 59.1% is obtained. This is higher than the traditional method of randomly selecting instance-level labels to initialize the RetinaNet detector (43.4% vs 59.1%). By introducing a multi-step reasoning process, the detection capability of the visual language model MiniGPT-v2 is enhanced, and the recall rate of MiniGPT-v2 is effectively improved. The detection results of the enhanced visual language model MiniGPT-v2 are used as pseudo-truths to train the RetinaNet detector, achieving a mAP of 62.6%, which exceeds the performance of the RetinaNet detector initialized by the detection results of the original MiniGPT-v2 output by 3.5%. In addition, in order to ensure the recall rate of pseudo-true values, the present invention further introduces the detection results of the weakly supervised object detector OD-WSCL and the detection results output by MiniGPT-v2 based on multi-step reasoning as pseudo-true values, and finally achieves a mAP of 62.9%.
[0082] Table 1 Analysis of the effectiveness of MiniGPT-v2 combined with weakly supervised object detectors in the multi-step reasoning process for the active learning initialization process
[0083]
[0084] 2) Analysis of the effectiveness of the number perception step and iteration step in the multi-step reasoning process on the detection ability of the visual language large model on the VOC2007 test set and the COCO2017 validation set. As shown in Table 2, it can be observed that the category number perception step improves the detection performance mAP from 46.4% to 50.4% on the VOC2007 test set and from 27.9% to 28.5% on the COCO2017 validation set. By further introducing the iteration step, it is improved from 50.4% to 51.5% on the VOC2007 test set and from 28.5% to 29.0% on the COCO2017 validation set. Based on the above experiments, the effectiveness of the number perception step and iteration step proposed in the present invention on the detection ability of the visual language large model is proved.
[0085] Table 2 Analysis of the effectiveness of the number perception step and iteration step in the multi-step reasoning process on the detection ability of the visual language large model
[0086]
[0087] 3) The impact of the proportion of image-level labels on the VOC20007 test set on the initialization performance of RetinaNet
[0088] As shown in Table 3, it can be observed that the present invention uses the first 5000 images and the corresponding image-level labels to initialize the RetinaNet detector, achieving a mAP of 62.9%, which is a significant improvement (62.9% vs 43.4%) compared to the traditional method of initializing the RetinaNet detector using 5% instance-level labels. In addition, in order to further verify the effectiveness of the present invention in initializing the fully supervised detector using images and corresponding image-level labels, the present invention compares the effects of 5%, 6%, 7%, 10%, and 20% image-level labels on the initialization performance of the fully supervised detector. As shown in the first and second rows of Table 3, the RetinaNet detector is initialized using pseudo-truth values obtained from image-level labels and instance-level labels corresponding to the same number of images, achieving a competitive result with the traditional method (as shown in the second row of Table 3) (42.7% vs 43.4%). It is worth noting that the initialization method of the present invention does not access any instance-level labels. Then, the present invention gradually increases the proportion of images and corresponding image-level labels. It can be observed that the pseudo-truth values obtained by initializing the RetinaNet detector using only 6% of the images and corresponding image-level labels have greatly exceeded the traditional method (as shown in the second row of Table 3). The present invention increases the proportion of images and corresponding image-level labels to further improve the initialization performance of active learning object detection technology. Considering the cost of image-level annotation and the initialization detection performance, the present invention selects the first 5,000 images and image-level labels (30%) on the VOC2007 / 2012 dataset.
[0089] Table 3 Impact of the ratio of image-level labels on the VOC20007 test set on the initialization performance of RetinaNet
[0090]
[0091]
[0092] 4) The impact of the proportion of image-level labels on the COCO2017 validation set on the initialization performance of RetinaNet:
[0093] As shown in Table 4, it can be observed that the traditional initialization method on the COCO dataset uses 2% of instance-level labels to initialize the RetinaNet detector, and obtains 7.4%, 15.9%, and 6.5% AP, AP50, and AP75. The performance of the RetinaNet detector is relatively weak when the pseudo-true values obtained by only using the first 2% of images (2400 images) and the corresponding image-level labels are initialized. This is because the first 2% of images are more difficult to cover the 80 categories in the COCO dataset, resulting in the absence of some categories in the pseudo-true values, and the detection performance of the RetinaNet detector cannot be fully exploited. Next, the present invention continuously increases the ratio of images and corresponding image-level labels. It is observed that when 4% of images and corresponding image-level labels are used, it exceeds the traditional initialization method and achieves excellent active learning initialization performance (8.4% AP, 16.8% AP50, 7.4% AP75). Similar to the conclusions obtained in Table 3, the present invention believes that continuously increasing images and corresponding image-level labels can improve the performance of active learning initialization. In order to balance the annotation cost of image-level labels and the initialization detection performance, the present invention uses 4% of the images (about 5,000 images) and corresponding image-level labels instead of the traditional method of initializing the RetinaNet detector with 2% instance-level labels.
[0094] Table 4 Impact of the ratio of image-level labels on the COCO2017 validation set on the initialization performance of RetinaNet
[0095]
[0096] 5) Visualization results of pseudo-true values and true values. Figure 4The figure shows some visualization results of pseudo-truth values and true values, where the red box represents the output pseudo-truth value bounding box and the green box represents the true value bounding box. The pseudo-truth value observed by the method of the present invention contains the complete object area or most of the object area. Even if there are dense objects, very small objects, and occluded objects in the scene, the bounding box output by the present invention through the fusion of the large visual language model and weakly supervised object detection is still very close to the true value bounding box. For example, in the first and second images of the third row, there are dense "herds of cattle" and "crowds", and the pseudo-truth value generated by the present invention does not have the problem of missed detection. For example, in the first image of the second row, although the "sheep" occupies very few pixels compared to the entire image, the pseudo-truth value generated by the present invention still compactly contains the complete "sheep". For example, in the second image in the first row and the third image in the second row, the "sheep" blocks the "front of the car" and the "bicycle" blocks the lower body of the "person", and the present invention still achieves high precision and recall. Thanks to the initialization method of active learning object detection technology proposed in the present invention, the high accuracy of pseudo-true values is ensured through the visual language large model MiniGPT-v2, and the high recall rate of pseudo-true values is achieved through the weakly supervised object detector. The visualization results further illustrate the effectiveness and robustness of the present invention.
[0097] Although the present invention is described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. It should therefore be understood that many modifications may be made to the exemplary embodiments and that other arrangements may be devised without departing from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the various dependent claims and features described herein may be combined in a manner different from that described in the original claims. It should also be understood that the features described in conjunction with a single embodiment may be used in other described embodiments.
Claims
1. An active learning initialization method guided by a large visual language model and a weakly supervised object detector, characterized in that include, The visual language model is used to obtain the number of categories to be detected in the input image through class-related prompts, and then the spatial coordinate detection results are obtained based on the number of categories to be detected; Repeat the iterative detection process to obtain the final spatial coordinate detection results; At the same time, a candidate region generation algorithm is used to obtain the candidate region of the input image; Then, the image-level labels and candidate regions of the input image are used to train the weakly supervised object detector, and the object location detection results of the category to be detected are obtained; The final spatial coordinate detection results and object positioning detection results are fused to obtain pseudo-true values, which are used to train the fully supervised object detector in the active learning initialization phase.
2. The active learning initialization method guided by a large visual language model and a weakly supervised object detector according to claim 1, characterized in that: The visual language model is the MiniGPT-v2 model.
3. The active learning initialization method guided by a large visual language model and a weakly supervised object detector according to claim 1, characterized in that: The visual question answering identifier [vqa] based on the visual language large model combines the class-related prompts designed based on the image-level labels of the input images, asks the number of categories to be detected in the input image, and obtains the number of categories to be detected.
4. The active learning initialization method guided by a large visual language model and a weakly supervised object detector according to claim 3, characterized in that: The visual language model outputs the spatial coordinate detection results of the categories to be detected based on the number of categories to be detected and the object detection identifier [detection].
5. The active learning initialization method guided by a large visual language model and a weakly supervised object detector according to claim 4, characterized in that: The number of categories to be detected obtained by the mth iteration detection n m for: n m =M vl (I,T1) Where M vl represents the visual language model, I represents the input image, T1 represents the input text, including the visual question answering identifier [vqa] and the class-related prompts designed by the image-level label of the input image.
6. The active learning initialization method guided by a large visual language model and a weakly supervised object detector according to claim 5, characterized in that: The spatial coordinate detection results obtained by the mth iteration detection for: Where T2 is the class-related hint designed for the image-level label of the input image, and M is the maximum number of iterations.
7. The active learning initialization method guided by a large visual language model and a weakly supervised object detector according to claim 6, characterized in that: The spatial coordinate detection result obtained in each next iterative detection is calculated by the intersection and union ratio with the spatial coordinate detection result obtained in the previous iterative detection. When the intersection and union ratio is lower than 0.5, the newly generated spatial coordinate detection result is retained and the current spatial coordinate detection result is updated. Otherwise, the newly generated spatial coordinate detection result is ignored until the final spatial coordinate detection result is obtained.
8. The active learning initialization method guided by a large visual language model and a weakly supervised object detector according to claim 7, characterized in that: The object location detection result of the to-be-detected category output by the weakly supervised object detector is expressed as Where M w represents a weakly supervised object detector and R represents a candidate region.
9. The active learning initialization method guided by a large visual language model and a weakly supervised object detector according to claim 8, characterized in that: The final spatial coordinate detection results and object positioning detection results are fused, and the obtained pseudo-true value is expressed as y: Where η1 is the weak supervision result retention coefficient; The weak supervision result retention coefficient η1 is selected based on the confidence of the weak supervision object detector output and the intersection and union ratio of the final spatial coordinate detection result output by the visual language large model and the object positioning detection result output by the weak supervision object detector.
10. The active learning initialization method guided by a large visual language model and a weakly supervised object detector according to claim 9, characterized in that: The calculation method of weak supervision result retention coefficient η1 is: Where C(·) represents the confidence of the output of the weakly supervised object detector, α is the confidence threshold, IoU represents the intersection over union ratio, and β represents the intersection over union ratio threshold.
Citation Information
Cited By
Multi-agent cooperative traffic signal control method, system, equipment and medium
CN121281292A
Target detection method based on unsupervised feature clustering and multi-modal large model collaborative iteration
CN121883895A