Open-world Object Detection Method Based on Region Perception and Unknown Information Enhancement
By introducing a method based on region-aware and unknown object information enhancement in open world object detection, the UCIE module and RARF module are used to solve the problem of false detection and missed detection of object detection in open world environments, and the balance of detection performance and accuracy recall is improved.
Patent Information
- Application Number
- CN202410660697.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-27
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-05-27
AI Technical Summary
Existing target detection technologies have problems of false detection and missed detection in open-world environments, and cannot effectively deal with unknown target categories.
The open-world object detection method based on region-awareness and unknown object information enhancement is adopted to improve detection performance through the UCIE module and the RARF module. The UCIE module contains unknown target object-score regression header, object-selective evaluation header, real box updater, and unknown class classifier for enhancing unknown class information. The RARF module uses the region-aware module and the Ncut algorithm to filter redundant boxes to reduce the number of unknown object prediction boxes.
It effectively reduces the false detection and missed detection rates of targets in an open world environment, improves the detection performance of unknown objects, and achieves better accuracy and recall balance.
Smart Images

Figure CN119942057B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular, to an open-world object detection method based on region perception and unknown information enhancement. Background Art
[0002] The task of object detection is to accurately and efficiently identify and locate a large number of object instances of predefined categories from images. Deep learning has accelerated the progress of object detection research. With the in-depth research and wide application of deep learning, people have achieved fruitful success in visual tasks. In recent years, researchers have made many improvements to the detection model based on deep learning methods and achieved remarkable progress. The accuracy and efficiency of object detection have been greatly improved. These model detection methods are all carried out under a strong assumption that all classes to be detected will be available during the training phase, in other words, all object categories are known during the training phase. However, in actual object detection tasks, this setting in the closed set training effect often cannot meet the requirements, which greatly limits the development and application of object detection technology. Therefore, in order to build an intelligent bridge between humans and machines and promote the practical application of object detection technology in the open world, it is of great significance to study the open-world object detection problem. Summary of the Invention
[0003] The purpose of the present invention is to provide an open-world object detection method based on region perception and unknown object information enhancement to solve the problem of how to reduce the false detection and missed detection of objects in the open-world environment.
[0004] The present invention adopts the following technical solutions:
[0005] An open-world object detection method based on region perception and unknown object information enhancement, comprising the following steps:
[0006] Step s1, unify the specifications of the pictures in the dataset, set their sizes to multiples of 32, and standardize and normalize the dataset data using the average and variance;
[0007] Step s2, use the backbone feature extraction network to extract features;
[0008] Step s3, use the feature information of each layer of the feature pyramid as the feature map and input it into the RPN for object location, and output a set of proposal boxes that may contain objects;
[0009] Step s4, process the proposal boxes obtained in step s3 and the feature map obtained in step s2 through the ROIPooling layer to obtain a set of proposal box feature maps with a fixed size;
[0010] Step S5: Input the feature map obtained in step S4 into the Unknown Class Information Enhancement (UCIE) module to obtain a set of ground truth boxes containing the unknown class bounding boxes;
[0011] UCIE module: It includes an unknown object objectness score regression head, an objectness evaluation head, a ground truth box updater, and an unknown class classifier:
[0012] First, the unknown object objectness score regression head consists of a fully connected neural network with an input of 1024 and an output of 1 and a set of linear operations, which evaluates the object confidence of the proposed feature map obtained in step S4. The objectness evaluation head calculates the intersection over union (IoU) score between the ground truth box and the predicted box using the IoU method. The IoU calculation formula is:
[0013]
[0014] where q i represents the predicted box, and b k represents the ground truth box. Use the obtained score to find the known object proposal boxes in the proposed feature map. The unknown classifier calculates the possibility that the proposed feature map contains unknown objects using the IOP, IOC, and IOK score methods. The specific calculation formulas for IOP, IOC, and IOK are:
[0015]
[0016]
[0017]
[0018] where w ij represents the horizontal distance between prediction box i and prediction box j, h ij represents the vertical distance between prediction box i and prediction box j, and p j represents the area size of q j . According to the score, initially obtain a set of unknown object proposal boxes, and then input the obtained known object proposal boxes and unknown object proposal boxes into the ground truth box updater to obtain a set of ground truth boxes containing more unknown information;
[0019] Step S6: Input the set of ground truth boxes obtained in step S5 into the classifier and regressor to perform classification and regression operations on the proposed feature map obtained in the step;
[0020] Step S7: Combine step S4, step S5, and step S6 to perform detection learning on the known objects and unknown objects in the dataset images. Then calculate the loss according to the ground truth, and use the optimization algorithm to update the weight parameters of the target feature extraction network to reach the convergence state;
[0021]
[0022] where is the classifier loss, calculated by a cross - entropy function, represents the total loss of the RPN layer, represents the smooth loss of proposal box regression, represents the loss of the unknown class enhancement module, calculated by an optimized loss function.
[0023]
[0024] where l represents a constant threshold, represents the set of unknown class proposal boxes, Ψ(f i ) represents the feature map feature f i .
[0025] the objectness confidence score of;
[0026] The method of the present invention proposes a region - aware based unknown redundant box filtering method (Region Aware Based Unknown Redundant Box Filtering, RARF). RARF optimizes the detection performance of the detection model for unknown objects and is a new type of post - processing mechanism for unknown objects. In this method, we design a region - aware module (Region Aware Module, RA), which can effectively reduce the number of unknown object prediction boxes that contain too many or too few unknown objects. Specifically, we use the normalized cut algorithm (Normalized cut, Ncut) to preliminarily filter the unknown object prediction boxes obtained in step s6, removing some prediction boxes with similar features. Then, they are sent to RA for region similarity calculation, and the prediction boxes with too high or too low RA scores are filtered out, effectively reducing the occurrence of the situation where a prediction box contains multiple unknown objects or only a small part of an unknown object. The region similarity evaluation formula:
[0027]
[0028] where q i represents the i - th prediction box, I represents the detected image, w i and h i respectively represent the width and height of the prediction box q i , and a I represents the region size of the image I;
[0029] Step s8: Using the open-world images that have not participated in training and the rectangular box annotations as inputs, perform forward operations on the inputs using the trained and converged target feature extraction network to generate prediction results, and send the prediction results into the RARF module to filter out redundant prediction boxes of unknown objects and generate the final detection results. Description of the Drawings
[0030] Figure 1 is the flowchart of the method of the present invention;
[0031] Figure 2 is the network structure diagram of the method of the present invention;
[0032] Figure 3 is the UCIE module;
[0033] Figure 4 is the RARF module; Detailed Embodiments
[0034] The present invention proposes an open-world object detection method based on region-aware unknown class information enhancement. The following further explains the detailed embodiments and basic principles of the present invention with reference to the accompanying drawings. As Figure 1 shown, it specifically includes the following processes:
[0035] Step s1: Data processing. Taking the PASCAL VOC dataset as an example, first unify the specifications of the pictures in the dataset, set their sizes to multiples of 32, and perform standardization and normalization processing on the dataset data using the average value and variance;
[0036] Step s2: Feature extraction. Using the network and the feature pyramid as the backbone feature network, save the operation results of each layer of the feature pyramid as the feature extraction results;
[0037] Step s3: Input the feature extraction results obtained in step s2 into the RPN. The RPN uses a 3×3 convolution to generate anchor boxes, and then uses 2 1×1 convolution layers and a softmax layer to generate positive anchor boxes and bounding box regression offsets, and then outputs a set of proposal boxes that may contain objects;
[0038] Step s4: Input the proposal boxes obtained in step s3 and the feature map obtained in step s2 into the ROIPooling layer for max pooling processing, and output a set of proposal box feature maps of a fixed size;
[0039] Step s5: Input the proposal box feature maps obtained in step s4 into the Unknown Class Information Enhancement (UCIE) module to extract the proposal boxes containing known objects and unknown objects, and update the ground truth boxes;
[0040] AsFigure 3 The UCIE module includes an unknown target objectiveness score regression head, an objectiveness evaluation head, a ground truth box updater, and an unknown class classifier. To avoid confusion between known classes, unknown classes, and non-object classes, we use thresholds ξ, γ, and σ to divide non-object classes and object classes based on objectiveness scores, known classes and unknown classes, and unknown classes and non-object classes, respectively. This method enhances the classification and localization capabilities of the classifier and regressor for targets in an open environment by adaptively updating the ground truth boxes;
[0041] Step s6: Feed the set of ground truth boxes obtained in step s5 into the classifier and regressor to perform classification and regression operations on the proposed feature map obtained in the step;
[0042] Step s7: Combine steps s4, s5, and s6 to perform detection learning on known and unknown objects in the dataset images. Then calculate the loss based on the ground truth values and use an optimization algorithm to update the weight parameters of the target feature extraction network to reach a converged state;
[0043] Step s8: Use the open-world images and rectangle box annotations that have not participated in training as inputs, perform forward operations on the inputs using the trained and converged target feature extraction network to generate prediction results. Feed the prediction results into the RARF module, and use the region-aware method to filter out prediction boxes that contain multiple unknown objects or only a small part of a single unknown object to generate the final detection results.
[0044] As Figure 4 The RARF module includes an NMS module, an Ncut module, a similarity judgment module, and an RA module. To effectively filter redundant boxes of known and unknown objects, we use the standard post-processing mechanism NMS to process known prediction boxes, use Ncut to filter similar features of unknown object prediction boxes, then use the objectiveness determination module to filter known object prediction boxes, and finally use the RA module to remove prediction boxes with too many or too few unknown objects.
[0045] Experimental settings:
[0046] This experiment was trained for a total of 18,000 epochs. The iterative training in the second stage started from the 12,000th epoch, using the Stochastic Gradient Descent (SGD) optimizer and the cosine annealing learning rate scheduling strategy with an initial learning rate of 0.02, and a weight decay of 10 -4 , and a momentum of 0.9.
[0047] The experimental results are shown in Table 1-2. On the COCO-OOD dataset, the U-F1 score of our invention is 3.3% higher than that of the second-place UnSniffer, and our U-PRE score is 9.6% higher than that of the second-place UnSniffer. However, there is also a corresponding cost. Our U-AP and U-RCE metrics are inferior to the best metrics respectively. On the COCO-Mix dataset, our proposed method is optimal in terms of U-AP, U-F1, and U-PRE metrics. Our U-AP metric is 1.6% higher than that of the second-place ORE, our U-F1 metric is 1.7% higher than that of the second-place UnSniffer, and our U-PRE metric is also 4.9% higher than it. These comparisons show that our invention is superior to existing methods in detecting unknown objects. By comparing Experiment 1 and Experiment 2, it can be seen that adding the unknown class information enhancement module can effectively improve the model's detection performance for unknown objects. Compared with the base class model, our invention achieved an increase of 3.3% in U-F1, 5.7% in U-PRE, and 0.6% in U-REC on the OOD dataset, and an increase of 1.4% in U-AP, 1% in U-F1, 0.9% in U-PRE, and 0.2% in U-REC on the Mix dataset. Comparing Experiment 1 and Experiment 3, it can be found that our method is superior to the base class method in terms of U-AP, U-F1, and U-PRE metrics. On the OOD dataset, our invention improved the model's U-AP, U-F1, and U-PRE metrics by 0.3%, 1.3%, and 4.2% respectively. On the Mix dataset, the U-AP, U-F1, and U-PRE metrics increased by 1.7%, 0.8%, and 1.7% respectively. This shows that using the region-aware unknown class redundant box filtering method is robust in the open-world object detection task and plays a positive role in improving the model performance. Finally, through the result comparison of Experiment 4, it can be found that after adding the UCIE and RABF modules, most of the metrics of the model on the two datasets reach the optimal, which proves that there is a mutually promoting effect between the two modules and fully proves the feasibility and effectiveness of the proposed method of the present invention. The above experiments show that the present method achieves a better balance between precision and recall.
[0048] Table 1 Comparison of the results of the method of the present invention and other methods
[0049]
[0050]
[0051] Table 2 Comparison of the results of using / without using UCIE and RABF respectively
[0052]
Claims
1. An open world target detection method based on region perception and unknown information enhancement, characterized in that: The method comprises: Step 1: Unify the specifications of the images in the dataset, set their sizes to multiples of 32, and use the mean and variance to standardize and normalize the dataset data; Step 2, using the backbone feature extraction network to perform feature extraction; Step 3: Input the feature information of each layer of the feature pyramid as a feature map into the RPN to locate the target object and output a set of proposal boxes that may contain the object; Step 4: Process the proposal box obtained in step 3 and the feature map obtained in step 2 through the ROIPooling layer to obtain a set of proposal box feature maps of fixed size; Step 5: Input the feature map obtained in step 4 into the unknown class information enhancement UCIE module to obtain a real box set containing the unknown class bounding box; the UCIE module is used to enhance the unknown object information during the training process. The UCIE module consists of a target object score regression head Ψ, an object evaluation head, a real box updater and an unknown class classifier; Ψ contains a fully connected neural network with an input of 1024 and an output of 1 and a set of linear operations. The object evaluation head uses the IOU algorithm to calculate the probability of containing a known object in the proposal box: where q i represents the prediction box, b k Represents the true box; the unknown classifier includes IOP, IOC and IOK functions. IOP and IOC extract the proposed box that may contain the unknown object by comparing the proposed box with the true box. IOK further extracts and filters the proposed box by calculating the similarity of the proposed box. where w ij represents the horizontal distance between prediction box i and prediction box j, h ij represents the vertical distance between prediction box i and prediction box j, p j Indicates q j The size of the region is then input as the Ψ object confidence evaluation to ensure that the unknown object proposal box really contains the unknown object. The real box updater uses an iterative algorithm to continuously update the bounding box information of the unknown object in the real box. Step 6: Send the real frame set obtained in step 5 to the classifier and regressor to perform classification and regression operations on the proposed feature map obtained in step 5; Step 7: Combine steps 4, 5, and 6 to detect and learn the known objects and unknown objects in the dataset image, then calculate the loss based on the true value, and use the optimization algorithm to update the weight parameters of the target feature extraction network to reach a convergence state; Step 8: Take the open world image and rectangular box annotation that are not involved in the training as input, use the trained converged target feature extraction network to perform forward operation on the input to generate a prediction result, and send the prediction result to the region-aware unknown class redundant box filtering RARF module to filter the redundant prediction boxes of the unknown object to generate the final detection result.
2. The open world target detection method based on region perception and unknown information enhancement according to claim 1, characterized in that: Step 8 introduces the RARF module to better filter out redundant prediction boxes on unknown objects. The module contains an NMS module, a normalized cut Ncut module, a similarity judgment module, and a region-aware RA module. The workflow is as follows: s1, using the softmax function of the standard post-processing mechanism NMS, only the predicted box with the highest matching score with the relevant real box is retained as the known class detection result; s2, use the Ncut module to calculate the similarity between each unknown object proposal box and remove the unknown object proposal boxes with too high similarity; s3, input the known proposal box obtained in s1 and the preliminary unknown proposal box obtained in s2 into the similarity judgment module to evaluate the similarity of known objects, and remove the unknown object proposal box with a score greater than the threshold θ; s4, input the unknown object proposal box obtained in s3 into the RA module, calculate the regional similarity between the proposal box and the entire image, and drop the unknown object proposal boxes with scores higher than the threshold ρ and lower than the threshold τ in the region to obtain a set of more accurate unknown object proposal boxes: