Open world target detection method based on region perception and unknown information enhancement
By introducing a method based on region-awareness and unknown object information enhancement in object detection technology, the error detection and missed detection problems of object detection in open world environments are solved, and higher detection accuracy and recall rate are achieved.
Patent Information
- Application Number
- CN202410660697.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-05-27
AI Technical Summary
Existing target detection technologies have problems of false detection and missed detection in open-world environments, and cannot effectively deal with unknown target categories.
The open-world object detection method based on region-awareness and unknown object information enhancement is adopted, and the model's detection ability of unknown classes is enhanced through the UCIE module and the RARF module, and the redundant boxes are filtered through region similarity calculation and Ncut algorithm.
It effectively reduces the target misdetection and missed detection rates in an open world environment, and improves the detection performance and accuracy of the model for unknown objects.
Smart Images

Figure CN119942057A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to an open-world target detection method based on region perception and unknown information enhancement. Background Art
[0002] The task of object detection is to accurately and efficiently identify and locate a large number of object instances of predefined categories from an image. Deep learning has accelerated the progress of object detection research. With the in-depth study and wide application of deep learning, people have achieved fruitful success in visual tasks. In recent years, researchers have made many improvements to the detection model based on deep learning methods and have made significant progress. The accuracy and efficiency of object detection have been greatly improved. These model detection methods are all based on a strong assumption that all classes to be detected will be available in the training phase. In other words, all target categories are known in the training phase. However, in actual object detection tasks, the training effect of this setting in a closed set often cannot meet the needs, which greatly limits the development and application of object detection technology. Therefore, in order to build an intelligent bridge between humans and machines and promote the practical application of object detection technology in the open world, it is of great significance to study the problem of open world object detection. Summary of the invention
[0003] The purpose of the present invention is to provide an open-world target detection method based on region perception and unknown object information enhancement, so as to solve the problem of how to reduce false detection and missed detection of targets in an open-world environment.
[0004] The present invention adopts the following technical solutions:
[0005] An open world object detection method based on region perception and unknown object information enhancement comprises the following steps:
[0006] Step s1, unify the specifications of the images in the data set, set their sizes to multiples of 32, and use the mean and variance to standardize and normalize the data set data;
[0007] Step s2, using the backbone feature extraction network to perform feature extraction;
[0008] Step s3: Input the feature information of each layer of the feature pyramid as a feature map into the RPN to locate the target object and output a set of proposal boxes that may contain the object;
[0009] Step s4, the proposal box obtained in step s3 and the feature map obtained in step s2 are processed by the ROIPooling layer to obtain a set of proposal box feature maps of fixed size;
[0010] Step s5, inputting the feature map obtained in step s4 into the Unknown Class Information Enhancement (UCIE) module to obtain a real box set containing the unknown class bounding box;
[0011] UCIE module: contains an unknown target object score regression head, an object evaluation head, a true box updater and an unknown class classifier:
[0012] First, the unknown target object score regression head consists of a fully connected neural network with an input of 1024 and an output of 1 and a set of linear operations to evaluate the object confidence of the proposed feature map obtained in step s4. The object evaluation head uses the IOU method to calculate the intersection-over-union score of the real box and the predicted box. The IOU calculation formula is:
[0013]
[0014] where q i represents the prediction box, b k Represents the true box, and uses the obtained score to find the known object proposal box in the proposed feature map. The unknown classifier uses the IOP, IOC and IOK scoring methods to calculate the possibility that the proposed feature map contains unknown objects. The specific calculation formulas of IOP, IOC and IOK are:
[0015]
[0016]
[0017]
[0018] where w ij represents the horizontal distance between prediction box i and prediction box j, h ij represents the vertical distance between prediction box i and prediction box j, p j Indicates q j Based on the scores, a set of unknown object proposal frames are initially obtained, and then the known object proposal frames and unknown object proposal frames are input into the real frame updater to obtain a set of real frames containing more unknown information;
[0019] Step s6, sending the true frame set obtained in step s5 to the classifier and regressor to perform classification and regression operations on the proposed feature map obtained in the step;
[0020] Step s7, combining steps s4, s5, and s6 to detect and learn the known objects and unknown objects in the dataset image, then calculating the loss according to the true value, and using the optimization algorithm to update the weight parameters of the target feature extraction network to reach a convergence state;
[0021]
[0022] in is the classifier loss, calculated by a cross entropy function, represents the total loss of the RPN layer, represents the smooth loss of proposal box regression, Represents the loss of the unknown class enhancement module, which is calculated by an optimization loss function.
[0023]
[0024] Where l represents a constant threshold, represents the set of unknown class proposal boxes, Ψ(f i ) represents the feature map feature f i .
[0025] The object confidence score of
[0026] The method of the present invention proposes a region-aware based unknown redundant box filtering method (RARF). RARF optimizes the detection performance of the detection model for unknown objects and is a new type of unknown object post-processing mechanism. In this method, we designed a region-aware module (RA), which can effectively reduce the number of unknown object prediction boxes that contain too many or too small unknown objects. Specifically, we use the normalized cut algorithm (Ncut) to perform preliminary filtering on the unknown object prediction box obtained in step s6 to remove some prediction boxes with similar features. Then it is sent to RA for regional similarity calculation to filter out prediction boxes with RA scores that are too high or too low, effectively reducing the occurrence of a prediction box containing multiple unknown objects or only a small part of an unknown object. Regional similarity evaluation formula:
[0027]
[0028] where q i represents the i-th prediction box, I represents the detected image, and w i and h i Represent the predicted box q i The width and height of a I Represents the area size of image I;
[0029] In step s8, the open world image and rectangular box annotation that have not participated in the training are taken as input, and the target feature extraction network that has converged in the training is used to perform forward operation on the input to generate a prediction result. The prediction result is sent to the RARF module to filter the redundant prediction boxes of the unknown object to generate the final detection result. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a flow chart of the method of the present invention;
[0031] Figure 2 It is a network structure diagram of the method of the present invention;
[0032] Figure 3 It is the UCIE module;
[0033] Figure 4 It is the RARF module; DETAILED DESCRIPTION
[0034] The present invention proposes an open world target detection method based on region-aware unknown class information enhancement. The specific implementation methods and basic principles of the present invention are further described below in conjunction with the accompanying drawings. Figure 1 As shown, the specific process includes the following:
[0035] Step s1: Data processing. Taking the PASCAL VOC dataset as an example, first, the images in the dataset are standardized and their sizes are set to multiples of 32. The mean and variance are used to standardize and normalize the dataset data.
[0036] Step s2: feature extraction, using the network and feature pyramid as the backbone feature network, and saving the operation results of each layer of the feature pyramid as the feature extraction result;
[0037] Step s3: The feature extraction result obtained in step s2 is input to RPN. RPN generates anchor boxes using a 3×3 convolution, and then uses two 1×1 convolution layers and a softmax layer to generate forward anchor boxes and bounding box regression offsets, and then outputs a set of proposal boxes that may contain objects;
[0038] Step s4: Input the proposal box obtained in step s3 and the feature map obtained in step s2 into the ROIPooling layer for maximum pooling processing, and output a set of proposal box feature maps of fixed size;
[0039] Step s5: Input the proposal box feature map obtained in step s4 into the Unknown Class Information Enhancement (UCIE) module to extract the proposal boxes containing known objects and unknown objects, and update the real boxes;
[0040] like Figure 3 ,The UCIE module contains an unknown target object score regression head, an object evaluation head, a ground truth box updater and an unknown classifier. In order to avoid confusion between known classes, unknown classes and non-object classes, we use a threshold ξ to divide non-object classes and object classes based on the object score, a threshold γ to divide known classes and unknown classes, and a threshold σ to divide unknown classes and non-object classes. This method enhances the classification and localization capabilities of classifiers and regressors for targets in open environments by adaptively updating the ground truth box;
[0041] Step s6, sending the true frame set obtained in step s5 to the classifier and regressor to perform classification and regression operations on the proposed feature map obtained in the step;
[0042] Step s7, combining steps s4, s5, and s6 to detect and learn the known objects and unknown objects in the dataset image, then calculating the loss according to the true value, and using the optimization algorithm to update the weight parameters of the target feature extraction network to reach a convergence state;
[0043] In step s8, the open world image and rectangular box annotation that have not participated in the training are taken as input, and the target feature extraction network that has converged in the training is used to perform forward operation on the input to generate a prediction result. The prediction result is sent to the RARF module, and the region-aware method is used to filter out the prediction box containing multiple unknown objects or only a small part of an unknown object to generate the final detection result.
[0044] like Figure 4 ,The RARF module consists of an NMS module, an Ncut module, a similarity judgment module, and an RA module.,In order to effectively filter the redundant frames of known and unknown objects,,we use the standard post-processing mechanism NMS to process the known,prediction frames, use Ncut to filter the similarity features of the unknown object prediction frames,,then use the object judgment module to filter the known object prediction frames,,and finally use the RA module to remove the prediction frames with too many or too few,unknown objects.
[0045] Experimental setup:
[0046] The experiment was trained for a total of 18,000 rounds. The second phase of iterative training started from 12,000 rounds, using the stochastic gradient descent (SGD) optimizer and a cosine decay learning rate scheduling strategy with an initial learning rate of 0.02 and a weight decay of 10. -4 , with a momentum of 0.9.
[0047] The experimental results are shown in Table 1-2. On the COCO-OOD dataset, our invented U-F1 score is 3.3% higher than the second-place UnSniffer, and our U-PRE score is 9.6% higher than the second-place UnSniffer. However, it also comes at a price. Our U-AP and U-RCE indicators are inferior to the best indicators. On the COCO-Mix dataset, our invented method is the best in U-AP, U-F1, and U-PRE indicators. Our U-AP indicator is 1.6% higher than the second-place ORE, our U-F1 indicator is 1.7% higher than the second-place UnSniffer, and our U-PRE indicator is 4.9% higher than it. These comparisons show that our invention is superior to existing methods in unknown object detection. By comparing Experiment 1 and Experiment 2, we can see that adding the unknown class information enhancement module can effectively improve the model's detection performance for unknown objects. Compared with the base model, our invention achieved 3.3% U-F1, 5.7% U-PRE and 0.6% U-REC improvement on the OOD dataset, and 1.4% U-AP, 1% U-F1, 0.9% U-PRE and 0.2% U-REC improvement on the Mix dataset. Comparing Experiment 1 and Experiment 3, it can be found that our method is superior to the base method in U-AP, U-F1 and U-PRE indicators. On the OOD dataset, our invention improves the model's U-AP, U-F1 and U-PRE indicators by 0.3%, 1.3% and 4.2% respectively. On the Mix dataset, the U-AP, U-F1 and U-PRE indicators are improved by 1.7%, 0.8% and 1.7% respectively. This shows that the use of region-aware unknown class redundant box filtering method is robust in open world object detection tasks and plays a positive role in improving model performance. Finally, by comparing the results of Experiment 4, it can be found that after adding the UCIE and RABF modules, most of the indicators of the model on the two data sets have reached the optimal level, which proves that the two modules have a mutually reinforcing effect and fully proves the feasibility and effectiveness of the method of the present invention. The above experiments show that the present method achieves a better balance between precision and recall.
[0048] Table 1 Comparison of the results of the method of the present invention with other methods
[0049]
[0050]
[0051] Table 2 Comparison of results with and without UCIE and RABF
[0052]
Claims
1. An open world target detection method based on region perception and unknown information enhancement, characterized in that: The method comprises: Step s1, unify the specifications of the images in the data set, set their sizes to multiples of 32, and use the mean and variance to standardize and normalize the data set data; Step s2, using the backbone feature extraction network to perform feature extraction; Step s3: Input the feature information of each layer of the feature pyramid as a feature map into the RPN to locate the target object and output a set of proposal boxes that may contain the object; Step s4, the proposal box obtained in step s3 and the feature map obtained in step s2 are processed by the ROI Pooling layer to obtain a set of proposal box feature maps of fixed size; Step s5, inputting the feature map obtained in step s4 into the Unknown Class Information Enhancement (UCIE) module to obtain a real box set containing the unknown class bounding box; Step s6, sending the true frame set obtained in step s5 to the classifier and regressor to perform classification and regression operations on the proposed feature map obtained in the step; Step s7, combining steps s4, s5, and s6 to detect and learn the known objects and unknown objects in the dataset image, then calculating the loss according to the true value, and using the optimization algorithm to update the weight parameters of the target feature extraction network to reach a convergence state; In step s8, the open world image and rectangular box annotation that have not participated in the training are taken as input, and the target feature extraction network that has converged in the training is used to perform forward operation on the input to generate a prediction result. The prediction result is sent to the Region Aware Based Unknown Redundant Box Filtering (RARF) module to filter the redundant prediction boxes of unknown objects and generate the final detection result.
2. The open world target detection method based on region perception and unknown information enhancement according to claim 1, characterized in that: In step s5, a UCIE module is proposed to enhance the unknown object information during training. The UCIE module consists of a target object score regression head, an object evaluation head, a ground truth box updater and an unknown class classifier. The target object score regression head consists of a fully connected neural network with an input of 1024 and an output of 1 and a set of linear operations. The object evaluation head uses the IOU algorithm to calculate the probability of containing a known object in the proposal box: where q i represents the prediction box, b k represents the real frame; The unknown classifier includes IOP, IOC and IOK functions. IOP and IOC extract proposal boxes that may contain unknown objects by comparing the proposal boxes with the real boxes. IOK further extracts and filters the proposal boxes by calculating the similarity of the proposal boxes. where w ij represents the horizontal distance between prediction box i and prediction box j, h ij represents the vertical distance between prediction box i and prediction box j, p j Indicates q j The target object score regression head is then input to perform object confidence evaluation to ensure that the unknown object proposal box really contains the unknown object. The real box updater uses an iterative algorithm to continuously update the bounding box information of the unknown object in the real box.
3. The open world target detection method based on region perception and unknown information enhancement according to claim 1, characterized in that: Step s8 introduces the RARF module to better filter out redundant prediction boxes on unknown objects. The module contains an NMS module, a normalized cut (Ncut) module, a similarity determination module and a region awareness (RA) module. The workflow is as follows: In the first step, the softmax function of the standard post-processing mechanism NMS is used to retain only the predicted box with the highest matching score with the relevant real box as the known class detection result; In the second step, the Ncut module is used to calculate the similarity between each unknown object proposal frame and remove the unknown object proposal frames with too high similarity; In the third step, the known proposal frames obtained in the first step and the preliminary unknown proposal frames obtained in the second step are input into the similarity determination module to evaluate the similarity of known objects, and the unknown object proposal frames with scores greater than the threshold θ are removed; In the fourth step, the unknown object proposal box obtained in the third step is input into the RA module, and the regional similarity between the proposal box and the entire image is calculated. The unknown object proposal boxes with scores higher than the threshold ρ and lower than the threshold τ are discarded to obtain a set of more accurate unknown object proposal boxes.
Citation Information
Patent Citations
Open set target detection and identification method based on deep neural network
CN114241260A
Target detection method and device, computer readable medium and electronic equipment
CN115115906A