An Open-World Object Detection Method Based on Few-Shot Learning

By designing OFDet, combining few-sample learning and open-world object detection technology, the problem that existing object detectors cannot recognize unknown objects in category agnostic scenarios is solved, and effective positioning of unknown objects and efficient detection of new categories is achieved.

CN116229101BActive Publication Date: 2025-06-13XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310198831.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-03
Publication Date
2025-06-13
Estimated Expiration
2043-03-03

AI Technical Summary

Technical Problem

Existing object detectors cannot effectively identify and locate unknown objects in category-agnostic open-world scenarios, and their scalability in new category detection is limited.

Method used

An open-world object detection method based on few-sample learning is designed, called OFDet, which combines Faster R-CNN and OLN, including category agnostic object detection module (CALM), basic classification module (BCM), new category detection module (NDM), and unknown category selection algorithm (UPS) to realize the positioning of unknown objects and detection of new categories.

Benefits of technology

OFDet performs well on multiple known and newly set tasks, effectively identifying and positioning unknown objects, and obtaining high average accuracy and average recall on new category detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116229101B_ABST
    Figure CN116229101B_ABST
Patent Text Reader

Abstract

An open-world object detection method based on few-shot learning, belonging to the field of image processing. For the open-world object detection task based on few-shot learning, a few-shot learning method is introduced in the object detection with unknown categories, providing a small number of samples of unknown categories to guide the network to achieve the detection of new categories and the localization of unknown categories. The network OFDet for open-world object detection based on few-shot learning is modeled on an object detector with unknown categories under a two-stage fine-tuning paradigm. OFDet consists of three modules: the category-agnostic object detection module CALM, the basic classification module BCM, and the detection module NDM for new categories. To select more accurate unknown objects, a selection algorithm based on unknown candidate boxes is proposed. It performs well on multiple existing tasks, and on the newly set OFOD task, it achieves the best effect on the average recall rate of unknown categories and obtains a relatively high average precision for new categories.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and relates to traditional object detection, few-shot object detection, and class-agnostic object detection, in particular to an open-world object detection method based on few-shot learning. Background Art

[0002] The object detection task is one of the fundamental and most widely used tasks in the field of computer vision. Given an image, object detection needs to locate the object bounding boxes containing the foreground according to the network and the image content, and give their specific categories. Most existing object detectors are built on a closed-set setting, that is, the categories in the test set are completely determined by the categories used in the training process. In a more realistic scenario, these detectors cannot recognize categories not seen in the training. In contrast, humans can recognize unknown objects similar to the existing categories in a new environment, regardless of their specific categories, which has gradually attracted people's attention to class-agnostic object detection in recent years (Konan S, Liang K J, Yin L. Extending One-Stage Detection with Open-World Proposals[J]. arXiv preprint arXiv:2201.02302, 2022). As a sub-problem of open-set learning, class-agnostic object detection aims to locate all possible objects in an image without classifying them. However, due to the inability to classify, the scalability of class-agnostic object detection in more downstream tasks is limited, and it also does not have the ability to further detect unknown objects of interest.

[0003] Since such a network can be used as a robust preprocessing model in detection to preliminarily screen foreground and background, an intuitive idea is to add a detection output module (R-CNN Module) behind this model for classification and regression operations, and use the annotation information of some objects with unknown classes to fine-tune the network, expecting the model to reduce the computational cost and be more flexible and convenient during the training process. This concept involves few-shot detection (Kang B, Liu Z, Wang X, et al. Few-shot object detection via feature reweighting[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2019:8420-8429). As a two-stage closed-set detection task, a few-shot object detector can, based on a model pre-trained with base samples, use only a small number of samples of new classes to achieve fast fine-tuning and generalization of new classes in scenarios with scarce data. However, directly combining few-shot object detection with class-agnostic object detection results in poor performance in both the accuracy of base classes (bAP) and the accuracy of new classes (nAP). The possible reason is that the class-agnostic object detection task (CAOD) trains the network using binary labels and lacks multi-class classification information, leading to underfitting for the original classes and impairing the generalization performance for new classes; at the same time, the desired network also needs to identify potential unknown objects in the second stage. Therefore, the original settings do not meet the requirements of the present invention. Summary of the Invention

[0004] The purpose of the present invention is to address the limited scalability of the existing class-agnostic object detection task. By combining few-shot learning, a new visual detection task, namely open-world object detection based on few-shot learning (Open-World Object Detection, OFOD), is designed, and OFDet is proposed to implement the open-world object detection task based on few-shot learning. While detecting base classes, through the method of few-shot learning, the network can locate and identify unknown objects.

[0005] The neural network OFDet is based on Faster R-CNN and OLN, and includes a backbone network, a region proposal network, a class-agnostic object detection module CALM, a base classification module BCM, a new class detection module NDM, and an unknown class selection algorithm UPS.

[0006] An open-world object detection method based on few-shot learning for detecting and locating known and unknown objects in target images; it is divided into two stages, including the following steps:

[0007] In the first stage, train the network with sufficient base classes:

[0008] Step 1, set the input image size to H×W×C, and preset the corresponding anchor boxes Anchor for the image;

[0009] Step 2, the image passes through the backbone network to obtain the corresponding feature F b ;

[0010] Step 3, F b Passes through the Region Proposal Network (RPN) to select all possible foreground objects as positive samples, and outputs the coordinate information of the candidate boxes and the localization confidence r for calculating the loss;

[0011] Step 4, the network obtains the features F of the proposals in the corresponding regions in F b through the RoI Align operation and the regression and localization information output by the RPN p ;

[0012] Step 5, F p Passes through the Class-Agnostic Object Detection Module (CALM) to output the regression information of the finally corrected proposals and the corresponding localization confidence, and calculates the loss. F p At the same time, it passes through the Base Classification Module (BCM) to output the classification information of the proposals and calculates the loss;

[0013] In the second stage, remove the BCM module, fine-tune the network with a small number of new classes, and at the same time use the unknown class selection algorithm to locate potential unknown objects:

[0014] Steps 1-5 in the second stage are the same as the corresponding steps in stage 1;

[0015] Step 6, F p Further corrects the candidate boxes according to the regression information of CALM, and comprehensively selects the features F of the top k proposals according to the two localization confidences (r, c) output by CALM and the RPN for each proposal p* , and then sends them to the NDM module;

[0016] Step 7, F p* Passes through the New Class Detection Module (NDM) to obtain the final classification information and regression information of the proposals and calculates the loss;

[0017] Step 8, during the testing process, the UPS method is used to select potential unknown objects that are not designated as positive samples, and finally, metrics (mAP and mAR) are used to evaluate the known and unknown categories.

[0018] The present invention has the following technical effects:

[0019] 1) The present invention designs a new task, namely open-world object detection (OFOD) based on few-shot learning, and attempts to perform localization and detection operations on unknown categories and new categories respectively.

[0020] 2) The present invention proposes OFDet to implement the open-world object detection task based on few-shot learning. This network is based on a category-free object detection network and adopts a two-stage fine-tuning paradigm to solve the OFOD task. In this network model, the present invention designs three simple and practical modules (CALM, BCM, and NDM) and an unknown proposal selection algorithm (UPS) for known and unknown classes.

[0021] 3) The present invention can perform well in multiple known and newly set tasks. The present invention can not only perform well in multiple existing tasks, but also achieve the best effect on the average recall rate (AR) of unknown categories and obtain a relatively high average precision (AP) of new categories in the newly set OFOD task. Brief Description of the Drawings

[0022] Figure 1 It is a diagram illustration of the open-world object detection task OFOD based on few-shot learning.

[0023] Figure 2 It is a diagram illustration of the network OFDet designed for the open-world object detection task based on few-shot learning. Detailed Embodiment

[0024] The following embodiments will describe in detail the technical solutions and beneficial effects of the present invention with reference to the accompanying drawings.

[0025] I. Setting of the Open-World Object Detection Task OFOD Based on Few-Shot Learning

[0026] As Figure 1 shown, the designed OFOD task of the present invention is as follows:

[0027] The present invention proposes a new task, open-world object detection based on few-shot learning; given an object detection dataset D and all the categories C it contains, first divide C into base categories C_b, new categories C n and unknown categories C n , and the relationship among the three categories is C b∪C n ∪C uk =C, Based on this, D is divided into three subsets, namely a set D containing sufficient basic category C b of b , a set D containing a small number of new category C n of n and a set D containing unknown categories C without labeled information uk of uk . The OFOD task objective of the present invention is to learn a network model to detect D from D b ∪D n and detect D b and D n . Compared with the existing category-agnostic object detection (CAOD), the present invention needs to classify C b , and use a small number of categories of C n to distinguish C n and C uk ; compared with the existing few-shot object detection (FSOD), the present invention not only needs to detect C b and C n , but also needs to locate potential unknown objects C through a set algorithm uk .

[0028] II. OFDet Model Implementation Process

[0029] As Figure 2 shown is the schematic diagram of the construction of the network OFDet designed for the open-world object detection task based on few-shot learning. The construction of OFDet is conducive to the implementation of the open-world object detection (OFOD) task proposed by the present invention based on few-shot learning

[0030] 2.1 General Description

[0031] The implementation of OFDet designed by the present invention includes the following steps

[0032] As Figure 2As shown, the OFDet designed in the present invention is based on Faster R-CNN (Girshick R. Fast r-cnn[C] / / Proceedings of the IEEE international conference on computer vision. 2015: 1440-1448) and OLN (Kim D, Lin T Y, Angelova A, et al. Learning Open-World Object Proposals without Learning to Classify[J]. IEEE Robotics and Automation Letters, 2022, 7(2): 5453-5460), and refers to the two-stage paradigm in few-shot learning (Wang X, Huang T E, Darrell T, et al. Frustratingly simple few-shot object detection[J]. arXiv preprint arXiv: 2003.06957, 2020). It includes a backbone network, a Region Proposal Network (RPN), a Class-Agnostic Localization Module (CALM), a Base Classification Module (BCM), a Novel Detection Module (NDM) in the second stage, and an Unknown Proposal Selection (UPS) algorithm.

[0033] 2.2 Stage 1: Basic Training

[0034] As Figure 2 shown in Figure (a) of

[0035] The category-agnostic localization module CALM and the basic classification module BCM are introduced separately below.

[0036] 2.2.1 Category-agnostic Object Detection Module CALM

[0037] The category-agnostic object detection module (CALM) is a variant of the R-CNN detection head, which includes a RoI head for localization (L-RoI head), a candidate box regressor, and a localization confidence evaluator. Different from the existing R-CNN, CALM does not distinguish the specific categories of candidate boxes proposals, that is, category-agnostic candidate box generation. CALM is used to calculate the coordinate information of candidate boxes and the localization confidence of whether they contain foreground. It plays an important role in identifying both known and unknown objects.

[0038] 2.2.2 Basic Classification Module BCM

[0039] As mentioned above, although the category-agnostic candidate box generator (CALM) has better localization capabilities for all foreground objects, it still lacks the classification ability of basic categories. To obtain these discriminative features in the first stage and also to avoid damaging the localization of unknown objects and the classification ability of known categories, the basic classification module BCM is added additionally in the first stage to decouple the localization and classification branches. BCM consists only of a classification RoI head (C-RoI) and a basic classifier. Since the regression features have been obtained in the CALM module, no additional supervision is required in this module, so the classification features of basic categories can be directly extracted.

[0040] 2.3 Stage 2: Few-shot Fine-tuning

[0041] As Figure 2 shown in Figure (b) of , in the second stage, the C-RoI weights of BCM and the classification output layer module are used as the pre-trained weights of the novel detection module (NDM), and then removed. Also, we propose an unknown proposal selection algorithm (UPS) to locate unknown objects during the testing process.

[0042] 2.3.1 Novel Detection Module NDM

[0043] As described in the OFOD task, in downstream tasks, the detector also needs to perform further operations on the located unknown objects. Therefore, a novel detection module (NDM) is designed to detect novel categories in unknown categories using only a small amount of data. As Figure 2As shown in Figure (b), NDM is behind the CALM module and is only used in the second stage. NDM consists of a detection RoI head (D-RoI), a candidate box regressor, and a classifier. CALM is used to generate candidate boxes for localization quality, and then the candidate boxes that meet the constraints are fed into the NDM module for multi-class classification and detection. Since we have a large number of base class samples and a small number of new class samples, we can learn good separation effects for different classes, avoid retraining the network from scratch, save training computational overhead, and improve the efficiency of the model.

[0044] 2.3.2 Unknown Class Candidate Box Selection Algorithm UPS

[0045] After detecting new classes through the NDM module, the model needs to locate potential unknown objects. Therefore, an unknown class candidate box selection algorithm UPS is proposed to obtain potential unknown objects in the CALM output. Specifically, the top-k candidate boxes are selected according to the localization confidence. Among them, the candidate boxes designated as positive samples in the NDM module are used for regression and classification of new classes, while the remaining unselected ones are used as candidate boxes for potential unknown classes. Then, UPS is continued to select candidate boxes with a higher localization confidence and a lower intersection over union (IoU) with the existing positive sample candidate boxes as candidate boxes for unknown classes, and finally, an evaluation is carried out. The following gives the process of the unknown class selection algorithm:

[0046]

[0047] 2.4 Model Testing Process

[0048] As Figure 2 shown in the network architecture diagram of the model, the size of the input RGB image is set to H×W×3, where H and W are the length and width of the image respectively. During the testing process, the confidence s of the localization of each proposal in the CALM module is obtained by taking the square root of the product of the localization quality r in the RPN and the localization quality c of CALM, that is In the second stage, for detecting new classes, the network selects the top k 1 proposals according to two constraints: the confidence s and non-maximum suppression (NMS) θ, and inputs them into the NDM module as candidate boxes, where k and θ are 5000 and 0.9 respectively; for identifying unknown objects, θ 1 = 0.4, θ 2 = 1.0 and K uk = 20 are used as the hyperparameters of UPS.

[0049] III. Model Training Process

[0050] 3.1 Calculation of Two-Stage Loss Function

[0051] The loss function in the first stage:

[0052]

[0053] The total loss function of the first-stage model consists of three parts, calculating the losses of the RPN, CALM, and BCM modules respectively, where L cls is the cross-entropy loss function for the base categories, and are both L1 loss functions for the regression of category-free candidate boxes. The hyperparameter λ 1 is set to 8.

[0054] The loss function in the second stage:

[0055]

[0056] The total loss function of the second stage consists of three parts. In addition to calculating the losses of the existing RPN and CALM modules, it is also necessary to calculate the loss of the NDM module, where and are used to calculate the classification and regression losses of the new categories respectively.

[0057] 3.2 Model training parameter settings

[0058] Through gradient backpropagation, the parameters of the network can be optimized during the training process. During the training process, the SGD optimizer is used, the hyperparameter moment = 0.9, and the batch size is set to 16. In the first stage, the initial learning rate is set to 0.02, while in the second stage, the initial learning rate is set to 0.01. In both stages, the model decouples the gradients backpropagated by the RPN and RCNN backpropagation modules to the backbone. For the training of the PASCAL VOC dataset, most layers are frozen in the second stage, and only the NDM module is fine-tuned. In addition, the data augmentation method of random-lighting is used in the second stage, and the hyperparameter scale is set to 1.

[0059] 3.3 Model training

[0060] For the training of the model, first in the first stage, the input images are input into the model to obtain the output results, and formula (1) is used to calculate the loss of the model. The gradient of the loss function is backpropagated to update the model parameters in the first stage, training the classification ability of the basic model and the localization ability of all foreground objects. Then in the second stage, a small amount of data of new classes is used to fine-tune the model, and formula (2) is used to calculate the loss of the model. The gradient of the loss function is backpropagated to update the model parameters in the second stage, training the detection ability of the network for a small number of new classes. Finally, the training of the entire two-stage model is completed.

[0061] IV. Experimental Results of the Model

[0062] Table 1 shows the experimental results of few-shot detection on the VOC dataset. Table 2 shows the few-shot detection experiments on the COCO dataset. Table 3 shows the open-world object detection experiments based on few-shot learning on the COCO dataset.

[0063] Table 1

[0064]

[0065] Table 2

[0066]

[0067] Table 3

[0068]

[0069] From Tables 1 to 3, it can be concluded that: (1) OFDet proposed in the present invention is significantly superior to the corresponding baseline network TFA in terms of performance on the original task setting; (2) At the same time, OFDet achieves the best effect on the average recall rate (uAR) of unknown classes in the new task setting (open-world object detection based on few-shot learning), and at the same time obtains a relatively high average precision (nAP) of new classes. Moreover, the proposed unknown class selection algorithm (UPS) can be used as a plug-and-play method (for example, FRCN-ALL w / UPS represents that the Faster R-CNN model uses the unknown class selection algorithm proposed in the present invention), helping the original few-shot detector to identify unknown objects, and the recall rate based on unknown classes is increased to 100%.

[0070] Specifically:

[0071] Table 1 shows the experimental evaluation results of OFDet for few-shot object detection on the VOC dataset; in Table 1, the bold numbers respectively represent the better performance under single-group settings (the same below). Under the three data partition sets (Novel Set 1, Novel Set 2, Novel Set 3), OFDet proposed by the present invention is superior to the baseline network TFA and most few-shot detectors under most settings. Specifically, for Novel Set 1, under the new class number settings of K = 1, 2, 3, 5, 10, the average precision of new classes (nAP 50 ) is improved by 3.2%, 4.9%, 4.9%, 8.4% and 3.3% compared with TFA respectively; similarly, on Novel Set 2 and Novel Set 3, nAP 50 exceeds TFA by 3.9%, 0.7%, 3.9%, 4.0% and 5.6%, and 4.8%, 9.0%, 0.8%, 1.9%, 4.7% respectively. It shows that the model proposed by the present invention has better generalization on new classes without adding too much computation and special modules.

[0072] Table 2 shows the experimental evaluation results of OFDet for few-shot object detection on the COCO dataset; the nAP of the OFDet model proposed by the present invention can reach 9.6%, 12.7%, 17.5% under the settings of 5-shot, 10-shot and 30-shot, and is improved by 2.2%, 2.7% and 3.8% compared with TFA respectively. At the same time, OFDet also has better performance than few-shot detectors based on the meta-learning paradigm under most settings. It shows that the method proposed by the present invention has stronger robustness and better generalization in more complex scenarios.

[0073] Table 3 shows the experimental evaluation results of OFDet for open-world object detection based on few-shot learning on the COCO dataset. OFDet is compared with multiple baseline models, including FRCN-ALL, TFA, and MetaR-CNN. Among them, FRCN-ALL represents the strategy of not using the corresponding frozen weight layer and directly fine-tuning the Faster R-CNN model trained with base classes. For the division of the number of new classes in the COCO dataset, referring to the settings of few-shot detection tasks, 10-shot and 30-shot are selected. As shown in Table 3, uAR represents the recall rate of unknown classes, and Unknown Set represents the set of different unknown classes. Although multiple existing baseline detectors can detect new classes, they are all unable to locate unknown objects under the settings of 10-shot and 30-shot. The unknown class picking algorithm (UPS) proposed in the present invention can well assist OFDet. To further verify the effectiveness of the method proposed in the present invention, the unknown class picking algorithm is applied to these models, which are named FRCN-ALL w / UPS, TFA w / UPS, and Meta R-CNN w / UPS in the table respectively. Specifically, the present invention selects candidate regions with relatively high confidence but not designated as positive samples from the Region Proposal Network (RPN) module of these models and applies the unknown class picking algorithm to achieve the recall of unknown classes. With the assistance of the unknown class picking algorithm, the uAR of these baseline models has been greatly improved. Finally, compared with the methods of all these existing models, the proposed OFDet achieves the best performance in the recall evaluation uAR of unknown classes. Among them, the uAR of OFDet under the 10-shot setting can reach 30.2%, 30.1%, and 24.4%; the uAR of OFDet under the 30-shot setting can reach 32.5%, 31.2%, and 24.5%. At the same time, OFDet also maintains quite excellent performance in the precision evaluation nAP of new classes, which can reach 11.1%, 11.2%, 12.6% and 15.7%, 16.7%, and 17.1% respectively in the corresponding settings. The results in Table 3 show that the method proposed in the present invention can be applied to open-world scenarios and can achieve a good balance in the tasks of detecting new classes and locating potential objects.

[0074] The above embodiments are only used to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any modifications made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the present invention.

Claims

1. An open-world object detection method based on few-shot learning, characterized in that it is used to detect and locate known and unknown objects in the target image; it is divided into two stages, including the following steps: In the first stage, the network is trained with sufficient base classes: Step 1, set the input image size to H×W×C, and preset the corresponding anchor boxes Anchor for the image; Step 2, the image passes through the backbone network to obtain the corresponding feature F b ; Step 3, F b Pass through the Region Proposal Network (RPN) to select all possible foreground objects as positive samples, output the coordinate information of the candidate boxes and the localization confidence r for calculating the loss; Step 4, the network obtains the features F of the proposals in the corresponding region in F through the RoI Align operation and the regression and localization information output by the RPN b ; p ; Step 5, F p Through the class-agnostic object detection module CALM, a variant of the R-CNN detection head, including a RoI head (L-RoI head) for localization, a bounding box regressor, and a localization confidence evaluator, calculates the coordinate information and the localization confidence c of the bounding box, outputs the regression information and the localization confidence c of the corrected proposals, and uses the L1 loss function to calculate the loss of class-agnostic bounding box regression and Meanwhile, F p Passes through the basic classification module BCM, which consists only of a classification RoI head (C-RoI) and a basic classifier, decouples the localization and classification branches, outputs the classification information of proposals, and calculates the loss L of the basic category using the cross-entropy loss function cls ; In the second stage, the BCM module is removed, and the network is fine-tuned with a small number of new classes. At the same time, an unknown class selection algorithm is used to locate potential unknown objects; Steps 1-4 in the second stage are the same as the corresponding steps in the first stage; Step 5, F p Output the regression information of the corrected proposals and the localization confidence c through the CALM module, and use the L1 loss function to calculate the loss of the regression of the candidate boxes without classes and At the same time, a new class detection module NDM is added behind the CALM module, which consists of a detection RoI head (D-RoI), a candidate box regressor, and a classifier, for multi-class classification and detection, and the C-RoI weights of the BCM and the classification output layer module are used as the pre-trained weights of the NDM module. Subsequently, the BCM is removed from the network; Step 6, F p Further correct the candidate boxes according to the regression information of CALM, and comprehensively select the features F of the top k proposals according to the two localization confidence levels (r, c) of each proposal in CALM and the RPN output p* , and then send them to the NDM module; Step 7, F p* Through the new category detection module NDM, obtain the classification information and regression information of the final proposals, and calculate the new category classification loss and regression loss Step 8, during the test process, the potential unknown objects in the non-positive samples are selected through the unknown class candidate box selection algorithm UPS. Finally, the known and unknown classes are evaluated using the metrics mAP and mAR. The two localization confidences (r, c) output by CALM and RPN for each proposal and the non-maximum suppression (NMS) θ are used to comprehensively select the top k1 proposals to obtain the potential unknown objects in the CALM output.