A confidence-driven incremental object detection method
Through the confidence-driven incremental object detection method, using the old model to generate pseudo-labels and confidence values, the catastrophic forgetting problem in single-stage detectors is solved, the old category detection capability is maintained, the calculation cost is reduced, and it is suitable for large-scale incremental learning.
Patent Information
- Application Number
- CN202411942834.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-12-27
AI Technical Summary
The prior art is difficult to effectively solve the catastrophic forgetting problem in incremental target detection in single-stage detectors, especially when introducing new categories, which leads to the model's detection ability of old category targets, and the calculation cost is high and the integration is difficult.
The confidence-driven incremental object detection method is used to generate area-level pseudo-labels and pixel-level confidence values through the old model, and filter the pseudo-labels with the non-maximum suppression method, and merge them with the real labels to correct the background area of the new model, and distillate them using the confidence loss function to maintain the detection ability of the old category, and learn the new category at the same time.
It effectively avoids background offset problems, maintains high detection accuracy, reduces calculation costs, is suitable for large-scale incremental learning scenarios, and is suitable for single-stage detectors.
Smart Images

Figure CN119380003B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a confidence-driven incremental target detection method. Background Art
[0002] Visual object detection is a fundamental task in computer vision. It plays a vital role in many fields, such as autonomous driving, multimedia information retrieval, and scene understanding. Thanks to the progress of deep learning, high detection performance can be achieved on both one-stage and two-stage detectors. For all these detectors, they are trained on a set of labeled samples with a fixed number of object categories. However, generally, new training samples from known or unknown categories are gradually added as the application scenarios gradually change. In order to adapt the trained model to the new scenario, there are two possible options, retraining the model from scratch or gradually updating the model. However, retraining the model from scratch requires a lot of computing resources, and the old model samples are difficult to obtain due to high maintenance costs, privacy issues, and intellectual property issues. Therefore, a way to gradually update the trained model should be considered.
[0003] Assuming that the newly added samples are from known categories, the model update can be done by simple fine-tuning. However, when the new samples come from new target categories, the problem becomes challenging. Simply fine-tuning the trained model will lead to a deterioration in the model's ability to recognize targets of old categories. This problem is widely known as the "catastrophic forgetting" problem. In fact, since most neural network designs assume that data is provided once, this problem is widespread in various deep neural networks. Suppose there is a deep learning task that is divided into multiple stages, and the network model is trained at each stage. The EWC (Elastic Weight Conservation) method selectively maintains the network weights that are important to the trained model in the previous stage, while allowing other parts to adapt to the new task. In essence, this is an incremental learning method based on regularization. However, experiments show that EWC does not perform as well as those based on knowledge distillation on incremental object detection.
[0004] In recent research, incremental object detection has been addressed through various knowledge distillation strategies, most of which are proposed based on two-stage detectors. However, due to the background shift problem, although the old model and the new model are aligned at different levels through different knowledge distillation losses, they still conflict with each other. Therefore, it is difficult for the new model to balance maintaining the knowledge of the old model and adapting to new target categories. In addition, since it is impossible to distill on candidate boxes, this distillation strategy is not feasible for single-stage detectors. Moreover, different from two-stage detectors, single-stage detectors share a localization head for all categories (both old and new categories). When distilling directly on the network output, the conflict between the new and old models becomes inevitable. In view of this, it is urgent to study an efficient incremental object detection method for single-stage detectors. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a confidence-driven incremental object detection method, which is applicable to the object detection scenario of gradually introducing new categories for training, effectively avoiding the background shift problem, reducing the computational cost at the same time, and being easy to integrate.
[0006] The present invention is implemented as follows: A confidence-driven incremental object detection method, the method comprising:
[0007] S1. Use the old model to generate region-level pseudo-labels for the new image to label the regions of the old category target objects in the new image;
[0008] S2. Combine the ground truth labels of the new image with the generated pseudo-labels to form an extended detection target category annotation for training the new model;
[0009] S3. Generate pixel-level confidence values for the new image through the old model and merge them with the extended labels to correct the background regions in the new model;
[0010] S4. Train the new model, and through region-level pseudo-labels and pixel-level confidence value distillation, maintain the detection of the old categories while learning the new categories.
[0011] Further, the step S1 specifically includes:
[0012] Step S11. Input the new image into the pre-trained old model to generate predicted , , , where represents the predicted bounding box, represents the confidence, represents the category;
[0013] Step S12: Use the non - maximum suppression method NMS to screen the pseudo - labels and obtain the screened prediction boxes. and the category , and keep the screened prediction boxes and the category as the pseudo - labels of the new image. Filter out unreliable region predictions by using non - maximum suppression.
[0014] Furthermore, the NMS threshold and the confidence threshold of the non - maximum suppression method NMS are set to 0.6 and 0.1 respectively.
[0015] Furthermore, the specific content of step S2 is as follows:
[0016] Step 21: Merge the pseudo - labels of the new image with the true labels to form an extended detection target category annotation, that is, the enhanced annotation for the new model training. The enhanced annotation is expressed as:
[0017]
[0018]
[0019] where and represent the bounding box and the category in the true label of the new image respectively. The and represent the bounding box and the category after the enhanced annotation;
[0020] Step 22: Change the location loss function and the classification loss function. The location loss function is updated to:
[0021]
[0022] The classification loss function is updated to:
[0023]
[0024] where represents the prediction box output by the new model for the new image, represents the category output by the new model for the new image, , , and c represent the hyperparameters in the loss function respectively. represents the intersection - over - union ratio of the prediction region and the true region, and n represents the number of samples. Specifically, represents the Euclidean distance between the center points of the predicted bounding box and the true bounding box, c represents the diagonal length of the smallest circumscribed rectangle containing the predicted bounding box and the true bounding box, α represents the weight hyperparameter, and v represents the consistency parameter of the aspect ratio of the predicted bounding box and the true bounding box.
[0025] Further, step S3 is specifically as follows:
[0026] By calculating and the intersection over union between them to generate a confidence score, and defining the confidence loss formula as:
[0027]
[0028]
[0029]
[0030]
[0031] Wherein, represents the confidence output by the new model for the new image, represents the predicted bounding box output by the new model for the new image, represents the class output by the new model for the new image, n represents the number of samples, represents the confidence after enhanced annotation, and the represents the indicator function, represents the confidence output by the old model;
[0032] After that, correct the network loss function and use the corrected network loss function to train the new model. The network loss function is expressed as:
[0033] .
[0034] Further, the new model is a single-stage detector and performs object detection tasks by sharing a localization head.
[0035] Further, the method further includes the step of gradually introducing different classes in multi-stage incremental object detection, wherein the pseudo-labels generated by the old model are used for all stages.
[0036] The present invention has the following advantages: By combining region-level pseudo-labels and pixel-level confidence values for distillation, the present invention can effectively solve the catastrophic forgetting problem in incremental object detection without replaying old-class samples. By generating pseudo-labels and combining them with the true labels of new classes, it ensures that the new model can retain the detection ability for old-class target objects, and corrects the judgment of the new model for the background region by correcting the confidence of the background region, preventing old-class target objects from being misjudged as the background in new images and avoiding the background shift problem. While maintaining high detection accuracy, the present invention has a low computational cost and is applicable to large-scale incremental learning scenarios. Brief Description of the Drawings
[0037] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0038] Figure 1 It is a flowchart of the execution of an incremental object detection method based on confidence-driven according to the present invention.
[0039] Figure 2 It is a framework diagram of an incremental object detection method based on confidence-driven according to the present invention. Specific embodiments
[0040] The technical solutions in the embodiments of the present application generally have the following ideas: Aiming at the problems existing in the prior art, the present invention uses distillation strategies at the region level and pixel level to ensure that the new model can learn new categories and avoid loss of detection performance for old categories. Similar to most incremental object detection methods, the update of the new model is based on knowledge distillation. The distillation at the region level is achieved through the pseudo-labels of old objects. To solve the background shift problem in the case where pseudo-labels are not available, another distillation is achieved through the confidence score map. Specifically, when performing knowledge distillation at the region level, the old model is used to predict the old category target objects in the new image to generate pseudo-labels, non-maximum suppression is used to filter unreliable predictions, and the pseudo-labels are combined with the true labels of the new category to form enhanced annotations; when performing distillation at the pixel level, the confidence scores generated by the old model are used to correct the background region judgment of the new model to avoid the background shift problem; the incremental object detection framework proposed by the present invention is applicable to single-stage detectors and has wide applicability; compared with existing incremental object detection methods, the present invention has the advantages of low computational cost and easy integration, and is applicable to large-scale incremental object detection scenarios.
[0041] To better understand the above technical solutions, the technical solutions of the present invention will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.
[0042] Please refer to Figure 1 and 2 As shown, the present invention provides an incremental object detection method based on confidence-driven, and the method includes:
[0043] S1. Use the old model to generate region-level pseudo-labels for the new image to mark the regions of old category target objects in the new image;
[0044] S2. Combine the true labels of the new image with the generated pseudo-labels to form an extended detection target category annotation for the training of the new model;
[0045] S3. Generate pixel-level confidence values for the new image through the old model and merge them with the extended labels to correct the background regions in the new model;
[0046] S4. Train the new model, and maintain the detection of old classes while learning new classes through region-level pseudo-labels and pixel-level confidence value distillation.
[0047] Preferably, the step S1 specifically includes:
[0048] Step S11. Input the new image into the trained old model to generate predicted , , , where represents the predicted bounding box, represents the confidence, represents the class;
[0049] Step S12. Use the non-maximum suppression method NMS to screen the pseudo-labels to obtain the screened predicted bounding box and class . Retain the screened predicted bounding box and class as the pseudo-labels of the new image, and filter out unreliable region predictions by using non-maximum suppression.
[0050] Preferably, the NMS threshold and confidence threshold of the non-maximum suppression method NMS are set to 0.6 and 0.1 respectively.
[0051] Preferably, the step S2 is specifically:
[0052] Step 21. Merge the pseudo-labels of the new image with the ground truth labels to form an augmented detection target class annotation, that is, the enhanced annotation for training the new model. The enhanced annotation is expressed as:
[0053]
[0054]
[0055] where and respectively represent the bounding box and class in the ground truth labels of the new image, and the and represent the bounding box and class after enhanced annotation;
[0056] Step 22. Change the location loss function and classification loss function. The location loss function is updated to:
[0057]
[0058] The classification loss function is updated to:
[0059]
[0060] where Denotes the predicted bounding box output by the new model for the new image, Denotes the class output by the new model for the new image, , , And c respectively represent the hyperparameters in the loss function, Denotes the intersection over union (IoU) between the predicted region and the ground truth region, and n represents the number of samples. Specifically, Denotes the Euclidean distance between the centers of the predicted bounding box and the ground truth bounding box, c represents the diagonal length of the smallest closed rectangle (i.e., the smallest bounding rectangle) containing the predicted bounding box and the ground truth bounding box, and α represents a weight hyperparameter used to balance the center point distance With other terms in the CIoU loss. It is usually a dynamic value that depends on the value of IoU to ensure that when IoU is high, the center point distance term has a greater impact on the loss, Denotes the consistency of the aspect ratio between the predicted bounding box and the ground truth bounding box. This term penalizes the case where the aspect ratios of the predicted bounding box and the ground truth bounding box are inconsistent.
[0061] Preferably, the step S3 is specifically as follows:
[0062] By calculating And To generate the confidence score by calculating the intersection over union between them, and the confidence loss formula is defined as:
[0063]
[0064]
[0065]
[0066]
[0067] Where n represents the number of samples, Denotes the confidence output by the new model for the new image, Denotes the predicted bounding box output by the new model for the new image, Denotes the class output by the new model for the new image, Denotes the confidence after enhanced annotation, and the Denotes the indicator function, Denotes the confidence output by the old model;
[0068] After that, the network loss function Is corrected, and the new model is trained using the corrected network loss function. The network loss function is expressed as:
[0069] .
[0070] Preferably, the new model is a single-stage detector, and performs object detection tasks by sharing a localization head.
[0071] Preferably, the method further includes the step of gradually introducing different classes in incremental object detection in multiple stages, where the pseudo-labels generated by the old model are used for all stages.
[0072] To better understand the above technical solution, the definitions involved in the incremental object detection process will be described in detail below.
[0073] First, the definition of general incremental object detection is as follows:
[0074] Set the object set as C, and these classes are gradually introduced into the detection model from the set C. The subset is introduced into the detection model at the t-th stage. For any , if , then . Let represent the image set containing annotated object instances from different classes in . The image ∈ can contain object instances from both new and old classes at the same time. However, for the old-class object instances, that is, , no annotations are provided. At the same time, replaying the image is not allowed.
[0075] The goal of incremental object detection is to build a series of detection models . Assuming that the current training stage is t, the student model (i.e., the new model) is trained based on the image set , given the trained model . Since the teacher model (i.e., the old model) is expected to perform well on all observed classes , only the teacher model is needed when training the student model . When t = 2, the incremental detection is usually called single-stage incremental detection; when t > 2, it is called multi-stage incremental detection.
[0076] For ease of understanding, the typical single-stage detector YOLOv7 is used for analysis. The key components of YOLOv7:
[0077] YOLOv7 gives the logits bounding boxes of the model , the confidence and the class , and their corresponding ground truths are respectively , and The training objective of YOLOv7 is to minimize the bounding box regression loss defined on the model logits , the object confidence loss and the positive sample classification loss . The network loss is defined as:
[0078]
[0079] The bounding box regression loss , the object confidence loss and the positive sample classification loss are expressed by the following formulas:
[0080]
[0081] In the above loss functions, is the loss between the predicted box and the ground truth bounding box calculated after positive sample matching, and are hyperparameters in the loss functions. Similarly, the confidence loss and the classification loss are calculated by cross-entropy.
[0082] 1. When performing the regional-level distillation of the present invention:
[0083] 1.1 Generating pseudo-labels for the input image
[0084] Assume that the old model can handle all observed classes well . The old-class target objects appearing in the image ∈ can be regarded as a convenient medium for transferring knowledge from to the new model . Basically, the old model is used as an automatic annotator to generate labels for the old classes on the image . By inputting the image ∈ into the model , we get:
[0085]
[0086] At this time, the generated predicted values specify all potential regions where old-class target objects may appear. To filter out unreliable predictions, we further adopt the non-maximum suppression (NMS) method After screening, and in the NMS operation, the NMS threshold and the confidence threshold are set to 0.6 and 0.1 respectively.
[0087] 1.2 Combine the pseudo - labels with the ground - truth labels to generate enhanced annotations
[0088] The predicted regions are retained as pseudo - labels on the image These pseudo - labels are merged with the ground - truth labels provided by the image to form enhanced annotations for the dataset After all images are processed in this way, the enhanced annotations can be expressed as:
[0089]
[0090] Since these enhanced annotations will be used to train the model . The location loss and the classification loss functions are modified to and . Since the labels of the old classes participate in the training, the presence of the old - class objects no longer interferes with the training of the new model. Instead, the new model will be compatible with the old - class objects.
[0091] 2 When performing the region - level distillation of the present invention:
[0092] The above - mentioned region - level distillation makes full use of the old - class objects that appear in the new sample images, and basically solves the background shift problem. However, these objects do not necessarily cover all the observed classes. In extreme cases, there may be no old - class objects in the new sample images. Therefore, the new model will inevitably lose the detection ability for these missing classes. As the number of training stages increases, this situation will gradually deteriorate.
[0093] For two - stage detectors, it is easy to perform direct distillation on the network output. However, for single - stage detectors such as YOLO, this is basically infeasible because the new and old classes share the same localization head . If we choose to distill the feature maps, we will face a scenario similar to the background shift. Therefore, other knowledge distillation methods from to are adopted and are applicable to single - stage detectors.
[0094] 2.1 Generate confidence scores
[0095] By the old model on the image The confidence scores generated above largely compensate for the deficiencies of pseudo-labels, especially in cases where pseudo-labels are not available. More importantly, these scores can be obtained in both single-stage and two-stage detectors. Therefore, in addition to pseudo-labels, we also distill confidence scores from to However, if we directly inherit the entire confidence score map of the old model it will be superimposed on the confidence scores generated by the enhanced annotations. To solve this problem, the confidence scores are generated by calculating the IoU between and since the enhanced annotations are available at this time. To transfer the confidence distillation from to the confidence loss formula is modified as:
[0096]
[0097] where I is the indicator function. The ground truth of the confidence score is obtained by calculating the intersection over union of the bounding box with the ground truth When it means there is an object at the current position. When it means the pixel belongs to the background. In this case, the confidence score generated by will be directly assigned to In this way, the knowledge in is distilled into We first mark the enhanced annotation area as the target area ( ), and then replace the background area ( ) with the confidence scores generated by the old model to correct the background area.
[0098] 2.2 Modify the network loss function
[0099] The modified confidence loss integrates region-level and pixel-level distillation. When we substitute into YOLOv7, its network loss function is modified as:
[0100]
[0101] Through the above region-level pseudo-label and pixel-level confidence value distillation, the detection of old classes is maintained during the training process of the new model, while learning new classes. The method of the present invention is simple and efficient, and is applicable to the confidence distillation of single-stage detectors. The knowledge of the teacher model is distilled at two levels. At the region level, Generate pseudo-labels and merge them with the real labels to generate enhanced real labels . At the pixel level, the confidence scores are required to be aligned with the confidence scores . To combine with region-level distillation, the regions corresponding to the enhanced real labels are retained. Since both region-level and pixel-level distillations are integrated into the confidence scores, our method is also called Confidence Score Distillation (CSD).
[0102] Next, experiments are conducted on the performance of the method (CSD) of the present invention:
[0103] In the first experiment, experiments are carried out on the currently popular object detection test dataset Pascal-VOC 2007, and the experimental results are shown in Table 1 below:
[0104] Table 1
[0105]
[0106] In this experiment, we verified the effectiveness of our method on single-stage learning tasks. The method CSD of the present invention, i.e., Confidence Score Distillation, was integrated into the single-stage detector YOLOv7 for testing. Similarly, other methods OnlineOD, EWC, and KD were also integrated into YOLOv7 respectively. Table 1 shows the performance results on VOC 2007. Three single-stage incremental learning settings were tested, namely 10+10, 15+5, and 19+1. The first 19, 15, and 10 categories in the training set were regarded as old categories respectively, corresponding to the three evaluation settings, and the remaining 1, 5, or 10 categories were regarded as new categories. It can be seen from Table 1 that the performance of CSD on YOLOv7 is significantly better than other methods.
[0107] In the second experiment, we studied the effectiveness of the method CSD of the present invention when integrated into the YOLOv7 model in multi-stage incremental object detection tasks. As shown in Table 2 below, the object categories in the Pascal VOC 2007 and COCO 2017 datasets were evenly divided into four groups: A, B, C, and D. The incremental training was divided into four stages, namely: starting from stage A, new categories were gradually added in stage B, and new object categories were gradually put in at stage C and stage D in this way for training. It can be seen from Table 2 that on the Pascal VOC 2007 and COCO 2017 datasets, the best performance comes from the CSD method based on YOLOv7. On the one hand, this is largely due to the excellent performance of YOLOv7, providing a good foundation for the incremental detector; on the other hand, this also shows that the distillation strategy we proposed successfully transfers the old knowledge to the new model.
[0108] Table 2
[0109]
[0110] Through experiments on the Pascal VOC 2007 and COCO 2017 datasets, it is verified that while maintaining the detection performance of new categories, the method of the present invention can significantly improve the detection accuracy of old categories compared with existing methods. Moreover, the present invention has strong adaptability in multi-stage incremental learning and has good feasibility and effectiveness for object detection scenarios in practical applications.
[0111] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.
Claims
1. An incremental object detection method based on confidence-driven, characterized in that: The method includes: S1. Using an old model to generate region-level pseudo-labels for a new image to label the regions of old-class target objects in the new image; S2. Combining the ground truth labels of the new image with the generated pseudo-labels to form an augmented detection target class annotation for training the new model; S3. Generating pixel-level confidence values for the new image through the old model and merging them with the augmented detection target class annotation to correct the background regions in the new model; S4. Training the new model, and through distillation of region-level pseudo-labels and pixel-level confidence values, maintaining the detection of old classes while learning new classes; The specific steps of step S1 include: Step S11: Input the new image into the trained old model to generate predicted , , , where represents the prediction box, represents the confidence level, represents the category; Step S12: Use the non-maximum suppression method NMS to screen the pseudo-labels to obtain the screened prediction boxes and the category , and retain the screened prediction boxes and the category as the pseudo-labels of the new image, and filter out unreliable regional predictions by using non-maximum suppression; The specific step of step S3 is: By calculating and to generate a confidence score by calculating the intersection over union between them, and define the confidence loss formula as: where n represents the number of samples, represents the confidence level output by the new model for the new image, represents the predicted bounding box output by the new model for the new image, and the and represent the bounding box and category after enhanced annotation, represents the confidence level after enhanced annotation, represents the confidence score, and the represents the indicator function, represents the confidence level output by the old model, I n represents the image; Confidence score The ground truth value is obtained by calculating the intersection over union (IoU) between the bounding box and the ground truth . When , it indicates that there is an object at the current position, and the enhanced annotation area is marked as the target area; when , it means that the pixel belongs to the background area, and the background area is replaced with the confidence score generated by the old model to correct the background area. In this case, the confidence score generated by the old model will be directly assigned to . In this way, the knowledge in the old model is distilled into the new model After that, the network loss function is corrected, and the new model is trained using the corrected network loss function. The network loss function is expressed as: 。 2. The method for incremental object detection based on confidence-driven according to claim 1, wherein: The NMS threshold and confidence threshold of the non-maximum suppression method NMS are set to 0.6 and 0.1 respectively.
3. A confidence-driven incremental object detection method according to claim 1, characterized in that: The specific step of step S2 is: Step 21. Merging the pseudo-labels and ground truth labels of the new image to form an augmented detection target class annotation, that is, the enhanced annotation for training the new model, and the enhanced annotation is expressed as: Among them, and respectively represent the bounding box and category in the true label of the new image, and the and represent the bounding box and category after enhanced annotation; Step 22. Changing the location loss function and classification loss function, and the location loss function is updated to: The classification loss function is updated to: Among them, represents the predicted bounding box output by the new model for the new image, represents the category output by the new model for the new image, represents the intersection over union of the predicted region and the ground truth region, represents the Euclidean distance between the centers of the predicted bounding box and the ground truth bounding box, c represents the diagonal length of the smallest bounding rectangle containing the predicted bounding box and the ground truth bounding box, α represents the weight hyperparameter, v represents the consistency parameter of the aspect ratio of the predicted bounding box and the ground truth bounding box, and n represents the number of samples.
4. A confidence-driven incremental object detection method according to claim 1, characterized in that: The new model is a single-stage detector and performs object detection tasks through a shared localization head.
5. A confidence-driven incremental object detection method according to claim 1, characterized in that: The method further includes the step of gradually introducing different classes in multi-stage incremental object detection, where the pseudo-labels generated by the old model are used in all stages.
Citation Information
Patent Citations
Aerial target detection incremental learning method and device and medium
CN118521818A