Incremental target detection method and system based on separate treatment of amnesia

By adopting the dividing and treating amnesia strategy DCA in incremental object detection, decoupling and identifying features, and using the semantic information of the pre-trained language model and the duplex classifier fusion mechanism, the problem of forgetting imbalance in incremental object detection is solved, and better performance and plasticity are achieved.

CN120164014APending Publication Date: 2025-06-17INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510190263.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing incremental object detection methods cannot access or store historical data under privacy and security considerations, resulting in forgetting problems; at the same time, the existing methods fail to effectively decouple the degree of forgetting in positioning and identification, resulting in excessive constraints on identification features and reduced generalization of positioning.

Method used

A strategy for dividing amnesia DCA is proposed, which decouples positioning features and recognition features, focuses on solving recognition forgetting, and provides a unified optimization direction for the incremental process through the pre-trained language model encoding class semantic information, and adopts duplex classifier fusion and query role embedding mechanism.

Benefits of technology

The positioning and identification tasks are effectively decoupled, the mutual interference of forgetting is reduced, and the generalization of positioning is maintained. At the same time, the catastrophic forgetting of recognition is focused on reducing the semantic consistency between new and old tasks is optimized, and the drift of old-class features in the incremental process is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164014A_ABST
    Figure CN120164014A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of information, and relates to an increment target detection method and system based on separate treatment of amnesia. The method comprises the following steps: extracting a visual feature sequence from an input picture; inputting the visual feature sequence and the position query into a positioning decoding module for decoupling to obtain the position embedding feature of the foreground object; inputting the position embedded features into a regression device module, and predicting a bounding box of the target object; inputting the visual feature sequence, the position embedding feature and the semantic feature into a decoupled semantic-guided recognition decoding module to obtain a category embedding feature of the foreground object; and inputting the category embedded feature into a duplex classifier, and predicting the category of the target object. According to the method, the phenomenon of unbalanced forgetting in incremental target detection is discovered and demonstrated, positioning and recognition tasks in incremental detection are effectively decoupled by adopting a divide-and-conquer forgetting strategy, mutual interference between the two tasks is reduced, generalization of positioning can be kept, and disastrous forgetting of recognition is focused on being reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology, and particularly relates to an incremental object detection method and system based on divide-and-conquer amnesia. Background Art

[0002] Incremental object detection is an important research direction in the field of object detection in recent years. Its task is to enable the model to continuously and effectively detect objects of new and old categories while introducing new categories, and at the same time prevent forgetting of the learned categories. With the rapid development of deep learning technology, incremental object detection has gradually attracted attention. Inspired by transfer learning and incremental learning, the current mainstream methods are techniques such as instance replay and knowledge distillation, which achieve the ability of the model to maintain the detection ability of old categories when learning new categories. Knowledge distillation is to transfer the knowledge of the already trained old model (teacher model) to the new model (student model), where the knowledge includes detection outputs, intermediate features, and various relationships between objects, etc. Instance replay refers to storing samples or features of old category objects and using them together with new category objects for training to alleviate the forgetting of old category knowledge.

[0003] Problems existing in the existing methods are as follows:

[0004] 1. Due to privacy and security considerations, the model often cannot access or store historical data, which hinders the replay of historical information. And the existing methods rely on storing old data to retain old knowledge.

[0005] 2. The existing methods do not consider the different degrees of forgetting between localization and recognition, but distill the coupled recognition and localization features. Excessive constraints on recognition features will limit the generalization of localization, thereby reducing plasticity.

[0006] 3. In the existing methods, the semantic objectives of each task can only be optimized independently, resulting in obvious drift or coverage of old category features. Summary of the Invention

[0007] The purpose of the present invention is to find the root cause of catastrophic forgetting and formulate corresponding strategies to alleviate forgetting in incremental object detection. Through analysis, it is found that there is an imbalance in forgetting between localization and recognition, where localization is category-agnostic and has less forgetting, while recognition has serious forgetting. Therefore, in order to get rid of the mutual entanglement of forgetting, the present invention proposes a divide-and-conquer amnesia strategy DCA to decouple localization features and recognition features and focus on solving recognition forgetting. In order to reduce the drift of recognition features, the present invention uses the category semantic information encoded by the pre-trained language model (PLMs) to provide a unified optimization direction for the entire incremental process, and the mechanism for embedding semantic knowledge includes duplex classifier fusion and query role embedding.

[0008] The technical solution adopted by the present invention is as follows:

[0009] An incremental object detection method based on divide-and-conquer amnesia, comprising the following steps:

[0010] Extract a visual feature sequence from the input image;

[0011] Input the visual feature sequence and the decoupled location query into a location decoding module to obtain the location embedding features of the foreground object;

[0012] Input the location embedding features into a regressor module to predict the bounding box of the target object;

[0013] Input the visual feature sequence, the location embedding features, and the semantic features into a decoupled semantic-guided recognition decoding module to obtain the category embedding features of the foreground object;

[0014] Input the category embedding features into a duplex classifier to predict the category of the target object.

[0015] Further, the step of extracting a visual feature sequence from the input image includes: using a residual network to extract rich visual features from the input image, and then using an encoder based on an attention mechanism to capture global information to obtain an enhanced feature sequence.

[0016] Further, the decoupled location decoding module is composed of multiple transformer layers, and each transformer layer is composed of self-attention and cross-attention; input a randomly initialized location query and the visual feature sequence, the location query infers the spatial relationship between multiple targets through the self-attention mechanism, and then performs cross-attention calculation with the visual feature sequence to query and aggregate the object information in the image to obtain the location embedding features of the foreground object.

[0017] Further, the decoupled semantic-guided recognition decoding module is composed of multiple transformer layers, and each layer is composed of self-attention and cross-attention; input the location embedding, the visual features, and the semantic features of a pre-trained language model, splice the location embedding and the semantic features together, perform self-attention calculation to embed the inter-class relationship, and then perform cross-attention calculation with the visual features to obtain the category embedding features of the foreground object.

[0018] Further, the duplex classifier module is composed of a linear classifier and a classifier based on semantic features; the classifier based on semantic features calculates the cosine similarity between the category embedding features and the semantic features to obtain the semantic classification probability, and the classification probability obtained by the linear classifier and the semantic classification probability are weighted to obtain the final classification probability for predicting the category of the target.

[0019] Further, the regressor module is composed of a feed-forward layer, processes the output of the decoupled location decoding module, and predicts the bounding box of the target.

[0020] Furthermore, optimization training is carried out by calculating the detection loss and the distillation loss; the detection loss consists of the classification loss and the bounding box loss, ensuring that the model can accurately predict the category and location; the distillation loss consists of output distillation and feature distillation, and the new model mimics the output and intermediate features of the old model for the old classes to alleviate catastrophic forgetting of the old classes.

[0021] An incremental object detection system based on divide-and-conquer amnesia, comprising:

[0022] A feature extraction module for extracting a visual feature sequence from an input picture;

[0023] A decoupled localization decoding module for obtaining the position embedding feature of the foreground object according to the input visual feature sequence and position query;

[0024] A regressor module for predicting the bounding box of the target object according to the input position embedding feature;

[0025] A decoupled semantic-guided recognition decoding module for obtaining the category embedding feature of the foreground object according to the input visual feature sequence, position embedding feature, and semantic feature;

[0026] A duplex classifier module for predicting the category of the target object according to the input category embedding feature.

[0027] The beneficial effects of the present invention are as follows:

[0028] Compared with the existing methods, the present invention discovers and demonstrates the forgetting imbalance phenomenon in incremental object detection. By adopting the divide-and-conquer forgetting strategy DCA, the localization and recognition tasks in incremental detection are effectively decoupled, reducing the mutual interference between the two tasks, enabling the localization to maintain generalization, and at the same time focusing on reducing the catastrophic forgetting of recognition. The decoupled recognition features are optimized by the semantic features extracted by the entire pre-trained language model, reducing the semantic drift of the old classes during the incremental process. Experiments show that the present invention can achieve better performance in various incremental settings of existing datasets, demonstrating the practicality and effectiveness of the proposed method, especially in long-term incremental learning scenarios. Description of the Drawings

[0029] Figure 1 is the framework of DCA of the present invention.

[0030] Figure 2 is the structure of the semantic-guided recognition decoder.

[0031] Figure 3 Before decoupling, the localization feature is affected by the recognition feature and becomes sensitive to the category. The method of the present invention decouples the category-independent localization feature and focuses on solving the catastrophic forgetting at the recognition end. Detailed implementation manners

[0032] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below through specific embodiments and drawings.

[0033] The present invention proposes an incremental object detection method DCA (Dividing and Conquering Amnesia) based on divide-and-conquer amnesia. As Figure 1 shown, DCA first redesigned the decoding process of incremental detection based on DETR (an end-to-end object detection network based on Transformer) into a process of first localization and then recognition. Structurally, it retains the backbone network and transformer encoder of DETR to extract visual features, and modifies the decoder into a decoupled localization decoder and a decoupled semantic-guided recognition decoder. To balance stability and plasticity, the present invention utilizes a duplex classifier fusion module, which newly introduces a semantics-based classification head to calculate the similarity between class features and semantic features, enabling all incremental tasks to share a unified semantic space and alleviating feature space overlap and distortion. The entire model consists of five parts: a feature extraction module, a decoupled localization decoding module, a decoupled semantic-guided recognition decoding module, a duplex classifier module, and a regressor module.

[0034] The feature extraction module consists of a 50-layer residual network (i.e., ResNet50 network) and an encoder based on the attention mechanism. The residual network can extract rich visual features Ve from the input image, and the encoder uses the attention mechanism to capture global information to obtain an enhanced feature sequence for the subsequent decoding process. Among them, visual features refer to the information extracted from the image that can represent the content of the image, including low-level features (basic features such as edges and textures), middle-level features (local shapes, partial object structures, etc.), and high-level features (complete object and scene semantics). The enhanced feature sequence not only contains local visual information but also incorporates global context information, making it more suitable for the subsequent decoding process.

[0035] The decoupled localization decoding module consists of 6 layers of transformer layers, and each layer is composed of self-attention (SA) and cross-attention (CA). Input the randomly initialized position query Q local and the visual feature sequence. The position query infers the spatial relationship between multiple objects through the self-attention mechanism, and then performs cross-attention calculation with the visual feature sequence to query and aggregate the object information in the image, obtaining the position embedding feature ε local of the foreground object. Among them, object information refers to the position information of the foreground object, usually represented by a bounding box, including the center point coordinates and width and height.

[0036] The decoupled semantic-guided recognition and decoding module consists of 6 transformer layers, each layer composed of self-attention and cross-attention, as Figure 2 shown. The input is the position embedding ε local obtained in the previous step, the visual feature Ve of the image, and the semantic feature Q of the pre-trained language model (PLMs). se First, the position embedding ε local and the semantic feature Q se are concatenated together to perform self-attention calculation to embed the inter-class relationship. Then, cross-attention calculation is performed with the visual feature Ve of the image to obtain the category embedding feature ε cls of the foreground object. Figure 2 In se , ε represents the semantic feature after passing through the decoder, xL represents that the semantic-guided recognition decoder has L attention blocks, and person is an example of the known category name of the current task.

[0037] The duplex classifier module consists of a linear classifier and a classifier based on semantic features. The classifier based on semantic features calculates the cosine similarity between the category embedding feature ε cls and the semantic feature ε se to obtain the semantic classification probability, and the classification probability obtained by the linear classifier is weighted to obtain the final classification probability for predicting the category of the target.

[0038] The regressor module consists of a feed-forward layer, processes the output of the decoupled localization and decoding module, and predicts the bounding box of the target.

[0039] The entire process of the method of the present invention is divided into the following steps:

[0040] 1. The input image passes through the backbone network and the transformer encoder to extract the visual feature sequence.

[0041] 2. The visual feature sequence and the position query pass through the decoupled localization and decoding module to obtain the position embedding of the foreground object.

[0042] 3. The position embedding is fed into the regressor module to predict the bounding box of the target object.

[0043] 4. The visual feature sequence, the position embedding, and the semantic feature are input into the decoupled semantic-guided recognition and decoding module to obtain the category embedding of the foreground.

[0044] 5. The category embedding is fed into the duplex classifier to predict the category probability of the target object.

[0045] 6. The entire model is optimized and trained by calculating the detection loss and the distillation loss.

[0046] ​Among them, the detection loss consists of classification loss and bounding box loss, ensuring that the model can accurately predict the class and location. The distillation loss consists of output distillation and feature distillation, mitigating catastrophic forgetting of old classes by having the new model mimic the output and intermediate features of the old model for old classes.

[0047] The formula for the detection loss function is as follows:

[0048]

[0049] Where and are the predicted class probabilities and bounding boxes respectively, c and b are the ground truth class and bounding box respectively, represents the best match between the predicted result and the ground truth obtained using the Hungarian matching algorithm. The bounding box loss L box includes the L1 loss and the Generalized IoU loss, measuring the difference between the predicted bounding box and the ground truth bounding box.

[0050] The formula for the distillation loss function is as follows:

[0051]

[0052] Where A ij is a binary mask at the instance level, indicating whether the predicted target belongs to an old class, 1 means belonging to an old class, and 0 means not belonging to an old class; is the intermediate feature of this target predicted by the new model; is the intermediate feature of this target predicted by the old model; is the output distillation; is the feature distillation; N old is the number of pseudo-labels of old classes, represent the predicted results of the new model and the old model respectively; L mse is the MSE loss; L box includes the L1 loss and the Generalized IoU loss.

[0053] 7. Using the trained feature extraction module, decoupled localization decoding module, decoupled semantic-guided recognition decoding module, duplex classifier module, and regressor module, perform object detection on the image to be detected. The class probability of the target object obtained in step 5 and the bounding box of the target object obtained in step 3 are used together as the final object detection result.

[0054] The key points of the present invention are:

[0055] 1. Discover the forgetting imbalance phenomenon in incremental object detection based on DETR.

[0056] 2. A divide-and-conquer amnesia strategy is proposed, which decouples incremental object detection into less-forgetting localization and more-forgetting recognition to reduce mutual interference.

[0057] 3. It is proposed to embed semantic knowledge in the pre-trained language model into the recognition process to alleviate feature drift, promote joint optimization between incremental tasks with a duplex classifier, and integrate semantic knowledge into the recognition decoder in the form of queries to alleviate recognition forgetting. Finally, the network achieves better performance and has the advantage of no sample replay.

[0058] Specific application scenarios of the present invention:

[0059] The method of the present invention has important application values in scenarios such as autonomous driving and intelligent monitoring. Autonomous vehicles need to detect and recognize new objects (such as new types of vehicles, road signs, or pedestrian behaviors) in a changing environment. This method can help the vehicle system gradually learn newly emerging object categories without retraining the entire model, while maintaining the recognition ability of historical categories. For example, when new types of electric vehicles or special traffic signs appear, the system can quickly adapt and recognize them. In the intelligent monitoring scenario, the monitoring system needs to detect and recognize newly emerging abnormal behaviors or objects (such as new types of dangerous goods, suspicious behaviors) in real time. The present invention can dynamically update the model to enable it to recognize new types of threats or abnormalities. For example, in airport security checks, the system can gradually learn the characteristics of new types of dangerous goods.

[0060] Effects of the present invention:

[0061] Extensive experiments are conducted on two commonly used datasets, PASCAL VOC and MS COCO, to evaluate the effects of DCA. The PASCAL VOC dataset contains 20 foreground categories, and MS COCO contains 80 object categories. Multiple incremental learning scenarios are simulated. For the VOC dataset, three different settings are considered, namely 10+10, 15+5, and 19+1, where one group of categories (10 categories, 5 categories, and 1 category) is incrementally introduced into the detector at one time. For the COCO dataset, experiments are conducted on the settings of 70+10, 60+20, 50+30, and 40+40. For the images in the current stage, only the annotations from the current category are retained, while the annotations from past and future categories are discarded. After the training of the current stage is completed, tests are conducted on the datasets of all the categories that have been seen.

[0062] Table 1 shows the effect comparison between various modules of the model DCA of the present invention. The results show that each module brings consistent performance gains, and the modules are complementary to each other. Tables 2 and 3 respectively show the performance comparison between the present invention and other incremental methods on the VOC dataset and the COCO dataset. The present invention achieves the best results under multiple incremental settings, proving the effectiveness of the present invention. Figure 3 Visualization of the localization features and recognition features of the present invention before and after decoupling is shown. It can be found that the present invention can decouple the class-agnostic localization features from the coupled features, so as to focus on alleviating recognition forgetting.

[0063] Comparison experiment of each module in Table 1

[0064]

[0065] Comparison between DCA and other methods on the VOC dataset in Table 2

[0066]

[0067] Comparison between DCA and other methods on the COCO dataset in Table 3

[0068]

[0069] Another embodiment of the present invention provides a computer device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing each step in the method of the present invention.

[0070] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a disk, an optical disc). The computer-readable storage medium stores a computer program, and when the computer program is executed by the computer, each step of the method of the present invention is implemented.

[0071] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention is subject to the scope defined by the claims.

Claims

1. An incremental target detection method based on divide-and-conquer amnesia, characterized in that: The following steps are involved: Extract visual feature sequences from input images; The localization decoding module decouples the visual feature sequence and the position query input to obtain the position embedding feature of the foreground object; The position embedding features are fed into the regressor module to predict the bounding box of the target object; The visual feature sequence, position embedding features and semantic features are input into the decoupled semantic-guided recognition decoding module to obtain the category embedding features of the foreground object; The category embedding features are fed into the duplex classifier to predict the category of the target object.

2. The method according to claim 1, characterized in that The step of extracting a visual feature sequence from an input image includes: A residual network is used to extract rich visual features from the input image, and then an encoder based on the attention mechanism is used to capture global information to obtain an enhanced feature sequence.

3. The method according to claim 1, characterized in that The decoupled positioning decoding module is composed of multiple layers of transformer layers, each of which is composed of self-attention and cross-attention. The randomly initialized position query and visual feature sequence are input. The position query uses the self-attention mechanism to infer the spatial relationship between multiple targets, and then performs cross-attention calculation with the visual feature sequence to query and aggregate the object information in the image to obtain the position embedding feature of the foreground object.

4. The method according to claim 1, characterized in that The decoupled semantic-guided recognition decoding module consists of multiple layers of transformer layers, each of which consists of self-attention and cross-attention; position embedding, visual features and semantic features of a pre-trained language model are input, the position embedding and semantic features are spliced ​​together, self-attention calculation is performed to embed inter-class relations, and then cross-attention calculation is performed with visual features to obtain the category embedding features of foreground objects.

5. The method according to claim 1, characterized in that The duplex classifier module consists of a linear classifier and a classifier based on semantic features; the classifier based on semantic features calculates the cosine similarity between the category embedding features and the semantic features to obtain the semantic classification probability, which is weighted with the classification probability obtained by the linear classifier to obtain the final classification probability for predicting the category of the target.

6. The method according to claim 1, characterized in that The regressor module consists of a feed-forward layer that processes the output of the decoupled localization decoder module to predict the bounding box of the target.

7. The method according to claim 1, characterized in that Optimization training is performed by calculating detection loss and distillation loss; the detection loss consists of classification loss and bounding box loss to ensure that the model can accurately predict categories and positions; the distillation loss consists of output distillation and feature distillation to alleviate catastrophic forgetting of old categories by having the new model imitate the output and intermediate features of the old model for the old categories.

8. An incremental target detection system based on divide-and-conquer amnesia, characterized in that include: Feature extraction module, used to extract visual feature sequences from input images; A decoupled localization decoding module is used to obtain the position embedding features of foreground objects based on the input visual feature sequence and position query; A regressor module is used to predict the bounding box of the target object based on the input position embedding features; A decoupled semantic-guided recognition decoding module is used to obtain the category embedding features of the foreground object based on the input visual feature sequence, position embedding features and semantic features; The duplex classifier module is used to predict the category of the target object based on the input category embedding features.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.