A target detection method based on incremental learning

CN117132785BActive Publication Date: 2026-09-25BEIHANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202311140472.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-06
Publication Date
2026-09-25
Estimated Expiration
2043-09-06

AI Technical Summary

Technical Problem

在现有技术中的增量检测任务中存在着严重的标签丢失的问题

Benefits of technology

(1)本发明基于CLIP模型进行增量检测目标图像,提高了模型对任意的新词汇的推理能力,在增量检测模型分类决策面上对齐了具有全局信息的语言嵌入空间,获得了更好的分类面,提高了模型在未来能够更好地纳入新任务类别特征;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117132785B_ABST
    Figure CN117132785B_ABST
Patent Text Reader

Abstract

The application relates to an incremental object detection method based on an incremental object detection model IODC and belongs to the technical field of image recognition. The incremental object detection model IODC is obtained by fusing a global perception category text model and a visual model based on a CLIP model, is used for incremental learning of pictures, can recognize unknown categories of objects in the pictures, improves the recognition rate of the pictures, and solves the problems of label loss, inability to detect potential objects in the target and long training period of a detection experiment in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, specifically to an incremental learning-based target detection method and an incremental detection model IODC. Background Technology

[0002] The real-world visual system is inherently incremental; people need to learn new knowledge through observation and integrate it into their existing visual knowledge base. While deep learning has achieved remarkable success in object detection tasks, it becomes extremely forgetful in incremental learning scenarios and suffers from catastrophic forgetting, causing a sharp decline in model performance for older tasks.

[0003] At present, the two technical challenges that need to be addressed in incremental detection models include: (1) Incremental detection models cannot identify potential unknown category objects in images. In the incremental learning stage, the training images in the incremental classification task must be non-overlapping, such as... Figure 1 As shown, however, in incremental detection tasks, the same image may have completely different annotation information, which is quite different from the dataset settings of incremental classification. Traditional detection models cannot identify categories other than the current task; (2) Traditional detectors are based on convolutional neural network frameworks such as Faster RCNN, and their classification layers can only output the categories covered in the current category set. For a category that has never appeared clearly, its classification prediction will always be 0. Even if the label of the unknown class object is given, it is still impossible to distinguish the sub-classes in a category, which hinders the refinement of the new class task of the incremental target.

[0004] Currently, Chinese patent CN113822368A proposes an anchorless incremental object detection method that improves the detection performance of new class test images under training with a large amount of richly labeled base class data (images) and a small number of labeled small-shot new classes. Chinese patent CN115546581A proposes a decoupled incremental object detection method that adds a channel-level decoupling module before RCNN and RPN respectively, enabling the backbone network to learn more generalizable and transferable features. The additional parameters provided by the decoupling module further enhance the model's learning ability, thus improving the trade-off between plasticity and stability. Wang et al. proposed a two-stage training method using Faster-RCNN as a framework, where only the classification and regression subnetworks are fine-tuned in the second stage, and the combined weights of features are readjusted to adapt to the novel class. Juan-Manuel et al. proposed a method borrowing from the CenterNet framework, introducing a feature extractor for image feature extraction, a target locator for object localization, and a ResNet-50 network to extract the weights corresponding to the image output of each class and use these weights to complete the detection of new classes.

[0005] Therefore, current technologies tend to improve the incremental characteristics of incremental detection models to further detect difficult targets in images, while controlling training time and resource costs. Most techniques neglect forward compatibility in incremental learning tasks, failing to allow the model to adapt to the incremental learning scenario beforehand. Existing incremental detection technologies suffer from severe label loss issues.

[0006] With the rapid development of the internet age, internet applications are active in all aspects of people's lives, and various sectors of society have accumulated massive amounts of data, which are also growing rapidly. How to quickly extract useful information from unlabeled or poorly labeled data and rapidly adapt it to existing models is a very challenging problem. New, unrecognizable object classes are more likely to appear in images trained for recognition; therefore, improving the forward compatibility of incremental detection models is crucial at this stage. Summary of the Invention

[0007] In view of the above problems, this invention provides an incremental learning-based object detection method and an incremental detection model IODC. Based on the CLIP model, the incremental object detection model IODC is obtained by fusing a globally perceptive category text model and a visual model. It is used for incremental learning of images and can identify objects of unknown categories in images, thereby improving the model's recognition rate. This solves the problems of label loss, failure to detect potential objects in the target, and long training cycles in detection experiments that occur in the incremental detection of existing technologies.

[0008] This invention provides a target detection method based on incremental learning, comprising: Step 1: Identify the category names of multiple objects in the training image, establish text features based on the category names of the objects, collect multiple text features to construct a text feature set, obtain the visual features of the image based on the image-text detection model, and establish a global perception category text feature model of the image based on the text feature set and visual features. Step 2: Based on the category text feature model described in Step 1, construct a visual model that can identify objects of unknown category in the image; Step 3: Integrate the global perception category text feature model described in Step 1 and the visual model described in Step 2 to establish an incremental object detection model IODC, and identify potential objects in the current task based on the incremental object detection model IODC.

[0009] Preferably, the image and text detection model in step 1 is the CLIP model; the visual features include: color features, shape features, texture features, and line features.

[0010] Preferably, step 1, establishing a globally aware category-based text feature model, specifically includes: Step 11: Identify the category names of all objects covered in multiple training images, construct a category text sentence for each identified category name, and establish a text feature set based on the fusion of language modalities of all category text sentences; Step 12: Input the training image described in Step 11 into the CLIP model, modify the output of the CLIP model classification layer to visual features, obtain the visual features of the image, perform similarity calculation on each visual feature and each text feature, match the visual feature with the highest similarity to the text feature to form a category element, and all category elements form a category set. Step 13: Add a generalized category to the category set described in Step 12, defining it as a generalized category set. Perform incremental training on the incremental detection model based on the generalized category set to obtain an updated incremental detection model. Step 14: Use the updated incremental detection model to detect new categories contained in the training images, establish a category mapping relationship between the generalized categories and the new categories, and build a globally perceptive category text feature model based on the category mapping relationship and the set of generalized categories. The new category is defined as a category that is not found in the generalized categories when the updated incremental detection model detects the object categories in the training images. Further, the category text feature sentence in Step 11 includes: sentence template and category name.

[0011] Furthermore, step 11, which involves constructing categorized text feature sentences and establishing a text feature set based on all categorized text sentences, specifically includes: Identify the category names of all objects covered in multiple training images; A language model is trained based on language modality information to obtain an updated language model, and a sentence template is constructed based on the updated language model. Each category name is entered into a sentence template to obtain a text sentence representing the category name; All text sentences are fed into the text encoder of the CLIP model to generate multiple text features, and a text feature set is established; the category name sentence template is "there is a {classname} in the scene".

[0012] The technical solution of this invention uses the sentence template "there is a {classname} in the scene", which is better adapted to the detection task. That is, the target object is usually located in a part of the scene. By utilizing the information of language modality, a better global feature classification space is obtained through a pre-trained language model. This makes the incremental detection model more predictable and malleable.

[0013] Furthermore, the visual features of the image described in step 12 and the category text features both have the same dimension, which is batchsize×512.

[0014] Furthermore, the similarity calculation in step 12 specifically includes: Input an image into the object detection model and use a detection network to obtain the visual features of the objects covered in the image. The visual features and the text features described in step 11 are normalized, and then the cosine similarity method is used to calculate the similarity between the normalized visual features and the text features. Obtain the predicted class probability logical values ​​of multiple objects in the image; input the cross-entropy loss function into the predicted class probability logical values ​​to obtain the classification loss; The similarity is corrected based on the classification loss to obtain a better similarity.

[0015] Furthermore, obtaining a globally perceptual category text feature model for recognizing image objects specifically includes: During the initial training phase, custom generalized categories are added based on the category set described in step 12, and defined as a generalized category set; the generalized category set includes categories that are not in the category set. The incremental detection model is trained based on the generalized category set to obtain an incremental detection model with global classification features; A dataset of new tasks is constructed by pre-setting multiple incremental training tasks for new classes; the new tasks The new task dataset is identified on the incremental detection model of the global classification features to identify multiple new task categories. The similarity between the multiple new task categories and multiple general categories in the generalized category set is calculated. The new task category with the highest similarity is mapped one-to-one with the generalized category, which is represented as the category mapping relationship between the generalized category and the new task category; Based on the category mapping relationship, the new task category is mapped to the position of the generalized category, generating text features of the new category. A globally perceptive category text feature model is established based on the text features of the new category and the set of generalized categories.

[0016] The technical solution of this invention incrementally trains an incremental object detection model with global classification features, thereby transferring the framework of object detection with global classification features into the incremental detection field. This allows the incremental detection model to enjoy the advantages of both fields. It can not only better alleviate the catastrophic forgetting problem from the global feature space and avoid the situation where traditional classification layers cannot predict unknown categories, but also enable the model to have a weak ability to predict unknowns in the early stages of learning, so as to better cope with future new category learning tasks.

[0017] This invention's technical solution uses the CLIP model to obtain category text features and then utilizes similarity calculation for classification, solving the problem of terminal neurons being suppressed and inactivated in traditional classification methods. By extracting category names from the task to obtain semantic information of the text, the traditional incremental object detection model aligns its classification decision surface with the language embedding space containing global information, thereby obtaining a better classification surface. This allows the model to better incorporate category features of new tasks that need to be identified in the future.

[0018] The technical solution of this invention can not only better alleviate the catastrophic forgetting problem from the global feature space and avoid the situation where the classification layer of the traditional incremental object detection model cannot predict unknown categories, but also enable the model to have a weak zero-shot object detection capability in the early learning stage, so as to better cope with future new class learning tasks.

[0019] Preferably, step 2, which involves constructing a visual model to identify unknown categories in an image, specifically includes: Step 21: Generate a set of candidate target boxes for training images based on the incremental detection model, obtain the predicted category of the candidate target boxes according to the classification module of the model, filter out the target boxes that are displayed as background category in the predicted category, and construct a set of background target boxes. Step 22: Input the set of background target boxes described in Step 21 into the CLIP model, and obtain the predicted category and prediction score of each background target box based on the visual encoder of the CLIP model; Step 23: Based on the predicted category and predicted score described in Step 22, detect whether there is an object displaying an unknown category in each background target box; If the predicted category of the target box is an unknown category and the predicted score exceeds the threshold, it is determined that there is an unknown object in the target box, which is characterized by the presence of an unknown object in the background information of the target box. Step 24: Based on the set of generalized categories, assign generalized category labels to the target boxes containing unknown objects in the background information one by one, and correct them to generalized category target boxes respectively; Step 25: Iterate through and calculate the confidence between each generalized category target box, and summarize the overlapping target boxes with high confidence, defining them as pseudo-target boxes; filter the size of the pseudo-target boxes, obtain updated pseudo-target boxes, and establish an updated pseudo-target box set; Step 26: Merge the updated pseudo-target box set described in Step 25 into the candidate target box set described in Step 21 to obtain the updated target box set; construct a visual model that can identify unknown category text features in the image based on the updated target box set and the text feature set described in Step 1.

[0020] Furthermore, step 26, which describes constructing a visual model capable of recognizing text features of unknown categories in an image, involves the following specific steps: The updated target box set is fused with the text feature set, so that each text feature in the text feature set corresponds to each target box in the updated target box set, thereby obtaining a visual model that can identify unknown categories of text features in an image.

[0021] Furthermore, when the incremental detection model identifies training images, the updated target box set can assign pseudo-target boxes to target boxes of unknown categories in the images for identification, and then match them with the text feature set to further identify potential objects.

[0022] Furthermore, the predicted category in step 22 is the text feature category and visual feature category corresponding to the background target box; The predicted score is the probability that the CLIP model identifies the text feature category corresponding to the background target box. Assuming that the background target box set only contains target boxes of the three categories of cat, dog and pig, then the predicted score is the probability that the background target box is classified as a cat, the probability that it is classified as a dog, and the probability that it is classified as a pig, and the sum of the three probabilities is 1.

[0023] Preferably, the potential object mentioned in step S3 is, for example, an image with a cat and a dog. In the initial learning stage, if the cat is taken as the target object, the potential object is the dog; in the subsequent incremental learning stage, if the dog is taken as the target object, the potential object is the cat.

[0024] The visual model acquired by the technical solution of this invention combines the capabilities of the CLIP model, providing explicit supervision signals for unknown categories. This helps the model correct background information of target boxes without any annotation information into generalized category target boxes. The technical solution of this invention adds an additional set of generalized category labels to the dataset, which can help the model better capture instances of unknown categories. By selecting a small number of target boxes through random sampling, the incremental detection model becomes aware of the existence of unknown categories in the image without requiring very precise localization information, allowing the model to learn the features of unknown categories in advance.

[0025] Compared with the prior art, the present invention has at least the following beneficial effects: (1) This invention uses the CLIP model to incrementally detect target images, which improves the model’s reasoning ability for any new words. It aligns the language embedding space with global information on the classification decision surface of the incremental detection model, obtains a better classification surface, and improves the model’s ability to better incorporate new task category features in the future. (2) The present invention obtains the category text model based on similarity calculation, which avoids the problem of terminal neurons being suppressed and inactivated in traditional classification methods; (3) This invention utilizes information from the language modality to build a text model, thereby obtaining a better global feature classification space, which makes the incremental detection model more predictable and adaptable. (4) This invention alleviates the catastrophic forgetting problem from the global feature space and avoids the situation where traditional classification layers cannot predict unknown categories; (5) The present invention obtains a visual model for recognizing unknown categories in images, which can both identify objects of unknown categories and locate objects of unknown categories, thereby improving the ability of incremental target detection models to perceive unknown categories. Attached Figure Description

[0026] The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of the invention.

[0027] Figure 1 This is a schematic diagram illustrating the label loss problem in the incremental detection task of this invention; Figure 2 This is a schematic diagram illustrating the knowledge transfer from a generalized category to a new category in the category mapping method of the present invention; Figure 3 A schematic diagram illustrating the construction of a set of category name sentences for this invention; Figure 4 This is a schematic diagram illustrating the identification of potential unknown objects in the background of an image according to the present invention; Figure 5 This is a schematic diagram of the flowchart for generating pseudo-target boxes using the CLIP model visual encoder of this invention. Detailed Implementation

[0028] To better understand the above-described objectives, features, and advantages of the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other. Furthermore, the present invention can be implemented in other ways different from those described herein; therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.

[0029] To illustrate the effectiveness of the method proposed in this invention, the following detailed description of the above technical solution is provided through a specific embodiment. The specific implementation steps are as follows: A specific embodiment of the present invention, such as Figure 1-5 A target detection method based on incremental learning is disclosed, including: Step 1: Identify the category names of multiple objects in the training image, establish text features based on the category names of the objects, collect multiple text features to construct a text feature set, obtain the visual features of the image based on the image-text detection model, and establish a global perception category text feature model of the image based on the text feature set and visual features. Step 2: Based on the category text feature model described in Step 1, construct a visual model that can identify objects of unknown category in the image; Step 3: Integrate the global perception category text feature model described in Step 1 and the visual model described in Step 2 to establish an incremental object detection model IODC, and identify potential objects in the current task based on the incremental object detection model IODC.

[0030] Preferably, the image and text detection model in step 1 is the CLIP model; the visual features include: color features, shape features, texture features, and line features.

[0031] Preferably, step 1, establishing a globally aware category-based text feature model, specifically includes: Step 11: Identify the category names of all objects covered in multiple training images, construct a category text sentence for each identified category name, and establish a text feature set based on the fusion of language modalities of all category text sentences; Step 12: Input the training image described in Step 11 into the CLIP model, modify the output of the CLIP model classification layer to visual features, obtain the visual features of the image, perform similarity calculation on each visual feature and each text feature, match the visual feature with the highest similarity to the text feature to form a category element, and all category elements form a category set. Step 13: Add a generalized category to the category set described in Step 12, defining it as a generalized category set. Perform incremental training on the incremental detection model based on the generalized category set to obtain an updated incremental detection model. Step 14: Use the updated incremental detection model to detect new categories contained in the training images, establish a category mapping relationship between the generalized categories and the new categories, and establish a global perception category text feature model based on the category mapping relationship and the set of generalized categories; the new category is defined as a category that is not found in the generalized categories when the updated incremental detection model is used to detect the object categories in the training images. That is, the category that was not identified in the detection task in the generalized categories.

[0032] Furthermore, the category text feature sentences mentioned in step 11 include: sentence template and category name.

[0033] Furthermore, the specific steps for constructing a category text feature sentence as described in step 11 are as follows: Identify the category names of all objects covered in multiple training images; A language model is trained based on language modality information to obtain an updated language model, and a sentence template is constructed based on the updated language model. Each category name is entered into a sentence template to obtain a text sentence representing the category name; All text sentences are fed into the text encoder of the CLIP model to generate multiple text features, and a text feature set is established; the category name sentence template is "there is a {classname} in the scene".

[0034] The technical solution of this invention uses the sentence template "there is a {classname} in the scene", which is better adapted to the detection task. That is, the target object is usually located in a part of the scene. By utilizing the information of language modality, a better global feature classification space is obtained through a pre-trained language model. This makes the incremental detection model more predictable and malleable.

[0035] Furthermore, the visual features of the image described in step 12 and the category text features both have the same dimension, which is batchsize×512.

[0036] Furthermore, the similarity calculation in step 12 specifically includes: Input an image into the object detection model and use a detection network to obtain the visual features of the objects covered in the image. The visual features and the text features described in step 11 are normalized, and then the cosine similarity method is used to calculate the similarity between the normalized visual features and the text features. Obtain the predicted class probability logical values ​​of multiple objects in the image; input the cross-entropy loss function into the predicted class probability logical values ​​to obtain the classification loss; The similarity is corrected based on the classification loss to obtain a better similarity.

[0037] Furthermore, obtaining a globally perceptual category text feature model for recognizing image objects specifically includes: During the initial training phase, custom generalized categories are added based on the category set described in step 12, and defined as a generalized category set; the generalized category set includes categories that are not in the category set. The incremental detection model is trained based on the generalized category set to obtain an incremental detection model with global classification features; A dataset of new tasks is constructed by pre-setting multiple incremental training tasks for new classes; the new tasks The new task dataset is identified on the incremental detection model of the global classification features described in step 132, and multiple new task categories are identified. The similarity between the multiple new task categories and multiple general categories in the general category set is calculated. The new task category with the highest similarity is mapped one-to-one with the generalized category, which is represented as the category mapping relationship between the generalized category and the new task category; Based on the category mapping relationship, the new task category is mapped to the position of the generalized category, generating text features of the new category. A globally perceptive category text feature model is established based on the text features of the new category and the set of generalized categories.

[0038] The technical solution of this invention incrementally trains an incremental object detection model with global classification features, thereby transferring the framework of object detection with global classification features into the incremental detection field. This allows the incremental detection model to enjoy the advantages of both fields. It can not only better alleviate the catastrophic forgetting problem from the global feature space and avoid the situation where traditional classification layers cannot predict unknown categories, but also enable the model to have a weak ability to predict unknowns in the early stages of learning, so as to better cope with future new category learning tasks.

[0039] This invention's technical solution uses the CLIP model to obtain category text features and then utilizes similarity calculation for classification, solving the problem of terminal neurons being suppressed and inactivated in traditional classification methods. By extracting category names from the task to obtain semantic information of the text, the traditional incremental object detection model aligns its classification decision surface with the language embedding space containing global information, thereby obtaining a better classification surface. This allows the model to better incorporate category features of new tasks that need to be identified in the future.

[0040] The technical solution of this invention can not only better alleviate the catastrophic forgetting problem from the global feature space and avoid the situation where the classification layer of the traditional incremental object detection model cannot predict unknown categories, but also enable the model to have a weak zero-shot object detection capability in the early learning stage, so as to better cope with future new class learning tasks.

[0041] Preferably, step 2, which involves constructing a visual model to identify unknown categories in an image, specifically includes: Step 21: Generate a set of candidate target boxes for training images based on the incremental detection model, obtain the predicted category of the candidate target boxes according to the classification module of the model, filter out the target boxes that are displayed as background category in the predicted category, and construct a set of background target boxes. Step 22: Input the set of background target boxes described in Step 21 into the CLIP model, and obtain the predicted category and prediction score of each background target box based on the visual encoder of the CLIP model; Step 23: Based on the predicted category and predicted score described in Step 22, detect whether there is an object displaying an unknown category in each background target box; If the predicted category of the target box is an unknown category and the predicted score exceeds the threshold, it is determined that there is an unknown object in the target box, which is characterized by the presence of an unknown object in the background information of the target box. Step 24: Based on the generalized category set described in step 131, assign generalized category labels to the target boxes containing unknown objects in the background information one by one, and correct them to generalized category target boxes respectively; Step 25: Iterate through and calculate the confidence between each generalized category target box, and summarize the overlapping target boxes with high confidence, defining them as pseudo-target boxes; filter the size of the pseudo-target boxes, obtain updated pseudo-target boxes, and establish an updated pseudo-target box set; Step 26: Merge the updated pseudo-target box set described in Step 25 into the candidate target box set described in Step 21 to obtain the updated target box set; construct a visual model that can identify unknown category text features in the image based on the updated target box set and the text feature set described in Step 1.

[0042] Furthermore, step 26, which describes constructing a visual model capable of recognizing text features of unknown categories in an image, involves the following specific steps: The updated target box set is fused with the text feature set, so that each text feature in the text feature set corresponds to each target box in the updated target box set, thereby obtaining a visual model that can identify unknown categories of text features in an image.

[0043] Furthermore, when the incremental detection model identifies training images, the updated target box set can assign pseudo-target boxes to target boxes of unknown categories in the images for identification, and then match them with the text feature set to further identify potential objects.

[0044] Furthermore, the predicted category in step 22 is the text feature category and visual feature category corresponding to the background target box; The predicted score is the probability that the CLIP model identifies the text feature category corresponding to the background target box. Assuming that the background target box set only contains target boxes of the three categories of cat, dog and pig, then the predicted score is the probability that the background target box is classified as a cat, the probability that it is classified as a dog, and the probability that it is classified as a pig, and the sum of the three probabilities is 1.

[0045] Preferably, the potential object mentioned in step S3 is, for example, an image with a cat and a dog. In the initial learning stage, if the cat is taken as the target object, the potential object is the dog; in the subsequent incremental learning stage, if the dog is taken as the target object, the potential object is the cat.

[0046] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A target detection method based on incremental learning, characterized in that, Includes the following steps: Step 1: Identify training images, construct a text feature set for the training images, obtain the visual features of the images, and establish a global perceptual category text feature model for the images based on the text feature set and visual features. Specifically, this includes: Step 11: Identify the category names of all objects covered in multiple training images, construct category text sentences for each identified category name, and establish a text feature set by fusing language modalities based on all category text sentences; Step 12: Input the image described in Step 11 into the incremental detection model to obtain the visual features of the image. Calculate the similarity between each visual feature and each text feature. Match the visual feature with the highest similarity to the text feature to form a category element. All category elements constitute a category set. Step 13: Add a generalized category to the category set described in Step 12, defining it as a generalized category set. Perform incremental training on the incremental detection model based on the generalized category set to obtain an updated incremental detection model. Step 14: Use the updated incremental detection model to detect new task categories, establish a category mapping relationship between generalized categories and new task categories, and establish a global perception category text feature model based on the category mapping relationship and the set of generalized categories; Step 2: Construct a visual model for identifying unknown category objects in images based on a globally perceptive category text feature model, specifically including: Step 21: Generate a set of candidate target boxes for training images based on the incremental detection model, obtain the predicted category of the candidate target boxes according to the classification module of the model, filter out the target boxes that are displayed as background category in the predicted category, and construct a set of background target boxes. Step 22: Input the set of background target boxes described in Step 21 into the CLIP model, and obtain the predicted category and prediction score of each background target box based on the visual encoder of the CLIP model; Step 23: Based on the predicted category and predicted score described in Step 22, detect whether there is an object displaying an unknown category in each background target box; If the predicted category of the target box is an unknown category and the predicted score exceeds the threshold, it is determined that there is an unknown object in the target box, which is characterized by the presence of an unknown object in the background information of the target box. Step 24: Based on the generalized category set, assign generalized category labels to the target boxes containing unknown objects in the background information one by one, and correct them to generalized category target boxes respectively; Step 25: Iterate through and calculate the confidence between each generalized category target box, and summarize the overlapping target boxes with high confidence, defining them as pseudo-target boxes; filter the size of the pseudo-target boxes, obtain updated pseudo-target boxes, and establish an updated pseudo-target box set; Step 26: Merge the updated pseudo-target box set described in Step 25 into the candidate target box set described in Step 21 to obtain the updated target box set; construct a visual model that can identify unknown category text features in the image based on the updated target box set and the text feature set described in Step 1; Step 3: Integrate the global perception category text feature model and the visual model to establish the incremental object detection model IODC, and identify potential objects in the current task image based on the incremental object detection model IODC.

2. The target detection method based on incremental learning according to claim 1, characterized in that, Step 11, which involves constructing categorized text sentences and establishing a text feature set based on all categorized text sentences, specifically includes: Identify the category names of all objects covered in multiple training images; A language model is trained based on language modality information to obtain an updated language model, and a sentence template is constructed based on the updated language model. Each category name is entered into a sentence template to obtain a text sentence representing the category name; All text sentences are fed into the text encoder of the CLIP model to generate multiple text features, thus establishing a text feature set.

3. The target detection method based on incremental learning according to claim 2, characterized in that, The sentence template is "there is a {classname} in the scene".

4. The target detection method based on incremental learning according to claim 3, characterized in that, Step 12, the similarity calculation, specifically includes: Input an image into the object detection model and use a detection network to obtain the visual features of the objects covered in the image. The visual features and the text features described in step 11 are normalized, and then the cosine similarity method is used to calculate the similarity between the normalized visual features and the text features. Obtain the predicted class probability logical values ​​of multiple objects in the image; input the cross-entropy loss function into the predicted class probability logical values ​​to obtain the classification loss; The similarity is corrected based on the classification loss to obtain a better similarity.

5. The target detection method based on incremental learning according to claim 4, characterized in that, The global perceptual category text feature model for obtaining image objects specifically includes: During the initial training phase, custom generalized categories are added based on the category set described in step 12, and defined as a generalized category set; the generalized category set includes categories that are not in the category set. The incremental detection model is trained based on the generalized category set to obtain an incremental detection model with global classification features; Incremental training is performed on the incremental detection model with global classification features to obtain the new class task of incremental training, and multiple new class tasks are collected to construct a new class task dataset; The new task dataset is identified on the incremental detection model with global classification features to identify multiple new task categories. The similarity between the multiple new task categories and multiple general categories in the general category set is calculated. The new task category with the highest similarity is mapped one-to-one with the generalized category, which is represented as the category mapping relationship between the generalized category and the new task category; Based on the category mapping relationship, the new task category is mapped to the position of the generalized category, generating text features of the new category. A globally perceptive category text feature model is established based on the text features of the new category and the set of generalized categories.

6. The target detection method based on incremental learning according to claim 5, characterized in that, The predicted categories in step 22 are the text feature category and the visual feature category corresponding to the background target box.

7. The target detection method based on incremental learning according to claim 6, characterized in that, The predicted score in step 22 is the probability that the CLIP model correctly identifies the text feature category corresponding to the background target box.

Citation Information

Patent Citations

  • Anchor-free incremental target detection method

    CN113822368A

  • Decoupled incremental target detection method

    CN115546581A

  • Target detection method and device based on incremental learning

    CN113205142A

  • Target detection model training method, target detection method and related equipment thereof

    CN113469176A