Incremental small sample target detection method

By adopting fusion embedding processing and review and update technology in the incremental small sample object detection method, the problem of new class representation error accumulation is solved, semantic aliasing is reduced, and detection performance is improved.

CN120088609APending Publication Date: 2025-06-03HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510196379.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing incremental small sample object detection method will accumulate errors in the new class update stage, resulting in inconsistency between semantic aliasing and knowledge representation.

Method used

By acquiring the pair of supported query images to be tested and inputting them into the pre-trained detection model, the feature map of the query images to be detected and the enhanced features of the supported images are obtained. The fusion embedding processing technology is used to process the original category knowledge based on the category knowledge of the new category and the base category knowledge of the known category, and the enhanced category knowledge is updated through review and update.

Benefits of technology

Reduce semantic aliasing, enhance semantic information related to new categories and avoid error accumulation of new class representations, and improve the detection performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088609A_ABST
    Figure CN120088609A_ABST
Patent Text Reader

Abstract

The invention discloses an incremental small sample target detection method, which belongs to the technical field of computer vision, and comprises the following steps of: inputting a to-be-detected query-supporting image pair into a pre-trained detection model, and obtaining a feature map of the to-be-detected query image and enhanced features of the query-supporting image, which reduce semantic aliasing and enhance category knowledge; according to the feature map of the to-be-detected query image and the enhanced feature of the support image, a new category target on the to-be-detected query image is detected, semantic aliasing is reduced, and semantic information of the new category and category related is enhanced; the problem that in an existing incremental small sample detection method, in the new class updating stage, new class representation will accumulate mistakenly is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an incremental few-shot object detection method, belonging to the technical field of computer vision. Background Art

[0002] Incremental few-shot object detection aims to incrementally recognize previously unseen objects given very few training examples. However, the overlapping semantics between the features corresponding to the base classes and the new classes inevitably lead to semantic aliasing when having similar semantic concepts.

[0003] Most existing methods use embedding vectors or parameters to represent class knowledge, which solves semantic aliasing and drift through the knowledge representation of the new class update scheme; however, the misrepresentation of classes will iteratively accumulate because they couple knowledge learning and embedding representation in each learning stage, thus inevitably hindering the maintenance of useful and consistent knowledge learned from old classes. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies in the prior art, provide an incremental few-shot object detection method, reduce semantic aliasing and enhance the semantic information related to new categories and classes, and solve the problem that the misrepresentation of new classes will accumulate in the new class update stage of existing incremental few-shot detection methods.

[0005] To achieve the above object / To solve the above technical problems, the present invention is implemented by adopting the following technical solutions: An incremental few-shot object detection method, comprising: Obtaining a pair of to-be-tested support query images; wherein the pair of to-be-tested support query images includes a to-be-detected query image and multiple support images having the same class label as the to-be-detected query image; Inputting the pair of to-be-tested support query images into a pre-trained detection model to obtain a feature map of the to-be-detected query image and enhanced features of the support images with reduced semantic aliasing and enhanced class knowledge; Detecting new category objects on the to-be-detected query image according to the feature map of the to-be-detected query image and the enhanced features of the support images.

[0006] Further, the training of the pre-trained detection model includes: Obtaining multiple support sample images with new class labels and known class labels from a pre-constructed training set; Inputting the support sample images into the detection model to obtain the original class knowledge of the new class, the class knowledge of the new class, and the base class knowledge of the known classes in the support sample images; Performing fusion embedding processing on the original class knowledge based on the class knowledge of the new class and the base class knowledge of the known classes to obtain enhanced features; The detection model is trained with enhanced features until the number of training rounds or the target loss function meets the preset threshold.

[0007] Furthermore, the original category knowledge is subjected to fusion embedding processing based on the category knowledge of the new category and the base category knowledge of the known categories to obtain enhanced features, including: Determine the union of the category knowledge of the new category and the known categories according to the category knowledge of the new category and the base category knowledge of the known categories; Based on the union of the category knowledge of the new category and the known categories, construct the first supercategory knowledge of the new category; Perform semantic de - overlapping processing on the first supercategory knowledge of the new category to obtain the second supercategory knowledge of the new category; Perform fusion embedding processing on the original category knowledge and the second supercategory knowledge to obtain enhanced features.

[0008] Furthermore, the construction of the first supercategory knowledge of the new category based on the union of the category knowledge of the new category and the known categories includes: Determine the similarity weights between the new category and the known categories according to the union of the category knowledge of the new category and the known categories; Construct the first supercategory knowledge of the new category according to the similarity weights and the union of the category knowledge.

[0009] Furthermore, the determination of the similarity weights between the new category and the known categories according to the union of the category knowledge of the new category and the known categories includes: Calculate the similarity weights through the following formula: ; Where: c Represents the category knowledge label; Represents the P th similarity weight between the new category and the known categories, representing the P th similarity weight between the new category and the known categories; Represents the original category knowledge of the new category; Represents calculating the cosine similarity; Represents the union of the category knowledge The p th knowledge element in p= , P Represents the total number of elements in the union of the category knowledge.

[0010] Furthermore, the construction of the first supercategory knowledge of the new category according to the similarity weights and the union of the category knowledge includes: Calculate the first supercategory knowledge through the following formula: ; Wherein: represents the first superclass knowledge, h represents the superclass knowledge marker; represents the P th similarity weight between the new category and the known category, c represents the category knowledge marker; represents the union of category knowledge the p th knowledge element in it, p= , P represents the total number of elements in the union of category knowledge.

[0011] Furthermore, performing semantic de - overlapping processing on the first superclass knowledge of the new category to obtain the second superclass knowledge of the new category includes: Taking the obtained first superclass knowledge as the superclass knowledge of the known category, and calculating the second superclass knowledge of the new category through the following formula: ; Wherein: represents the second superclass knowledge, h represents the superclass knowledge marker; represents the first superclass knowledge; represents the union of the superclass knowledge of the new category and the known category; represents the L2 norm.

[0012] Furthermore, before training the detection model with enhanced features, the enhanced features need to be updated; The enhanced features include the enhanced category knowledge and enhanced superclass knowledge of the new category; Updating the enhanced category knowledge in a way of review and update to obtain the updated enhanced category knowledge; Taking the enhanced superclass knowledge and the updated enhanced category knowledge as the updated enhanced features.

[0013] Furthermore, the way of updating the enhanced category knowledge in a way of review and update to obtain the updated enhanced category knowledge includes: Obtaining the old knowledge features in the background of the support sample image; Based on the old knowledge features and the base - class knowledge of the known category, determining the update direction of the enhanced category knowledge and the safe displacement in the update direction; Updating the enhanced category knowledge according to the safe displacement to obtain the updated enhanced category knowledge.

[0014] Furthermore, the way of updating the enhanced category knowledge according to the safe displacement to obtain the updated enhanced category knowledge includes: Updating the enhanced category knowledge by using the following formula: ; Wherein: represents the updated enhanced category knowledge; represents the enhanced category knowledge of the new category; represents the i th safe displacement in the update direction of the enhanced category knowledge, , represents a total of I update directions, represents the safe displacement; ; Wherein: represents the linear transformation function for projecting the original representation into the new feature space; represents the linear transformation function for projecting the original representation into the new feature space; represents the similarity function for calculating the inner product between two vectors; represents the i th old knowledge feature in the background of the support sample image; c represents the category knowledge marker, represents the b th base class knowledge of the known category in the support sample image.

[0015] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: 1. The present invention inputs the to-be-tested support query image pair into a pre-trained detection model to obtain the feature map of the to-be-detected query image and the enhanced features of the support image that reduce semantic aliasing and enhance category knowledge. According to the feature map of the to-be-detected query image and the enhanced features of the support image, new category targets on the to-be-detected query image are detected, reducing semantic aliasing and enhancing semantic information related to the new category and the category, and solving the problem that the representation of the new category will accumulate errors during the new category update stage in the existing incremental few-shot detection method.

[0016] 2. The present invention updates the enhanced category knowledge in a way of review and update to obtain the updated enhanced category knowledge, and uses the enhanced superclass knowledge and the updated enhanced category knowledge as the updated enhanced features to train the detection model, alleviating semantic transfer and improving the model detection performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flowchart of an incremental few-shot object detection method provided by an embodiment of the present invention; Figure 2 is a training flowchart of the pre-trained detection model provided by an embodiment of the present invention; Figure 3 It is a principle framework diagram of a detection model of an incremental small sample target detection method provided by an embodiment of the present invention; Figure 4 is a schematic diagram of obtaining enhanced features in an incremental small sample target detection method provided by an embodiment of the present invention; Figure 5 It is a schematic diagram of detection results of an incremental small sample target detection method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.

[0019] In the description of the present invention, it is to be understood that the terms "first", "second", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first", "second", etc. may explicitly or implicitly include one or more of the feature. In the description of the present invention, unless otherwise specified, "plurality" means two or more. Example

[0020] like Figure 1 As shown, an incremental small sample target detection method includes: Obtain a pair of support query images to be tested; wherein the pair of support query images to be tested includes a query image to be detected and a plurality of support images having the same category label as the query image to be detected; Input the query image to be tested into the pre-trained detection model to obtain the feature map of the query image to be tested and the enhanced features of the support image that reduce semantic aliasing and enhance category knowledge. Specifically: like Figure 2 As shown, the training of the pre-trained detection model includes: Obtain multiple support sample images with new category labels and known category labels from a pre-built training set; Input the support sample image into the detection model to obtain the original category knowledge of the new category in the support sample image, the category knowledge of the new category and the base category knowledge of the known category; Based on the category knowledge of the new category and the base category knowledge of the known category, the original category knowledge is fused and embedded to obtain enhanced features, including: According to the category knowledge of the new category and the base category knowledge of the known categories, determine the union of the category knowledge of the new category and the known categories; Construct the first superclass knowledge of the new class based on the union of the class knowledge of the new class and the known classes. Among them, according to the union of the class knowledge of the new class and the known classes, determine the similarity weight between the new class and the known classes, and calculate the similarity weight through the following formula: ; c Represents the class knowledge marker; Represents the P th similarity weight between the new class and the known classes; Represents the original class knowledge of the new class; Represents calculating the cosine similarity; Represents the union of class knowledge In the p th knowledge element, p= , P Represents the total number of elements in the union of class knowledge; Calculate the first superclass knowledge through the following formula: ; Represents the first superclass knowledge, h Represents the superclass knowledge marker; Perform semantic de - overlapping processing on the first superclass knowledge of the new class to obtain the second superclass knowledge of the new class. Among them, take the obtained first superclass knowledge as the superclass knowledge of the known classes, and calculate the second superclass knowledge of the new class through the following formula: ; Represents the second superclass knowledge, h Represents the superclass knowledge marker; Represents the union of the superclass knowledge of the new class and the known classes; Represents the L2 norm.

[0021] Perform fusion embedding processing on the original class knowledge and the second superclass knowledge to obtain enhanced features; among them, the enhanced features include the enhanced class knowledge and enhanced superclass knowledge of the new class; Use the enhanced features to train the detection model until the number of training rounds or the target loss function meets the preset threshold; Detect the new class target on the to - be - detected query image according to the feature map of the to - be - detected query image and the enhanced features of the support image. Embodiment

[0022] As Figure 1 shown, an incremental few - shot object detection method includes: Obtain a support query image pair to be measured; the support query image pair to be measured includes a query image to be detected and multiple support images with the same category label as the query image to be detected; Input the support query image pair to be measured into a pre-trained detection model to obtain the feature map of the query image to be detected and the enhanced features of the support images with reduced semantic aliasing and enhanced category knowledge. Specifically: As Figure 2 shown, the training of the pre-trained detection model includes: The detection model includes a feature extraction module, a mutual perception embedding module, a knowledge update module, and a detection module. Among them, the feature extraction module is used to extract image features and knowledge, including the feature map of the query image, the original category knowledge of the new category in the support sample image, the category knowledge of the new category, the base category knowledge of the known category, and the old knowledge features in the background of the support sample image; The mutual perception embedding module is used to perform fusion embedding processing on the original category knowledge and the second super-category knowledge to obtain enhanced features; The knowledge update module is used to update the enhanced category knowledge in a way of review and update to obtain the updated enhanced category knowledge; The detection module includes a cosine classifier and a class-agnostic regressor. Among them, the class-agnostic regressor is used to locate the target object, and the class-agnostic regressor is used to classify the located object; Obtain multiple support sample images with new category labels and known category labels from a pre-constructed training set; As Figure 3 and Figure 4 shown, input the support sample image into the detection model to obtain the original category knowledge of the new category, the category knowledge of the new category, and the base category knowledge of the known category in the support sample image; Based on the category knowledge of the new category and the base category knowledge of the known category, perform fusion embedding processing on the original category knowledge to obtain enhanced features, including: According to the category knowledge of the new category and the base category knowledge of the known category, determine the union of the category knowledge of the new category and the known category; Based on the union of the category knowledge of the new category and the known category, construct the first super-category knowledge of the new category. Among them, according to the union of the category knowledge of the new category and the known category, determine the similarity weight between the new category and the known category, and calculate the similarity weight through the following formula: ; c represents the category knowledge label; represents the P th similarity weight between the new category and the known category; represents the original category knowledge of the new category; Indicates calculating cosine similarity; Indicates the union of category knowledge in the p th knowledge element, p= , P Indicates the total number of elements in the union of category knowledge; Calculate the first superclass knowledge through the following formula: ; Indicates the first superclass knowledge, h Indicates the superclass knowledge label; perform semantic de - overlapping processing on the first superclass knowledge of the new category to obtain the second superclass knowledge of the new category. Among them, take the already obtained first superclass knowledge as the superclass knowledge of the known category, and calculate the second superclass knowledge of the new category through the following formula: ; Indicates the second superclass knowledge, h Indicates the superclass knowledge label; Indicates the first superclass knowledge; Indicates the union of the superclass knowledge of the new category and the known category; Indicates the L2 norm.

[0023] Perform fusion embedding processing on the original category knowledge and the second superclass knowledge. First, send the original category knowledge and the second superclass knowledge to a pair of two - layer fully - connected (FC) layers respectively. The sigmoid activation function attached after the fully - connected layer converts the vector value into the important weight of the channel. The embeddings in the two branches are fused through element - wise multiplication, and finally, embedding fusion is used to weight the original category knowledge and the second superclass knowledge to generate enhanced features; among them, the enhanced features include the enhanced category knowledge of the new category and the enhanced superclass knowledge ; Update the enhanced category knowledge in the way of rehearsal update to obtain the updated enhanced category knowledge. Specifically: Store the enhanced category knowledge and the enhanced superclass knowledge of the new category in the memory pool of the knowledge update module and fix the strong superclass knowledge; Obtain the old knowledge features in the background of the support sample image; Based on the old knowledge features and the base - class knowledge of the known category, determine the update direction of the enhanced category knowledge and the safe displacement in the update direction; Update the enhanced category knowledge according to the safe displacement to obtain the updated enhanced category knowledge, including: Update the enhanced category knowledge using the following formula: ; Represents updated enhanced category knowledge; Represents enhanced category knowledge of a new category; Represents the safety displacement in the i th update direction of the enhanced category knowledge, , Represents a total of I update directions, Represents the safety displacement; ; Wherein: Represents a linear transformation function for projecting the original representation into a new feature space; Represents a linear transformation function for projecting the original representation into a new feature space; Represents a similarity function for calculating the inner product between two vectors; Represents the i th old knowledge feature in the background of the support sample image; c Represents a category knowledge marker, Represents the base category knowledge of the b th known category in the support sample image; Take the enhanced supercategory knowledge and the updated enhanced category knowledge as updated enhanced features, and use the updated enhanced features to train the detection model. Use the Adam optimizer, set the learning rate to 0.001, the size of a single batch in the small batch training mode is 32. During the model training, calculate the corresponding loss values of the training set and the validation set after each round of iteration. Use the model after 100 iterations as the trained detection model; As Figure 5 shown, detect the new category target on the query image to be detected according to the feature map of the query image to be detected and the enhanced features of the support image.

[0024] The above is only the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.

Claims

1. An incremental small sample target detection method, characterized in that: include: Get the supported query image pair to be tested; The support query image pair to be tested includes a query image to be detected and a plurality of support images having the same category label as the query image to be detected; Input the to-be-tested support query image pair into a pre-trained detection model to obtain a feature map of the to-be-tested query image and an enhanced feature of the support image that reduces semantic aliasing and enhances category knowledge; According to the feature map of the query image to be detected and the enhanced features of the support image, new category targets on the query image to be detected are detected.

2. The incremental small sample target detection method according to claim 1, characterized in that: The training of the pre-trained detection model includes: Obtain multiple support sample images with new category labels and known category labels from a pre-built training set; Input the support sample image into the detection model to obtain the original category knowledge of the new category in the support sample image, the category knowledge of the new category and the base category knowledge of the known category; Based on the category knowledge of the new category and the base category knowledge of the known category, the original category knowledge is fused and embedded to obtain enhanced features; The detection model is trained with enhanced features until the number of training rounds or the target loss function meets the preset threshold.

3. The incremental small sample target detection method according to claim 2, characterized in that: The process of fusing and embedding the original category knowledge based on the category knowledge of the new category and the base category knowledge of the known category to obtain enhanced features includes: According to the category knowledge of the new category and the base category knowledge of the known categories, determine the union of the category knowledge of the new category and the known categories; Based on the union of the category knowledge of the new category and the known categories, the first supercategory knowledge of the new category is constructed; Perform semantic de-overlapping processing on the first super-class knowledge of the new category to obtain the second super-class knowledge of the new category; The original category knowledge is fused and embedded with the second supercategory knowledge to obtain enhanced features.

4. The incremental small sample target detection method according to claim 3, characterized in that: The step of constructing first superclass knowledge of the new class based on the union of class knowledge of the new class and known classes includes: Determine the similarity weight between the new category and the known category according to the union of the category knowledge of the new category and the known category; According to the similarity weight and the union of category knowledge, the first supercategory knowledge of the new category is constructed.

5. The incremental small sample target detection method according to claim 4, characterized in that: The determining the similarity weight between the new category and the known category according to the category knowledge union of the new category and the known category includes: The similarity weight is calculated by the following formula: ; in: represents the normalization function; c Represents category knowledge markers; Indicates the difference between new categories and known categories. P Similarity weights; Original category knowledge representing new categories; Indicates the calculation of cosine similarity; Represents the union of category knowledge Middle p knowledge elements, p= , P Represents the total number of elements in the category knowledge set.

6. The incremental small sample target detection method according to claim 4, characterized in that: The step of constructing the first superclass knowledge of the new class according to the similarity weight and the class knowledge union includes: The first superclass knowledge is calculated by the following formula: ; in: represents the first superclass knowledge, h Indicates superclass knowledge tag; Indicates the difference between new categories and known categories. P Similarity weights, c Represents category knowledge markers; Represents the union of category knowledge Middle p knowledge elements, p= , P Represents the total number of elements in the category knowledge set.

7. The incremental small sample target detection method according to claim 3, characterized in that: The performing semantic de-overlapping processing on the first super-class knowledge of the new category to obtain the second super-class knowledge of the new category includes: The acquired first superclass knowledge is used as the superclass knowledge of the known category, and the second superclass knowledge of the new category is calculated by the following formula: ; in: represents the second superclass knowledge, h Indicates superclass knowledge tag; represents the first superclass knowledge; Represents the union of the new category with the supercategory knowledge of known categories; represents the L2 norm.

8. The incremental small sample target detection method according to claim 3, characterized in that: Before using enhanced features to train the detection model, the enhanced features need to be updated; The enhanced features include enhanced category knowledge and enhanced supercategory knowledge of the new category; The enhanced category knowledge is updated by reviewing and updating to obtain updated enhanced category knowledge; The enhanced superclass knowledge and the updated enhanced category knowledge are used as updated enhanced features.

9. The incremental small sample target detection method according to claim 8, characterized in that: The method of updating the enhanced category knowledge by reviewing and updating to obtain updated enhanced category knowledge includes: Acquire old knowledge features in the background of supporting sample images; Based on the old knowledge features and the base class knowledge of known categories, determine the update direction of enhanced category knowledge and the safe displacement in the update direction; The enhanced category knowledge is updated according to the safety displacement to obtain updated enhanced category knowledge.

10. The incremental small sample target detection method according to claim 9, characterized in that: The updating of the enhanced category knowledge according to the safety displacement to obtain updated enhanced category knowledge includes: The following formula is used to update the enhanced category knowledge: ; in: represents updated enhanced category knowledge; Enhanced category knowledge to represent new categories; The first i The safe displacement in the update direction, , Indicates shared I Update direction, Indicates safe displacement; ; in: represents the normalization function; represents the linear transformation function used to project the original representation into the new feature space; represents the linear transformation function used to project the original representation into the new feature space; represents a similarity function that computes the inner product between two vectors; Indicates that the sample image background is supported. i Gejiu knowledge characteristics; c Represents the category knowledge tag, Indicates that the sample image supports b Base class knowledge of known categories.