A Zero-Shot Object Detection Method Based on DETR and Meta-Learning

By fusing semantic vectors in the DETR architecture and performing class-by-category optimal matching loss function optimization, the problem of low recall and category confusion of class objects in zero-sample object detection is solved, and a high-accuracy zero-sample object detection is achieved.

CN116958741BActive Publication Date: 2025-07-11FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310832459.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-07
Publication Date
2025-07-11
Estimated Expiration
2043-07-07

AI Technical Summary

Technical Problem

The existing zero-sample object detection algorithm has low recall in the detection of unseen objects and is easy to confuse background and unseen objects. Especially in the detection of endangered animals, training samples are difficult to obtain. Traditional methods require a large amount of calibration data.

Method used

The zero-sample object detection method based on the DETR architecture is fused into the query vector, and the detection results are directly decoded through category-by-category optimal matching and loss function optimization, and the DETR detector of the transformer architecture is trained.

Benefits of technology

It improves the detection recall of objects of unknown class, reduces the confusion between background classes and unknown class, and achieves high accuracy detection without training samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958741B_ABST
    Figure CN116958741B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of machine learning, and specifically relates to a zero-shot object detection method based on DETR and meta-learning. The method of the present invention is constructed based on the DETR detector of the transformer architecture. A zero-shot learning mechanism is introduced into the DETR deep object detection framework, and the class semantic vector is directly incorporated into the query vector of DETR, and the result is directly predicted by the decoder. During the training process, the training is completed through optimal matching and loss calculation for each category. The method framework of the present invention is simple, easy to use, highly scalable, and highly interpretable. The results of zero-shot object detection on mainstream visual attribute datasets show that the performance of this method is significantly better than existing methods. The present invention provides algorithm support for object detection technology in the industrial application field and can also be easily extended to other zero-shot learning tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning, and particularly relates to an object detection method based on DETR and meta-learning. Background Art

[0002] Object detection technology is a basic task in computer vision tasks, which aims to locate and classify object category objects in images. The application scope of object detection technology is extensive, and it provides basic support for some downstream tasks, such as instance segmentation, scene understanding, pose estimation and other tasks. Existing deep object detection models have achieved good accuracy in some categories, but they rely heavily on large-scale calibrated datasets. However, in real-world scenarios, problems such as unbalanced data sample distribution and unsupervised samples are faced. Therefore, how to perform effective object detection without training samples has become an open problem in the field of computer vision. Zero-shot learning provides a systematic framework for solving this type of problem, that is, by using a large amount of visible class data and semantic vectors for training, predictions can be made on unseen classes only relying on unseen class semantic vectors without any training data of unseen classes.

[0003] DETR is an object detection algorithm based on the transformer architecture. Due to the successful application of the transformer (a neural network structure using the attention mechanism) architecture, the current best-performing object detection algorithms are all modified based on DETR, such as DINO, Deformable DETR, etc.

[0004] Most existing zero-shot object detection algorithms are modified based on the Faster-RCNN object detection framework, and these methods all have the following limitations:

[0005] (1) The candidate boxes generated by RPN are difficult to cover unseen class objects, resulting in a low recall rate for unseen classes;

[0006] (2) Since the background class of the second-stage classifier is trained on visible classes, unseen class objects and the background class are prone to confusion.

[0007] The application scenarios of zero-shot object detection are numerous. The present invention focuses on the field of endangered animal detection. For endangered animals, it is often difficult to obtain their training samples, resulting in the inability of traditional object detection models to detect these endangered animals. Zero-shot object detection only needs to provide the semantic vectors of these animals to detect them. Summary of the Invention

[0008] The purpose of the present invention is to provide a zero-shot object detection method based on DETR and meta-learning with high detection accuracy.

[0009] Zero-shot Object Detection Problem Definition: In zero-shot object detection, the dataset is divided into visible classes and unseen classes. The visible classes are used for training, and the unseen classes are used for test verification. At the same time, a semantic vector is provided for each class in the visible and unseen classes, and this semantic vector is a description of the class. In existing methods, most use the RPN network of Faster RCNN to generate candidate boxes, and then use cosine similarity to calculate the similarity between the candidate box features and the class features to complete the classification of the candidate boxes. This invention is based on the DETR series of detectors, fuses the semantic vector into the query vector, and directly decodes to obtain the result.

[0010] The zero-shot object detection method based on DETR and meta-learning provided by this invention has the following specific steps:

[0011] S1, Sample image data and class semantic vectors; For images, in the training set, randomly sample a training image I. For class semantic vectors, first randomly sample the classes included in image I and randomly sample the classes not included in image I and together constitute the class set At the same time, image I undergoes feature extraction through the backbone network and the transformer encoder to obtain the feature x of image I I .

[0012] S2, Perform a linear mapping on the semantic vectors corresponding to the class set; For the class set the corresponding semantic vectors are denoted as The semantic vectors are mapped through the linear layer to obtain the mapped semantic vectors That is:

[0013]

[0014] S3, Add the mapped semantic vectors and the object query vectors of DETR; Add the mapped semantic vectors and the object query vectors of DETR to obtain the query vector after fusing the semantic vectors That is:

[0015]

[0016] S4, Decode the query vector; Input the query vector after fusing the semantic vectors into the decoder g of DETR θ to obtain the predicted calibration box result That is:

[0017]

[0018] Benefits of the above design: Traditional DETR can only detect the categories that appear in the training set. Then, for a query vector, it outputs a prediction result, and the query vector here is category-independent, that is, for each query vector, it can output the bounding box predictions of any category. In the present invention, the query vector is fused with the category semantic vector, turning the category-independent query vector into a category-specific query vector. After this query vector is fused with a category, this query vector is responsible for detecting this category, only predicting the position of the bounding box of the fused-in category, and the probability that the predicted bounding box belongs to this category, that is, the confidence. In this way, it is possible to detect any category by fusing the semantic vectors of any category, even if this category does not appear in the training set.

[0019] To enable the model to achieve the above purpose, the design of the loss function is very important:

[0020] For the loss function of the traditional DETR detector for general object detection, the traditional DETR detector performs optimal bipartite matching, that is, it matches the predicted bounding box and the ground truth bounding box with the minimum cost, and optimizes based on the matching result. The matching process of the traditional DETR series detectors can be expressed as where, is the classification loss, is the localization loss. Here, the classification loss and the localization loss are added together to represent the cost under this matching. is a permutation of N numbers. Here, it means to find a permutation to minimize the subsequent cost, that is, to minimize the total loss. Among them, is the ground truth bounding box, c represents the category, b represents the position, is the predicted bounding box. If the number of ground truth bounding boxes is less than the number of predicted bounding boxes, the ground truth bounding boxes will be filled with the empty set

[0021] Based on this matching result The traditional DETR series detectors will perform subsequent loss function optimization:

[0022] The calculation of the loss function is based on the above matching result is also composed of the classification loss and the localization loss.

[0023] The loss function design principle of the zero-shot object detection algorithm proposed by the present invention is as follows:

[0024] In the zero-shot object detection algorithm, the optimal matching and loss function optimization are carried out for each category: that is, only one category (denoted as in the current sampled category set ) is used for optimal matching and loss function calculation, and then all the categories are added up to obtain the final loss function. The reason for calculating the loss for each category is as follows: In the traditional DETR, each query vector generates a prediction for any category. However, in the zero-shot object detection of the algorithm of the present invention, each query vector fuses a semantic vector of a specific category to become a category-specific query vector, and only generates a prediction box for that category. Therefore, the matching process changes from the multi-category matching of the traditional DETR to a two-category matching for each category, specifically:

[0025] S5, optimal matching for each category; for the prediction result First, perform optimal matching for each category. For each category its matching target is:

[0026]

[0027]

[0028] where c i is the category of the ground truth bounding box b i in the image I. Through the above formula, the matching target is modified to two categories, 0 and 1, according to whether the category is the same as ;

[0029] The result of its optimal matching is expressed as:

[0030]

[0031] where T τ is the number of query vectors, which is also equal to the number of prediction results; is one of the full permutations of T τ elements, is the classification loss function, is the localization loss function, is the matching target, is the category prediction result output by the decoder, indicating the probability that the corresponding bounding box prediction belongs to the incorporated semantic category.

[0032] S6, based on the result found by the optimal matching For the category its loss function is defined as:

[0033]

[0034] in, is the classification loss, is the regression loss, is the reconstruction loss for comparison; is the corrected label in equation (5), The category prediction result output by the decoder indicates the corresponding calibration box prediction the probability of belonging to the incorporated semantic category;

[0035] for Its implementation is as follows:

[0036]

[0037]

[0038] Among them, N pos For prediction The number of labeled boxes for the category, is the output feature of the penultimate layer of the decoder, h ρ is a linear layer that maps from visual space to semantic space, and κ is a temperature control hyperparameter; For Category The semantic vector of k is the category of the semantic vector integrated into the k-th query vector. In the above contrast reconstruction loss First, we will By h ρ The mapping is done to the semantic space, and then the similarity is calculated with the semantic vector. Through the contrast reconstruction loss, on the one hand, the visual space is supervised to be reconstructed back to the semantic space, so that the visual features can retain the semantic information as much as possible during the training process. On the other hand, the reconstructed semantic vector is compared with the semantic vectors of other categories to calculate the contrast loss, so as to narrow the distance with the semantic vectors of the same category and widen the distance with the semantic vectors of other categories, thus enhancing the generalization of the model.

[0039] S7, the total loss function is the category set Add up all categories in:

[0040]

[0041] The zero-sample target detection algorithm of the present invention has a wide range of application scenarios.

[0042] For example, it can be applied to the field of endangered animal protection. By relying solely on the descriptive semantic vectors of endangered animals, endangered animals can be detected, and the descriptive semantic vectors can be converted from the text descriptions of endangered animals through a pre-trained language model. If traditional object detection is used, a large number of training samples of endangered animals are required to train a better object detector, and the training samples of endangered animals are often difficult to obtain.

[0043] The advantages and beneficial effects of the present invention are as follows:

[0044] Most of the existing zero-shot object detection works are modified based on Faster RCNN, and its process is generally as follows: First, use the RPN network to generate a lot of candidate boxes, and then use the visual-semantic alignment module to classify and screen the candidate boxes. The visual-semantic alignment module here can generally be divided into two categories. One is to map the visual features and semantic features to the same space and then calculate the similarity; the other is to train a generation network to learn to generate visual features from semantic vectors, so as to convert zero-shot object detection into supervised object detection.

[0045] The method of the present invention is constructed based on the DETR detector of the transformer architecture. The zero-shot learning mechanism is introduced into the DETR deep object detection framework. The category semantic vectors are directly integrated into the query vectors of DETR, and the results are directly predicted through the decoder. During the training process, the training is completed through optimal matching and loss calculation for each category. The method framework of the present invention is simple, easy to use, highly scalable, and highly interpretable. The results of zero-shot object detection on the mainstream visual attribute datasets show that the performance of this method is significantly better than the existing methods. The present invention provides algorithm support for object detection technology in the industrial application field and can also be easily extended to other zero-shot learning tasks. Brief Description of the Drawings

[0046] Figure 1 It is a schematic diagram of the network structure of the present invention.

[0047] Figure 2 It is a comparison between the method of the present invention and other methods. Detailed Embodiments

[0048] The present invention will be further described below through embodiments in conjunction with the drawings.

[0049] For testing, resnet50 is used as the convolutional network feature extraction module, and the parameters pre-trained on ImageNet are used for weight initialization.

[0050] As Figure 1 shown, the zero-shot object detection method based on DETR and meta-learning has the following steps:

[0051] 1. Sample the image data and class semantic vectors. For an image, randomly sample a training image I from the training set. For the class semantic vectors, first randomly sample the classes included in image I and randomly sample the classes not included in image I and together to form the class set Meanwhile, image I undergoes feature extraction through a backbone network and a Transformer encoder to obtain x I . Here, we select the MS COCO dataset for training and testing. We use 48 out of the 80 classes in MS COCO as the visible classes, and the other 17 classes as the unseen classes. All of the following come from the 48 visible classes. During model testing, simply replace with all the unseen classes. Here, the backbone network is the ResNet50 network, and the Transformer encoder is a 6 - layer Transformer encoder.

[0052] 2. For the class set its corresponding semantic vectors are denoted as Pass the semantic vectors through a linear layer for mapping to obtain the mapped semantic vectors i.e., Here, the semantic vectors are obtained through the CLIP pre - trained model, and is a single - layer fully - connected network. has a dimension of 512, and after being mapped by the dimension becomes 256. The CLIP model comes from: https: / / arxiv.org / abs / 2103.00020.

[0053] 3. Add the mapped semantic vectors and the object query vectors of DETR to obtain the query vectors after fusing the semantic vectors i.e., Here, the query vectors are the query vectors in the DETR detector, a total of 900. Each query vector is a vector of learnable parameters with a dimension of 256 and is automatically updated during the training process.

[0054] 4. Input the query vectors after fusing the semantic vectors into the decoder g of DETR θ to obtain the predicted bounding box results i.e., Here, the decoder g θIt consists of a 6-layer Transformer decoder.

[0055] 5. For the prediction result First, perform class-by-class optimal matching. For each class Its matching target is Among them, Among them, c i is the class of the ground truth bounding box b in the image I i The result of its optimal matching is expressed as:

[0056]

[0057] Through the above formula, the matching target is modified to two classes, 0 and 1, according to whether the class is the same as . Among them, T τ is the number of query vectors and is also equal to the number of prediction results. is one of the full permutations of T τ elements. is the classification loss function, is the localization loss function. is the matching target, is the class prediction result output by the decoder, indicating the probability that the corresponding bounding box prediction belongs to the incorporated semantic class. Here, is implemented with focal loss, is implemented with L1 loss.

[0058] 6. Further, based on the result searched by the optimal matching For the class Its loss function is defined as:

[0059]

[0060] Among them, is the classification loss, specifically the focal loss function, is the regression loss, specifically the L1 loss function, is the contrast reconstruction loss; among them, is the corrected label in formula (5), is the class prediction result output by the decoder, indicating the probability that the corresponding bounding box prediction belongs to the incorporated semantic class.

[0061] For Its implementation form is as follows:

[0062]

[0063]

[0064] Among them, N pos is the number of calibration boxes predicted as categories, is the output feature of the penultimate layer of the decoder, h ρ is the linear layer that maps from the visual space to the semantic space, and κ is the temperature control hyperparameter; is the semantic vector of the category c. k is the category of the semantic vector incorporated by the k-th query vector. In the above contrast reconstruction loss , we first map the in the visual space through h ρ to the semantic space, and then calculate the similarity with the semantic vector. Here, h ρ is a one-layer fully connected network that maps from a dimension of 256 to a dimension of 512; through the contrast reconstruction loss, on the one hand, the visual space is supervised to be reconstructed back to the semantic space so that the visual features can retain semantic information as much as possible during the training process. On the other hand, the reconstructed semantic vector needs to calculate the contrast loss with the semantic vectors of other categories to narrow the distance from the semantic vectors of the same category and widen the distance from the semantic vectors of other categories, enhancing the generalization of the model.

[0065] 7. Further, the total loss function is the sum of all categories in the category set ,

[0066]

[0067] 8. The method of the present invention is tested and verified on the MS COCO dataset. The specific process is as follows:

[0068] MS COCO dataset: It is an object detection benchmark dataset, which contains a total of 80 categories of calibrated objects, among which 48 categories are visible categories and 17 categories are unseen categories.

[0069] And the present invention is compared with the following existing few-shot object detection methods: DSES, TD, PL, BLC, ZSDTR, Robust-Syn, ContrastZSD. The experimental results on the MS COOC dataset are shown in Table 1 below:

[0070] Table 1: Performance comparison on the MS COCO dataset

[0071]

[0072] As shown in Table 1, the backbone networks adopted by different methods are listed. It can be seen that the present invention exceeds the SOTAs, and the present invention achieves the best results in both recall rate and mAP. These data prove the effectiveness of the present invention.

[0073] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A zero-shot object detection method based on DETR and meta-learning. Definition of the zero-shot object detection problem: In zero-shot object detection, the dataset is divided into visible classes and unseen classes. The visible classes are used for training, and the unseen classes are used for test verification. At the same time, a semantic vector is provided for each class of the visible classes and unseen classes, and this semantic vector is a description of the class. The semantic vector is fused into the query vector, and the result is directly decoded. The specific steps are as follows: S1, Sample image data and class semantic vectors; For an image, in the training set, randomly sample a training image I; For the class semantic vector, first randomly sample the classes contained in image I and randomly sample the classes not contained in image I and together constitute the class set Meanwhile, image I undergoes feature extraction through the backbone network and the transformer encoder to obtain the feature x of image I I ; S2. Perform a linear mapping on the semantic vectors corresponding to the category set; for the category set The corresponding semantic vector is denoted as The semantic vector Pass through the linear layer Perform the mapping to obtain the mapped semantic vector That is: S3. Add the mapped semantic vector to the object query vector of DETR; Add the mapped semantic vector and the object query vector of DETR to obtain a query vector after fusing the semantic vector That is: S4, decode the query vector; input the query vector after fusing the semantic vector into the decoder g of DETR θ to obtain the predicted bounding box results That is: S5, perform optimal matching for each row category; for the prediction result First, perform optimal matching for each category The matching target is: Among them, c i is the category of the true calibration box b in the image I i , according to the above formula, the matched targets are modified to two categories, 0 and 1, depending on whether the category is the same as ; The result of its optimal match is expressed as: Among them, T τ is the number of query vectors, which is also equal to the number of prediction results; is one of the full permutations of T τ elements; is the classification loss function, is the localization loss function; is the target of matching, is the class prediction result output by the decoder, indicating the probability that the corresponding calibration box prediction belongs to the incorporated semantic category.

2. The zero-shot object detection method based on DETR and meta-learning according to claim 1, characterized in that Results found based on optimal matching For the category Its loss function is defined as: Among them, is the classification loss, is the regression loss, is the contrastive reconstruction loss; is the corrected label in formula (5), is the class prediction result output by the decoder, indicating the predicted bounding box belongs to the probability of the incorporated semantic class; For The implementation form is as follows: Among them, N pos is the number of calibration boxes predicted as categories, is the output feature of the penultimate layer of the decoder, h ρ is the linear layer that maps from the visual space back to the semantic space, and κ is the temperature control hyperparameter; is the semantic vector of the category , c k is the category of the semantic vector incorporated into the k-th query vector; in the above contrast reconstruction loss , first, the in the visual space is mapped to the semantic space through h ρ , and then the similarity is calculated with the semantic vector; through the contrast reconstruction loss, on the one hand, the visual space is supervised to be reconstructed back to the semantic space, and during the training process, the visual features are made to retain semantic information as much as possible. On the other hand, the reconstructed semantic vector and the semantic vectors of other categories are used to calculate the contrast loss to narrow the distance from the semantic vectors of the same category and widen the distance from the semantic vectors of other categories, enhancing the generalization of the model.

3. The zero-shot object detection method based on DETR and meta-learning according to claim 2, wherein The total loss function is the sum of all classes in the class set as follows:

Citation Information

Patent Citations

  • Zero sample learning method and system based on semantic attribute attention redistribution mechanism

    CN110163258A

  • Multi-mark zero sample learning method based on deep mutual learning

    CN114998613A