An open vocabulary object detection algorithm based on cross-validation identification mechanism

CN117671246BActive Publication Date: 2026-09-29DALIAN UNIV OF TECH +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311700186.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2026-09-29
Estimated Expiration
2043-12-12

AI Technical Summary

Benefits of technology

[0016](1)能够在传统的开放词表目标检测任务设置下实现最为先进的性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117671246B_ABST
    Figure CN117671246B_ABST
Patent Text Reader

Abstract

The application belongs to the field of machine learning, target detection and open vocabulary target segmentation, and discloses an open vocabulary target detection algorithm based on a cross-validation identification mechanism. In order to improve the identification accuracy of the model under a large-scale vocabulary expression scene, a cross-validation identification module is designed, which mainly comprises two cross-validation identification branches and a multi-branch voting component. The two cross-validation identification branches preliminarily predict through the combination of the image-level and region-level classifiers and the label predictor, and then the final prediction mixed confidence is obtained according to the certainty and complementarity of the two branches for different category predictions through the multi-branch voting component, so that the detection precision is improved, the most advanced performance is achieved under the traditional open vocabulary target detection task setting, and the optimal result is achieved under the relatively complex large-scale open vocabulary target detection task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of machine learning, object detection, and open vocabulary object segmentation, and involves algorithms such as the pre-trained transferable model CLIP, visual Transformer, image-label model RAM, and object detection model DETR. Specifically, it is an open vocabulary object detection algorithm based on a cross-validation recognition mechanism. Background Technology

[0002] Object detection, a fundamental task in computer vision, aims to identify and locate different categories of objects in images and is a core technology for understanding complex scenes. Early deep learning object detection models included R-CNN, Fast R-CNN, and Faster R-CNN, which used Region Proposal Networks (RPNs) to generate candidate regions and combined them with Convolutional Neural Networks (CNNs) for object classification and bounding box regression. Subsequently, with the widespread application of attention mechanisms in deep learning, attention-based object detection models have also gained attention. Among them, DETR is an end-to-end object detection model based on the Transformer architecture. It achieves object detection through a global self-attention mechanism, avoiding the complex designs of traditional anchor boxes and non-maximum suppression (NMS).

[0003] However, the aforementioned methods for solving the object detection problem all face the bottleneck of only being able to detect a small number of predefined categories. To address this issue, Zareian et al. proposed the concept of open-vocabulary object detection and the first open-vocabulary object detection model, Open Vocabulary R-CNN (OVR-CNN), in their paper "Open-Vocabulary Object Detection Using Captions." Subsequently, with Harold et al. introducing a large-scale contrastive learning pre-trained model into the open-vocabulary object detection task in "Grounded Language-Image Pre-training," open-vocabulary object detection algorithms have gradually become a research hotspot. Open-vocabulary object detection refers to a method in object detection tasks that does not require predefined target categories, is not limited by the number of categories, and can discover, identify, and locate targets of any category during the training and inference phases. Compared with traditional object detection methods, open-vocabulary object detection algorithms enable models to automatically adapt to new target categories, thus becoming more flexible and applicable to constantly changing scenarios in the real world. OV-DETR (Open-vocabulary DETR) utilizes a large pre-trained model CLIP to generate conditional queries containing textual information, and for the first time incorporates open vocabulary information into the DETR object detection model. CORA reduces the domain difference between image and region features through region hints and pre-matching strategies, thus improving detection performance. Although these works have promoted the progress and development of open-vocabulary object detection, they are still plagued by problems such as overfitting. When inputting large-scale open-vocabulary category text, i.e., when facing large-scale open-vocabulary tasks that are closer to real-world scenarios, they struggle due to a large amount of confused category interference, resulting in a significant performance drop.

[0004] Recently, thanks to a more robust multimodal fusion approach, RAM (Random Access Detection) has achieved significant breakthroughs in large-scale image-level labeling tasks. RAM-like models can assign multiple corresponding semantic labels to an image, demonstrating the potential to solve large-scale open-vocabulary tasks—a potential that has not yet been fully explored in open-vocabulary object detection tasks. Preliminary experiments show that labeling models, due to their overemphasis on global semantic information, suffer a decline in their ability to recognize region-level semantic features, failing to guarantee accurate target localization. Simultaneously using labeling models and open-vocabulary object detection models can address the problem to some extent, but this approach is computationally expensive and suffers from error accumulation due to the use of multiple models. Therefore, how to introduce labeling capabilities into an open-vocabulary object detection model and enable its object detection and labeling capabilities to mutually reinforce each other, thereby reducing computational costs and improving accuracy, is a key issue for open-vocabulary tasks, especially large-scale ones. Summary of the Invention

[0005] To address the aforementioned issues, this invention designs a cross-validation recognition module. By combining the superior generalization ability of the CLIP model with the strong image-level discrimination capability of the label model, labeling capabilities are introduced into the open vocabulary object detection model, thereby improving the model's recognition accuracy in large-scale vocabulary scenarios. Based on this module, this invention proposes an open vocabulary object detection algorithm, LOV-DETR (Large-and-Open Vocabulary DETR), for large-scale vocabulary input.

[0006] The technical solution of this invention:

[0007] An open vocabulary object detection algorithm based on cross-validation recognition mechanism, called the LOV-DETR model:

[0008] The LOV-DETR model is a DETR-type framework that includes an image feature extractor, a text feature extractor, and a detector. The image feature extractor consists of a CLIP image encoder and a Transformer encoder, used to extract image features aligned with text features. The text feature extractor is composed of a CLIP text encoder. The detector consists of a cross-validation recognition module used twice and an anchor box refinement module. The anchor box refinement module introduces a class pre-matching mechanism on the basis of the DAB-DETR structure with dynamic anchor boxes. By cooperating with the cross-validation recognition module, it avoids repeated inference for each class and accelerates the convergence speed of model training.

[0009] The framework of the LOV-DETR model is expressed as follows: Where x represents the input image, V L V indicates a larger vocabulary. B This represents a smaller known vocabulary, and y represents the detection result. It is in V B The model trained on it; in addition, V L and V B There are no matching categories, and V L >>V B ;

[0010] Given an image as input, image and text features are extracted using an image feature extractor and a text feature extractor, respectively. These features are then passed to a detector, where they are shared by the detector's cross-validation module and anchor box refinement module. The anchor box refinement module takes anchor boxes and their corresponding pre-matched category text as input and outputs more precise bounding boxes and their corresponding confidence scores. The cross-validation module is used twice: first, before anchor box refinement, it performs pre-matching through region classification, assigning a corresponding category text to each anchor box; second, after anchor box refinement, it performs re-identification through region classification, verifying whether the matched category exists in the image. The LOV-DETR model outputs the bounding boxes and their corresponding confidence scores.

[0011] The cross-validation identification module consists of two cross-validation identification branches and a multi-branch voting component, and is composed of expressions. To represent; among which, This represents the confidence level of the match between region r and category c. and This represents the prediction results of the two branches, and the number α represents the weight coefficient of the two branch predictions;

[0012] The two cross-validation recognition branches are constructed by combining classifiers and label predictors at different levels (image-level and region-level). Specifically, one branch consists of a CLIP-based image-level classifier and a region-level label predictor; the other branch consists of an image-level label predictor and a CLIP-based region-level classifier. The specific algorithm process is as follows: The cross-validation recognition branch uses an image-level classifier or label predictor to filter through a large-scale vocabulary to obtain a selected subset V. S The region classification confidence Conf(c,r) for all categories in the vocabulary is obtained from a region-level classifier or label predictor; when category c∈V S At that time, the final prediction confidence result P (c,r) For γ + ×Conf(c,r), when the category At that time, P (c,r) For γ _ ×Conf(c,r). γ + and γ _These are hyperparameters, named reward factor and penalty factor, respectively, reflecting the degree of influence of image-level screening on the final prediction result. The classifier is based on CLIP, outputting the predicted category based on the cosine similarity. The label predictor predicts the presence or absence of a category based on the confidence level output by the label prediction head. The label prediction head is a two-layer Bert Transformer with the self-attention mechanism removed, retaining only the cross-attention mechanism to save computational costs. The classifier and label predictor use different types of models to make predictions in different ways. Extensive experimental verification shows that they have good complementarity. Combining them for cross-validation and screening can improve the accuracy of the model in open object detection.

[0013] The aforementioned cross-validation identification branches have a strong filtering effect, but simply taking the intersection or union of the prediction results of two cross-validation identification branches will lead to an increase in the number of false samples. Therefore, this invention proposes a multi-branch voting component, which reduces the number of false samples by obtaining a mixed confidence level based on the certainty of the predictions of two cross-validation identification branches for different categories. Specifically, if both cross-validation branches predict the existence of a certain category, their confidence scores are multiplied to obtain the final predicted mixed confidence score. If both cross-validation branches predict the non-existence of a certain category, it is directly filtered out. If one cross-validation branch filters out a category while the other does not, the voting weights of the two branches are dynamically adjusted based on the original size of the region corresponding to that category. Then, the voting weight of each branch is multiplied by the predicted confidence score to obtain the confidence score of each branch with voting weight. The confidence scores with voting weight are multiplied to obtain the final mixed confidence score of that category. If the mixed confidence score of that category is the highest among all input categories, the classification result of that category as a region is retained; otherwise, it is filtered out. After the above process, the voting strategy selects the highest mixed confidence score as the final region classification result. The multi-branch voting component allows the two branches to complement each other, reducing the number of false samples and improving the final prediction accuracy.

[0014] Compared to baseline methods, the open vocabulary object detection algorithm based on cross-validation can improve the detection accuracy by 1.1 percentage points in traditional open vocabulary object detection task settings, and by 8.2 percentage points in more complex large-scale open vocabulary object detection task settings, providing a new solution for open vocabulary object detection tasks, especially large-scale open vocabulary object detection tasks.

[0015] The beneficial effects of this invention are:

[0016] (1) It can achieve the most advanced performance under the traditional open vocabulary target detection task settings.

[0017] (2) By using a specially designed cross-validation recognition module and different levels of feature interaction, an efficient and unified network architecture was obtained, which not only reduced the computational cost, but also improved the accuracy of distinguishing and detecting confused categories. It achieved the best results in large-scale open vocabulary tasks, effectively solved the unique challenges of large-scale open vocabulary tasks, and had good zero-sample transfer capability. Attached Figure Description

[0018] Figure 1 This is a flowchart of the LOV-DETR algorithm, an open vocabulary object detection network based on cross-validation recognition mechanism.

[0019] Figure 2 This is the algorithm flowchart for the cross-validation recognition module. Detailed Implementation

[0020] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and technical solutions.

[0021] Figure 1 This is a flowchart of the LOV-DETR open vocabulary object detection network based on cross-validation. The LOV-DETR model is a DETR-type framework, containing an image feature extractor, a text feature extractor, and a detector. The image feature extractor consists of a CLIP image encoder and a Transformer encoder, used to extract image features aligned with text features; the text feature extractor is composed of a CLIP text encoder; and the detector consists of a cross-validation module used twice and an anchor box refinement module.

[0022] Figure 2 This is the algorithm flowchart for the cross-validation recognition module. It consists of two cross-validation recognition branches and a multi-branch voting component. The two cross-validation recognition branches are constructed by combining classifiers and label predictors at different levels (image-level and region-level). Each branch is selected from a large-scale vocabulary by either an image-level classifier or a label predictor, resulting in a selected subset V. S The region classification confidence scores conf(c,r) for all categories in the vocabulary are obtained from a region-level classifier or label predictor. When category c∈V S At that time, the prediction confidence result P (c,r) For γ + ×Conf(c,r), when the category At that time, P (c,r) For γ _ ×Conf(c,r). γ + and γ _These are hyperparameters, named reward factor and penalty factor, respectively, reflecting the degree of influence of image-level screening on the final prediction result. The multi-branch voting component obtains a mixed confidence score from the confidence scores of the two branches and selects the highest mixed confidence score as the final region classification result.

[0023] In the cross-validation recognition module of LOV-DETR, a 2-layer Transformer is used as the core of the label predictor. This invention removes the self-attention of the 2-layer Transformer and retains only the cross-attention. The bounding box refinement module of LOV-DETR adopts the structure of DAB-DETR and is trained according to the traditional open vocabulary object detection task settings. During testing, the confidence corresponding to unknown categories is not amplified. The specific training and parameter settings are as follows: A phased training method is adopted. First, the label predictor of LOV-DETR is trained for 5 rounds on the pseudo-label dataset generated in RAM with a base learning rate of 1e-4. The learning rate decays by a factor of 0.1, and an asymmetric loss function is used. Then, the label predictor, CLIP image encoder, and CLIP text encoder are fixed, and the other trainable modules of the model are trained for 35 rounds, with the learning rate maintained at 1e-4 and no decay. The optimizer is AdamW with a batch size of 8 and a weight decay of 1e-4 is added. During inference, the reward coefficient Y in the cross-validation recognition module is... + Penalty coefficient Υ _ The weight coefficients α for the two branches are set to 1.05, 0.9 and 0.35 respectively. The top k labels with the highest confidence are taken as positive classes and the rest as negative classes. k is set to one-quarter of the total number of classes in the dataset, for example, 20 in the COCO dataset. Non-maximum suppression (NMS) with an IoU threshold of 0.5 is applied.

Claims

1. An open vocabulary target detection algorithm based on cross-validation recognition mechanism, referred to as the LOV-DETR model; characterized in that, The LOV-DETR model is a DETR-type framework, comprising an image feature extractor, a text feature extractor, and a detector. The image feature extractor consists of a CLIP image encoder and a Transformer encoder, used to extract image features aligned with text features. The text feature extractor is composed of a CLIP text encoder. The detector consists of a cross-validation recognition module used twice and an anchor box refinement module. The framework of the LOV-DETR model is expressed as follows: ,in, Represents the input image. Indicates a larger vocabulary. This represents the test results. Is The model trained on it; in addition, and There are no matching categories, and ; Given an image as input, image and text features are extracted using an image feature extractor and a text feature extractor, respectively. These features are then passed to a detector, shared by the detector's cross-validation module and anchor box refinement module. The anchor box refinement module takes anchor boxes and their corresponding pre-matched category text as input and outputs more precise bounding boxes and their corresponding confidence scores. The cross-validation module is used twice: first, before anchor box refinement, it performs pre-matching through region classification, assigning a corresponding category text to each anchor box; second, after anchor box refinement, it performs re-identification through region classification, verifying the existence of the matched category in the image. The LOV-DETR model outputs the bounding boxes and their corresponding confidence scores. The cross-validation identification module consists of two cross-validation identification branches and a multi-branch voting component, and is composed of expressions. To represent; among which, Indicates for the region and categories Match confidence, and This represents the prediction results of the two branches. Represents the weighting coefficients of the two-branch predictions; The two cross-validation recognition branches are constructed by combining classifiers and label predictors at different levels: one branch consists of a CLIP-based image-level classifier and a region-level label predictor; the other branch consists of an image-level label predictor and a CLIP-based region-level classifier. The specific algorithm process is as follows: the cross-validation recognition branch uses either the image-level classifier or the label predictor to filter the vocabulary, obtaining a selected subset. The regional classification confidence scores for all categories in the vocabulary are obtained from a regional classifier or label predictor. When category At that time, the final prediction confidence result When category hour, ; and These are hyperparameters, named reward factor and penalty factor respectively, reflecting the degree of influence of image-level screening on the final prediction result; the classifier is based on CLIP and outputs the predicted category according to the cosine similarity; the label predictor predicts the existence of the category by the confidence of the label prediction head outputting the category; the label prediction head is a two-layer BertTransformer with the self-attention mechanism removed and only the cross-attention mechanism retained. The multi-branch voting component reduces the number of false samples by obtaining a mixed confidence score based on the certainty of the predictions of two cross-validation branches for different categories. Specifically: if both cross-validation branches predict the existence of a certain category, their confidence scores are multiplied to obtain the final predicted mixed confidence score; if both cross-validation branches predict the non-existence of a certain category, it is directly filtered out; if one cross-validation branch filters out a certain category while the other does not, the voting weights of the two branches are dynamically adjusted based on the original size of the region corresponding to that category, and then the voting weight of each branch is multiplied by the predicted confidence score to obtain the confidence score of each branch with voting weight. The confidence scores with voting weight are multiplied to obtain the final mixed confidence score of that category. If the mixed confidence score of that category is the highest among all the mixed confidence scores of the input categories, the category is retained as a region classification result; otherwise, it is filtered out. After the above process, the voting strategy selects the highest mixed confidence score as the final region classification result.

Citation Information

Patent Citations

  • Open set target detection and identification method based on deep neural network

    CN114241260A

  • Assertion Detection in Multi-Labelled Clinical Text using Scope Localization

    US20210174027A1