A unified open vocabulary detection method based on language-aware selective query fusion

By constructing a unified open vocabulary detection model based on language perception based on selective query fusion, the problems of insufficient generalization ability of new categories and insufficient fusion of language visual modalities in the existing technology are solved, and higher detection accuracy and correlation are achieved.

CN118673910BActive Publication Date: 2025-05-16SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410809649.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2025-05-16
Estimated Expiration
2044-06-21

AI Technical Summary

Technical Problem

The prior art lacks generalization ability when facing new categories, the fusion of language modalities and visual modalities is not effective enough, and it is difficult to process image data containing rich language descriptions, resulting in low detection accuracy and correlation.

Method used

A unified open vocabulary detection method based on language perception is proposed. By constructing a detection model including image encoder, text encoder, transformer encoder and transformer decoder, the language-aware query selection module and query fusion module are used to dynamically select and fuse images and text features to generate the final detection result.

Benefits of technology

It improves the model's ability to generalize new categories, enhances the fusion effect of language and visual features, can better process image data containing rich language descriptions, and improves the accuracy and relevance of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118673910B_ABST
    Figure CN118673910B_ABST
Patent Text Reader

Abstract

The invention discloses a unified open vocabulary detection method based on language-aware selective query fusion. First, a unified open vocabulary detection model based on language-aware selective query fusion is constructed, including an image encoder, a text encoder, a transformer encoder and a transformer decoder. The transformer decoder includes a language-aware query selection module, a language-aware query fusion module, a category mapping layer and a target box regression layer. The image encoder extracts image features, the text encoder extracts language features, and the fusion module effectively fuses the image features and the language features to generate a final detection result. The invention can better understand and identify objects in images and improve the accuracy of model detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and artificial intelligence, and more specifically, to a unified open vocabulary detection method based on language-aware selective query fusion. Background Art

[0002] Traditional object detection methods based on fixed vocabulary can usually only recognize pre-defined categories, and cannot effectively recognize unknown or unseen categories. This method is limited in complex application scenarios.

[0003] The problems existing in the existing technology mainly include:

[0004] 1. The lack of an end-to-end training process results in the model’s insufficient generalization ability when faced with new categories.

[0005] 2. The fusion of language modality and visual modality is not effective enough, resulting in the model being unable to fully utilize natural language descriptions to improve detection accuracy and relevance.

[0006] 3. The processing capabilities of large-scale data sets are limited, especially for image data containing rich language descriptions. Existing technologies find it difficult to achieve efficient learning and reasoning. Summary of the invention

[0007] In order to overcome the defects of the above-mentioned prior art models in that they have insufficient generalization ability and low detection accuracy for new categories, the present invention provides a unified open vocabulary detection method based on language-aware selective query fusion.

[0008] In order to solve the above technical problems, the technical solution of the present invention is as follows:

[0009] The present invention proposes a unified open vocabulary detection method based on language-aware selective query fusion, comprising:

[0010] S1: obtaining different types of data sources and preprocessing them to obtain unified type data, wherein the unified type data includes input images and unified language text, and constructing a unified open vocabulary detection model for language-aware selective query fusion, wherein the model includes an image encoder, a text encoder, a transformer encoder and a transformer decoder, wherein the transformer decoder includes a language-aware query selection module, a language-aware query fusion module, a category mapping layer and a target box regression layer;

[0011] S2: Input the input image into the image encoder to obtain an image feature vector, and input the unified language text into the text encoder to obtain a text feature vector;

[0012] S3: Obtain a refined image feature vector according to the image feature vector;

[0013] S4: inputting the refined image feature vector and the text feature vector into a language-aware query selection module to obtain a dynamically selected content query feature vector and a statically learnable content query feature vector;

[0014] S5: inputting the refined image feature vector, the dynamically selected content query feature vector, and the static learnable content query feature vector into a language-aware query fusion module to obtain a fused content query feature vector;

[0015] S6: Input the fused content query feature vector into the category mapping layer and the target box regression layer respectively to obtain the corresponding predicted category and target box predicted coordinates;

[0016] S7: Construct a total loss function based on the predicted category and the predicted coordinates of the target box, train the model, and obtain a trained language-aware, selective query-fused unified open vocabulary detection model;

[0017] S8: Obtain the data to be detected, input it into the trained language-aware selective query fusion unified open vocabulary detection model, and output the classification results and target box coordinates.

[0018] Preferably, in S1, obtaining different types of data sources and preprocessing to obtain uniform type of data includes:

[0019] The data sources include detection type data, positioning type data and image-text type data. The detection type data includes input image, image target frame category label and image target frame coordinates. The positioning type data includes input image, image target frame description text and image target frame coordinates. The image-text type data includes input image, image text description and image target frame coordinates. The target frame coordinates of the image-text data are the edge coordinates of the image.

[0020] For detection type data, the image target box category label is processed into a sentence as a unified language text, the image target box description text of the positioning type data is directly used as the unified language text, and the image text description of the image-text type data is directly used as the unified language text;

[0021] Using triple data format of uniform type data Represents different types of data sources, where x∈R 3×H×W Represents the input image, H and W are the height and width of the image respectively, is the image target box coordinate, represents the unified language text, and n is the number of target boxes.

[0022] Preferably, in S3, obtaining a refined image feature vector according to the image feature vector includes:

[0023] The position feature vector corresponding to the image feature vector is obtained, and the image feature vector and the position feature vector are input into the transformer encoder to obtain a refined image feature vector.

[0024] Preferably, the transformer encoder comprises N transformer encoder layers connected in sequence.

[0025] Preferably, in S4, inputting the refined image feature vector and the text feature vector into a language-aware query selection module, and obtaining a dynamically selected content query feature vector and a statically learnable content query feature vector comprises:

[0026] The initial dynamically selected content query feature vector and the statically learnable content query feature vector are obtained, the similarity between the refined image feature vector and the text feature vector is calculated, and the similarity is sorted in descending order. The first K most relevant refined image feature vectors are taken to update the initial dynamically selected content query feature vector to obtain the dynamically selected content query feature vector, and the first K most relevant feature vector indexes are taken to update the initial statically learnable content query feature vector to obtain the statically learnable content query feature vector.

[0027] Preferably, the language-aware query fusion module comprises M sequentially connected post-fusion query fusion submodules;

[0028] Each of the post-fusion query fusion submodules includes a first self-attention layer, a first cross-attention layer, a second cross-attention layer, a first tanh gating layer, a first splicing layer, a first feedforward neural network layer, a second tanh gating layer, a second splicing layer, and a second feedforward neural network layer;

[0029] The output end of the first self-attention layer is connected to the third input end of the first cross-attention layer, and the output ends of the first cross-attention layer are connected to the third input end of the second cross-attention layer and the input end of the first concatenation layer;

[0030] The second cross attention layer, the first tanh gating layer, the first splicing layer, the first feedforward neural network layer, the second tanh gating layer, the second splicing layer, and the second feedforward neural network layer are connected in sequence;

[0031] The output end of the first splicing layer is also connected to the input end of the second splicing layer;

[0032] The input end of the first self-attention layer of the first post-fusion query fusion submodule is connected to the second output end of the language-aware query selection module; the input end of the first self-attention layer of the jth post-fusion query fusion submodule is connected to the output end of the second feedforward neural network layer of the j-1th post-fusion query fusion submodule, j=2,…,M;

[0033] The first input terminal and the second input terminal of the first cross-attention layer of the i-th post-fusion query fusion submodule are both connected to the output terminal of the transform encoder; the first input terminal and the second input terminal of the second cross-attention layer of the i-th post-fusion query fusion submodule are both connected to the first output terminal of the language-aware query selection module; i=1,2,…,M.

[0034] Preferably, the language-aware query fusion module comprises M sequentially connected middle-fusion query fusion submodules;

[0035] Each of the fusion query fusion submodules includes a first self-attention layer, a first cross-attention layer, a first tanh gating layer, a first splicing layer, a first feedforward neural network layer, a second tanh gating layer, a second splicing layer, a second cross-attention layer, and a second feedforward neural network layer;

[0036] The output end of the first self-attention layer is connected to the third input end of the first cross-attention layer and the input end of the first concatenation layer respectively;

[0037] The first cross attention layer, the first tanh gating layer, the first splicing layer, the first feedforward neural network layer, the second tanh gating layer, and the second splicing layer are connected in sequence; the output end of the first splicing layer is connected to the input end of the second splicing layer, the output end of the second splicing layer is connected to the third input end of the second cross attention layer, and the output end of the second cross attention layer is connected to the input end of the second feedforward neural network layer;

[0038] The input end of the first self-attention layer of the first fusion query fusion submodule is connected to the second output end of the language-aware query selection module; the input end of the first self-attention layer of the j-th fusion query fusion submodule is connected to the output end of the second feedforward neural network layer of the j-1-th fusion query fusion submodule, j=2,…,M;

[0039] The first input terminal and the second input terminal of the second cross-attention layer of the i-th fused query fusion submodule are both connected to the output terminal of the transform encoder; the first input terminal and the second input terminal of the first cross-attention layer of the i-th fused query fusion submodule are both connected to the first output terminal of the language-aware query selection module; i=1,2,…,M.

[0040] Preferably, the language-aware query fusion module comprises M sequentially connected pre-fusion query fusion submodules;

[0041] Each of the front-fusion query fusion submodules includes a first cross attention layer, a first tanh gating layer, a first splicing layer, a first feedforward neural network layer, a second tanh gating layer, a second splicing layer, a first self-attention layer, a second cross attention layer, and a second feedforward neural network layer;

[0042] The first cross attention layer, the first tanh gating layer, the first splicing layer, the first feedforward neural network layer, the second tanh gating layer, and the second splicing layer are connected in sequence; the input end of the first cross attention layer is connected to the input end of the first splicing layer, the output end of the first splicing layer is connected to the input end of the second splicing layer, the output end of the second splicing layer is connected to the input end of the first self-attention layer, the output end of the first self-attention layer is connected to the third input end of the second cross attention layer, and the output end of the second cross attention layer is connected to the input end of the second feedforward neural network layer;

[0043] The first and second input ends of the first cross attention layer of the first front-fusion query fusion submodule are both connected to the first output end of the language-aware query selection module; the third input end of the first cross attention layer of the j-th front-fusion query fusion submodule is connected to the output end of the second feedforward neural network layer of the j-1-th front-fusion query fusion submodule, j=2,…,M;

[0044] The first input terminal and the second input terminal of the second cross-attention layer of the i-th front-fusion query fusion submodule are both connected to the output terminal of the transform encoder; the third input terminal of the first cross-attention layer of the i-th front-fusion query fusion submodule is connected to the second output terminal of the language-aware query selection module; i=1,2,…,M.

[0045] Preferably, in S7, the method for determining the total loss function includes:

[0046] L=L cls +α*L box +β*L giou +γ*L dn

[0047] Among them, L cls represents the classification loss, L box represents the box regression loss, L giou represents the giou loss, L dn represents the denoising loss, α, β and γ represent L box , L giou and L dn The weight factor of

[0048] The method for determining the classification loss, box regression loss, GIOU loss and denoising loss includes:

[0049] L cls =SigmoidFocal(S,GT cls )

[0050] L box =||B-GTbox ||

[0051] L giou =GIoU(S,GT cls )

[0052] L dn =Denoisy(S noisy ,GT cls )+Denoisy(B noisy ,GT box )

[0053] Among them, SigmoidFocal represents the SigmoidFocal loss function, S∈R QxC represents the predicted category alignment score, Q is the number of filtered text feature vectors, C represents the number of unified language text categories, O represents the predicted category, E t represents the text feature vector, Indicates E t The transpose of GT cls ∈R Q×C represents the true category, B represents the target box prediction coordinates, GT box ∈R Q×4 represents the real coordinates of the target box, S noisy and B noisy It is the predicted category and target box predicted coordinates after adding noise; GIoU represents the giou processing function, and Denoisy represents the denoising function.

[0054] The present invention also provides a unified open vocabulary detection system based on language-aware selective query fusion, which is used to implement the above-mentioned unified open vocabulary detection method, including:

[0055] A data acquisition module is used to acquire different types of data sources and perform preprocessing to obtain unified type data, wherein the unified type data includes input images and unified language text, and construct a unified open vocabulary detection model for language-aware selective query fusion, wherein the model includes an image encoder, a text encoder, a transformer encoder and a transformer decoder, wherein the transformer decoder includes a language-aware query selection module, a language-aware query fusion module, a category mapping layer and a target box regression layer;

[0056] A text and image encoding module, used for inputting an input image into an image encoder to obtain an image feature vector, and inputting a unified language text into a text encoder to obtain a text feature vector;

[0057] A refined image feature vector acquisition module, used to obtain a refined image feature vector according to the image feature vector;

[0058] A query selection processing module, used for inputting the refined image feature vector and the text feature vector into the language-aware query selection module to obtain a dynamically selected content query feature vector and a statically learnable content query feature vector;

[0059] A query fusion processing module, used for inputting the refined image feature vector, the dynamically selected content query feature vector and the static learnable content query feature vector into the language-aware query fusion module to obtain a fused content query feature vector;

[0060] The model prediction module is used to input the fused content query feature vector into the category mapping layer and the target box regression layer respectively to obtain the corresponding predicted category and target box predicted coordinates;

[0061] A model training module is used to construct a total loss function based on the predicted category and the target box predicted coordinates, train the model, and obtain a trained language-aware, selective query-fused unified open vocabulary detection model;

[0062] The classification result and target box coordinate acquisition module is used to obtain the data to be detected, input it into the trained language-aware selective query fusion unified open vocabulary detection model, and output the classification result and target box coordinates.

[0063] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0064] The present invention proposes a unified open vocabulary detection method based on language-aware selective query fusion. First, a unified open vocabulary detection model based on language-aware selective query fusion is constructed, including an image encoder, a text encoder, a transformer encoder and a transformer decoder. The transformer decoder includes a language-aware query selection module, a language-aware query fusion module, a category mapping layer and a target box regression layer. The image encoder is responsible for extracting image features, the text encoder extracts language features, and the fusion module is responsible for effectively fusing image features and language features to generate the final detection result. The present invention can better understand and identify objects in images and improve the accuracy of model detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 This is a flow chart of the unified open vocabulary detection method described in Example 1;

[0066] Figure 2 This is a system block diagram of the unified open vocabulary detection method described in Example 1;

[0067] Figure 3 This is a schematic diagram of different types of data sources described in Example 1;

[0068] Figure 4This is a schematic diagram of the structure of the language-aware selective query fusion module described in Example 1;

[0069] Figure 5 This is a diagram showing the effect of the model described in Example 2 being applied to common open vocabulary object detection;

[0070] Figure 6 This is a diagram showing the effect of applying the model described in Example 3 to long-tail category open vocabulary target detection. DETAILED DESCRIPTION

[0071] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;

[0072] In order to better illustrate the present embodiment, some parts in the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product;

[0073] It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0074] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0075] Example 1

[0076] This embodiment provides a unified open vocabulary detection method based on language-aware selective query fusion, such as Figure 1 As shown, including:

[0077] S1: obtaining different types of data sources and preprocessing them to obtain unified type data, wherein the unified type data includes input images and unified language text, and constructing a unified open vocabulary detection model for language-aware selective query fusion, wherein the model includes an image encoder, a text encoder, a transformer encoder and a transformer decoder, wherein the transformer decoder includes a language-aware query selection module, a language-aware query fusion module, a category mapping layer and a target box regression layer;

[0078] Data sources include detection type data, positioning type data and graphic type data, such as Figure 3 The following is a schematic diagram of three different types of data. Each type of data has its own annotation format. The detection type data includes the input image, the image target box category label and the image target box coordinates. The positioning type data includes the input image, the image target box description text and the image target box coordinates. The image-text type data includes the input image, the image text description and the image target box coordinates. The target box coordinates of the image-text data are the edge coordinates of the image.

[0079] For detection data, the image target box category label is processed into a sentence as a unified language text, such as "a photo of {category}". The image target box description text of positioning data is directly used as a unified language text, and the image text description of image-text data is directly used as a unified language text;

[0080] Unified data integrates different data sources, integrating detection data, positioning data and image data to achieve end-to-end training model. Use triple data format of unified type data Represents different types of data sources, where x∈R 3×H×W Represents the input image, H and W are the height and width of the image respectively, is the image target box coordinate, represents the unified language text, and n is the number of target boxes.

[0081] S2: Input the input image into the image encoder to obtain an image feature vector, and input the unified language text into the text encoder to obtain a text feature vector;

[0082] S3: Obtain a refined image feature vector according to the image feature vector;

[0083] The position feature vector corresponding to the image feature vector is obtained, and the image feature vector and the position feature vector are input into the transformer encoder to obtain a refined image feature vector.

[0084] The transformer encoder comprises N transformer encoder layers connected in sequence.

[0085] S4: inputting the refined image feature vector and the text feature vector into a language-aware query selection module to obtain a dynamically selected content query feature vector and a statically learnable content query feature vector;

[0086] The initial dynamically selected content query feature vector and the statically learnable content query feature vector are obtained, the similarity between the refined image feature vector and the text feature vector is calculated, and the similarity is sorted in descending order. The first K most relevant refined image feature vectors are taken to update the initial dynamically selected content query feature vector to obtain the dynamically selected content query feature vector, and the first K most relevant feature vector indexes are taken to update the initial statically learnable content query feature vector to obtain the statically learnable content query feature vector.

[0087] The first output terminal of the language-aware query selection module outputs a dynamically selected content query feature vector, and the second output terminal outputs a statically learnable content query feature vector.

[0088] S5: inputting the refined image feature vector, the dynamically selected content query feature vector, and the static learnable content query feature vector into a language-aware query fusion module to obtain a fused content query feature vector;

[0089] like Figure 4 As shown, the language-aware query fusion module has three connection modes: post-fusion, mid-fusion and pre-fusion.

[0090] For post-fusion structures, such as Figure 4 As shown in (a), the language-aware query fusion module includes M sequentially connected post-fusion query fusion submodules;

[0091] Each of the post-fusion query fusion submodules includes a first self-attention layer, a first cross-attention layer, a second cross-attention layer, a first tanh gating layer, a first splicing layer, a first feedforward neural network layer, a second tanh gating layer, a second splicing layer, and a second feedforward neural network layer;

[0092] The output end of the first self-attention layer is connected to the third input end of the first cross-attention layer, and the output ends of the first cross-attention layer are connected to the third input end of the second cross-attention layer and the input end of the first concatenation layer;

[0093] The second cross attention layer, the first tanh gating layer, the first splicing layer, the first feedforward neural network layer, the second tanh gating layer, the second splicing layer, and the second feedforward neural network layer are connected in sequence;

[0094] The output end of the first splicing layer is also connected to the input end of the second splicing layer;

[0095] The input end of the first self-attention layer of the first post-fusion query fusion submodule is connected to the second output end of the language-aware query selection module; the input end of the first self-attention layer of the jth post-fusion query fusion submodule is connected to the output end of the second feedforward neural network layer of the j-1th post-fusion query fusion submodule, j=2,…,M;

[0096] The first input terminal and the second input terminal of the first cross-attention layer of the i-th post-fusion query fusion submodule are both connected to the output terminal of the transform encoder; the first input terminal and the second input terminal of the second cross-attention layer of the i-th post-fusion query fusion submodule are both connected to the first output terminal of the language-aware query selection module; i=1,2,…,M.

[0097] For the medium fusion structure, such as Figure 4 As shown in (b), the language-aware query fusion module includes M sequentially connected middle fusion query fusion submodules;

[0098] Each of the fusion query fusion submodules includes a first self-attention layer, a first cross-attention layer, a first tanh gating layer, a first splicing layer, a first feedforward neural network layer, a second tanh gating layer, a second splicing layer, a second cross-attention layer, and a second feedforward neural network layer;

[0099] The output end of the first self-attention layer is connected to the third input end of the first cross-attention layer and the input end of the first concatenation layer respectively;

[0100] The first cross attention layer, the first tanh gating layer, the first splicing layer, the first feedforward neural network layer, the second tanh gating layer, and the second splicing layer are connected in sequence; the output end of the first splicing layer is connected to the input end of the second splicing layer, the output end of the second splicing layer is connected to the third input end of the second cross attention layer, and the output end of the second cross attention layer is connected to the input end of the second feedforward neural network layer;

[0101] The input end of the first self-attention layer of the first fusion query fusion submodule is connected to the second output end of the language-aware query selection module; the input end of the first self-attention layer of the j-th fusion query fusion submodule is connected to the output end of the second feedforward neural network layer of the j-1-th fusion query fusion submodule, j=2,…,M;

[0102] The first input terminal and the second input terminal of the second cross-attention layer of the i-th fused query fusion submodule are both connected to the output terminal of the transform encoder; the first input terminal and the second input terminal of the first cross-attention layer of the i-th fused query fusion submodule are both connected to the first output terminal of the language-aware query selection module; i=1,2,…,M.

[0103] For the pre-fusion structure, such as Figure 4 As shown in (c), the language-aware query fusion module includes M sequentially connected pre-fusion query fusion sub-modules;

[0104] Each of the front-fusion query fusion submodules includes a first cross attention layer, a first tanh gating layer, a first splicing layer, a first feedforward neural network layer, a second tanh gating layer, a second splicing layer, a first self-attention layer, a second cross attention layer, and a second feedforward neural network layer;

[0105] The first cross attention layer, the first tanh gating layer, the first splicing layer, the first feedforward neural network layer, the second tanh gating layer, and the second splicing layer are connected in sequence; the input end of the first cross attention layer is connected to the input end of the first splicing layer, the output end of the first splicing layer is connected to the input end of the second splicing layer, the output end of the second splicing layer is connected to the input end of the first self-attention layer, the output end of the first self-attention layer is connected to the third input end of the second cross attention layer, and the output end of the second cross attention layer is connected to the input end of the second feedforward neural network layer;

[0106] The first and second input ends of the first cross attention layer of the first front-fusion query fusion submodule are both connected to the first output end of the language-aware query selection module; the third input end of the first cross attention layer of the j-th front-fusion query fusion submodule is connected to the output end of the second feedforward neural network layer of the j-1-th front-fusion query fusion submodule, j=2,…,M;

[0107] The first input terminal and the second input terminal of the second cross-attention layer of the i-th front-fusion query fusion submodule are both connected to the output terminal of the transform encoder; the third input terminal of the first cross-attention layer of the i-th front-fusion query fusion submodule is connected to the second output terminal of the language-aware query selection module; i=1,2,…,M.

[0108] S6: Input the fused content query feature vector into the category mapping layer and the target box regression layer respectively to obtain the corresponding predicted category and target box predicted coordinates;

[0109] S7: Construct a total loss function based on the predicted category and the predicted coordinates of the target box, train the model, and obtain a trained language-aware, selective query-fused unified open vocabulary detection model;

[0110] The method for determining the total loss function includes:

[0111] L=L cls +α*L box +β*L giou +γ*L dn

[0112] Among them, L cls represents the classification loss, L box represents the box regression loss, L giou represents the giou loss, L dn represents the denoising loss, α, β and γ represent L box , L giou and L dn The weight factor of

[0113] The method for determining the classification loss, box regression loss, GIOU loss and denoising loss includes:

[0114] L cls =SigmoidFocal(S,GT cls )

[0115] L box =||B-GT box ||

[0116] L giou =GIoU(S,GT cls )

[0117] L dn =Denoisy(S noisy ,GT cls )+Denoisy(B noisy ,GT box )

[0118] Among them, SigmoidFocal represents the SigmoidFocal loss function, S∈R QxC represents the predicted category alignment score, Q is the number of filtered text feature vectors, C represents the number of unified language text categories, O represents the predicted category, E t represents the text feature vector, Indicates E t The transpose of GT cls ∈R Q×C represents the true category, B represents the target box prediction coordinates, GT box ∈R Q×4 represents the real coordinates of the target box, S noisy and B noisy It is the predicted category and target box predicted coordinates after adding noise; GIoU represents the giou processing function, and Denoisy represents the denoising function.

[0119] S8: Obtain the data to be detected, input it into the trained language-aware selective query fusion unified open vocabulary detection model, and output the classification results and target box coordinates.

[0120] In the specific implementation process, the multi-type data is first processed into a unified type to obtain the unified features of multiple data sources. Then, the input image and the unified language text are fed into the corresponding image encoder Φ I and text encoder Φ T Extract image feature vector E i and text feature vector E t After the feature vector is extracted and flattened, the image feature vector and the corresponding position feature vector E pInput to the converter encoder Φ Enc The refined image feature vector E is obtained enc In order to ensure the relevance between the image feature vector and the unified language text, this embodiment adopts a language-aware query selection module Φ QS , this module can select the image feature vector related to the text feature vector. The selected image feature vector is used as the dynamically selected content query feature vector Q sc The initialization features and the static learnable content query feature vector Q lc Input together into the language-aware query fusion module Φ QF The fusion is performed and the output fusion content query feature vector is input into the target box regression layer F r and the category mapping layer F c , to predict the target box category and its corresponding target box coordinates.

[0121] In the language-aware query fusion module, this embodiment adopts a gated fusion mechanism to gradually fuse the content query feature vector to enhance its semantic perception ability while retaining the original semantic information of the content query. enc and the content query feature vector Q selected in the previous step sc And a static learnable content query feature vector Q lc As input, the fused content query feature vector Q is then obtained through the gated cross attention layer and the feedforward neural network layer. sf The language-aware query fusion module is repeated M times to gradually learn to fuse the embedding relationship between language and image.

[0122] The language-aware query selection module in the unified open vocabulary detection model of language-aware selective query fusion proposed in this embodiment can be expressed by the following formula:

[0123] E qs ,E ps =TopRank(Q,E enc ,E t )

[0124] Q sc =E qs

[0125] Among them, TopRank is a parameter-free sorting operation, which is based on the encoded image feature vector E enc and text feature vector E t The similarity from E enc Select the most similar Q image feature vectors E qs , and the corresponding position feature vector E ps , respectively for Qsc and Q lc Initialize.

[0126] The processing process of the language-aware query fusion module of the post-fusion structure is expressed as follows:

[0127]

[0128] in, represents the output result of the first cross attention layer of the i-th post-fusion query fusion submodule, represents the output result of the first concatenation layer of the i-th post-fusion query fusion submodule, represents the output result of the second concatenation layer of the i-th post-fusion query fusion submodule, represents the output result of the second feedforward neural network layer of the i-th post-fusion query fusion submodule. The language-aware query fusion module of the post-fusion structure includes M post-fusion query fusion submodules, that is, the processing process will be repeated M times. tanh represents the linear activation function, Φ Att1 represents the self-attention layer processing function, Φ Att2 represents the cross attention layer processing function, Φ FFW represents the feedforward neural network layer processing function, α attn and α ffw represents a learnable parameter. In order to ensure the consistency of training with the original detector framework and gradually incorporate language-aware context into the query, α attn and α ffw Initialized to zero.

[0129] like Figure 2 As shown in FIG. 1 , it is a system block diagram of the open vocabulary detection method proposed in this embodiment. The whole process of the unified open vocabulary detection model reasoning proposed in this embodiment is expressed by the formula:

[0130]

[0131] E enc =Φ Enc (E i )

[0132] Q sc =Φ QS (E enc ,E t ),Q sf =Φ QF (E enc ,Q sc ,Q lc )

[0133] O=F c (Q sf),B=F r (Q sf )

[0134]

[0135] in, Indicates E t During the training and optimization process of the model, the predicted category alignment score S∈R QxC By calculating the predicted category O and The similarity is obtained. The real coordinates of the target box GT box ∈R Q×4 and the true category GT cls ∈R Q×C Constructed by bipartite graph matching algorithm. Classification loss L cls Use the predicted category alignment score S and the true category GT cls ∈R Q×C Calculated. Regression loss L reg Use the target box to predict the coordinates B∈R Q×4 And the real coordinates of the target box GT box Calculate the regression loss L reg Including box loss L box and giou loss L giou .

[0136] Example 2

[0137] Based on Example 1, this example uses a large-scale object detection dataset Objects365 to pre-train the model for common categories of open vocabulary object detection. Figure 5 As shown in the figure, the model is applied to the open vocabulary object detection effect of common categories. The first column is the visualization result of the image and annotation, and the second column is the prediction result of the OV-DINO method proposed in this embodiment. It can be seen that the OV-DINO proposed in this embodiment can accurately detect the objects in the image and can discover additional real categories.

[0138] Example 3

[0139] This embodiment adds the GoldG dataset for pre-training on the basis of Embodiment 2. The GoldG dataset contains rich category labels and images, which helps the model learn a wider range of semantic information and detect open vocabulary objects for long-tail categories. The introduction of the GoldG dataset enhances the model's understanding of fine-grained categories. Through the positioning information in the GoldG dataset, the model can improve the ability to locate objects in images.

[0140] like Figure 6As shown in the figure, the model is applied to the open vocabulary object detection effect of the long-tail category, showing the detection effect of more than 1200 categories of images. The detection effect of the four pictures shows that the OV-DINO proposed in this embodiment has accurate long-tail category detection capabilities. It can be seen that the model shows better performance when processing categories with long-tail distribution, and can more accurately identify and locate diverse objects.

[0141] The same or similar reference numerals correspond to the same or similar components;

[0142] The terms used in the drawings to describe positional relationships are only used for illustrative purposes and should not be construed as limiting this patent;

[0143] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A unified open vocabulary detection method based on language-aware selective query fusion, characterized in that: include: S1: obtaining different types of data sources and preprocessing them to obtain unified type data, wherein the unified type data includes input images and unified language text, and constructing a unified open vocabulary detection model for language-aware selective query fusion, wherein the model includes an image encoder, a text encoder, a transformer encoder and a transformer decoder, wherein the transformer decoder includes a language-aware query selection module, a language-aware query fusion module, a category mapping layer and a target box regression layer; S2: Input the input image into the image encoder to obtain an image feature vector, and input the unified language text into the text encoder to obtain a text feature vector; S3: Obtain a refined image feature vector according to the image feature vector; S4: inputting the refined image feature vector and the text feature vector into a language-aware query selection module to obtain a dynamically selected content query feature vector and a statically learnable content query feature vector; S5: inputting the refined image feature vector, the dynamically selected content query feature vector, and the static learnable content query feature vector into a language-aware query fusion module to obtain a fused content query feature vector; S6: Input the fused content query feature vector into the category mapping layer and the target box regression layer respectively to obtain the corresponding predicted category and target box predicted coordinates; S7: Construct a total loss function based on the predicted category and the predicted coordinates of the target box, train the model, and obtain a trained language-aware, selective query-fused unified open vocabulary detection model; S8: Obtain the data to be detected, input it into the trained language-aware selective query fusion unified open vocabulary detection model, and output the classification results and target box coordinates.

2. The unified open vocabulary detection method based on language-aware selective query fusion according to claim 1, characterized in that: In S1, the obtaining of different types of data sources and preprocessing to obtain uniform type of data includes: The data sources include detection type data, positioning type data and image-text type data. The detection type data includes input image, image target frame category label and image target frame coordinates. The positioning type data includes input image, image target frame description text and image target frame coordinates. The image-text type data includes input image, image text description and image target frame coordinates. The target frame coordinates of the image-text data are the edge coordinates of the image. For detection type data, the image target box category label is processed into a sentence as a unified language text, the image target box description text of the positioning type data is directly used as the unified language text, and the image text description of the image-text type data is directly used as the unified language text; Using triple data format of uniform type data Represents different types of data sources, where x∈R 3×H×W Represents the input image, H and W are the height and width of the image respectively, is the image target box coordinate, represents the unified language text, and n is the number of target boxes.

3. The unified open vocabulary detection method based on language-aware selective query fusion according to claim 2, characterized in that: In S3, obtaining a refined image feature vector according to the image feature vector includes: The position feature vector corresponding to the image feature vector is obtained, and the image feature vector and the position feature vector are input into the transformer encoder to obtain a refined image feature vector.

4. The unified open vocabulary detection method based on language-aware selective query fusion according to claim 3, characterized in that: The transformer encoder comprises N transformer encoder layers connected in sequence.

5. The unified open vocabulary detection method based on language-aware selective query fusion according to claim 3, characterized in that: In S4, the refined image feature vector and the text feature vector are input into a language-aware query selection module to obtain a dynamically selected content query feature vector and a statically learnable content query feature vector, including: The initial dynamically selected content query feature vector and the statically learnable content query feature vector are obtained, the similarity between the refined image feature vector and the text feature vector is calculated, and the similarity is sorted in descending order. The first K most relevant refined image feature vectors are taken to update the initial dynamically selected content query feature vector to obtain the dynamically selected content query feature vector, and the first K most relevant feature vector indexes are taken to update the initial statically learnable content query feature vector to obtain the statically learnable content query feature vector.

6. The unified open vocabulary detection method based on language-aware selective query fusion according to claim 5, characterized in that: The language-aware query fusion module includes M sequentially connected post-fusion query fusion submodules; Each of the post-fusion query fusion submodules includes a first self-attention layer, a first cross-attention layer, a second cross-attention layer, a first tanh gating layer, a first splicing layer, a first feedforward neural network layer, a second tanh gating layer, a second splicing layer, and a second feedforward neural network layer; The output end of the first self-attention layer is connected to the third input end of the first cross-attention layer, and the output ends of the first cross-attention layer are connected to the third input end of the second cross-attention layer and the input end of the first concatenation layer; The second cross attention layer, the first tanh gating layer, the first splicing layer, the first feedforward neural network layer, the second tanh gating layer, the second splicing layer, and the second feedforward neural network layer are connected in sequence; The output end of the first splicing layer is also connected to the input end of the second splicing layer; The input end of the first self-attention layer of the first post-fusion query fusion submodule is connected to the second output end of the language-aware query selection module; the input end of the first self-attention layer of the jth post-fusion query fusion submodule is connected to the output end of the second feedforward neural network layer of the j-1th post-fusion query fusion submodule, j=2,…,M; The first input terminal and the second input terminal of the first cross-attention layer of the i-th post-fusion query fusion submodule are both connected to the output terminal of the transform encoder; the first input terminal and the second input terminal of the second cross-attention layer of the i-th post-fusion query fusion submodule are both connected to the first output terminal of the language-aware query selection module; i=1,2,…,M.

7. The unified open vocabulary detection method based on language-aware selective query fusion according to claim 5, characterized in that: The language-aware query fusion module includes M sequentially connected middle-fusion query fusion submodules; Each of the fusion query fusion submodules includes a first self-attention layer, a first cross-attention layer, a first tanh gating layer, a first splicing layer, a first feedforward neural network layer, a second tanh gating layer, a second splicing layer, a second cross-attention layer, and a second feedforward neural network layer; The output end of the first self-attention layer is connected to the third input end of the first cross-attention layer and the input end of the first concatenation layer respectively; The first cross attention layer, the first tanh gating layer, the first splicing layer, the first feedforward neural network layer, the second tanh gating layer, and the second splicing layer are connected in sequence; the output end of the first splicing layer is connected to the input end of the second splicing layer, the output end of the second splicing layer is connected to the third input end of the second cross attention layer, and the output end of the second cross attention layer is connected to the input end of the second feedforward neural network layer; The input end of the first self-attention layer of the first fusion query fusion submodule is connected to the second output end of the language-aware query selection module; the input end of the first self-attention layer of the j-th fusion query fusion submodule is connected to the output end of the second feedforward neural network layer of the j-1-th fusion query fusion submodule, j=2,…,M; The first input terminal and the second input terminal of the second cross-attention layer of the i-th fused query fusion submodule are both connected to the output terminal of the transform encoder; the first input terminal and the second input terminal of the first cross-attention layer of the i-th fused query fusion submodule are both connected to the first output terminal of the language-aware query selection module; i=1,2,…,M.

8. The unified open vocabulary detection method based on language-aware selective query fusion according to claim 5, characterized in that: The language-aware query fusion module includes M sequentially connected pre-fusion query fusion submodules; Each of the front-fusion query fusion submodules includes a first cross attention layer, a first tanh gating layer, a first splicing layer, a first feedforward neural network layer, a second tanh gating layer, a second splicing layer, a first self-attention layer, a second cross attention layer, and a second feedforward neural network layer; The first cross attention layer, the first tanh gating layer, the first splicing layer, the first feedforward neural network layer, the second tanh gating layer, and the second splicing layer are connected in sequence; the input end of the first cross attention layer is connected to the input end of the first splicing layer, the output end of the first splicing layer is connected to the input end of the second splicing layer, the output end of the second splicing layer is connected to the input end of the first self-attention layer, the output end of the first self-attention layer is connected to the third input end of the second cross attention layer, and the output end of the second cross attention layer is connected to the input end of the second feedforward neural network layer; The first and second input ends of the first cross attention layer of the first front-fusion query fusion submodule are both connected to the first output end of the language-aware query selection module; the third input end of the first cross attention layer of the j-th front-fusion query fusion submodule is connected to the output end of the second feedforward neural network layer of the j-1-th front-fusion query fusion submodule, j=2,…,M; The first input terminal and the second input terminal of the second cross-attention layer of the i-th front-fusion query fusion submodule are both connected to the output terminal of the transform encoder; the third input terminal of the first cross-attention layer of the i-th front-fusion query fusion submodule is connected to the second output terminal of the language-aware query selection module; i=1,2,…,M.

9. The unified open vocabulary detection method based on language-aware selective query fusion according to claim 5, characterized in that: In S7, the method for determining the total loss function includes: L=L cls +a*L box +β*L giou +γ*L dn Among them, L cls represents the classification loss, L box represents the box regression loss, L giou represents the giou loss, L dn represents the denoising loss, α, β and γ represent L box , L giou and L dn The weight factor of The method for determining the classification loss, box regression loss, GIOU loss and denoising loss includes: L cls =SigmoidFocal(S,GT cls ) IT box =||B-GT box || L giou =GIoU(S,GT cls ) L dn =Denoisy(S noisy ,GT cls )+Denoisy(B noisy ,GT box ) Among them, SigmoidFocal represents the SigmoidFocal loss function, S∈R QxC represents the predicted category alignment score, Q is the number of filtered text feature vectors, C represents the number of unified language text categories, O represents the predicted category, E t represents the text feature vector, Indicates E t The transpose of GT cls ∈R Q×C represents the true category, B represents the target box prediction coordinates, GT box ∈R Q ×4 represents the real coordinates of the target box, S noisy and B noisy It is the predicted category and target box predicted coordinates after adding noise; GIoU represents the giou processing function, and Denoisy represents the denoising function.

10. A unified open vocabulary detection system based on language-aware selective query fusion, used to implement the unified open vocabulary detection method according to any one of claims 1 to 9, characterized in that: include: A data acquisition module is used to acquire different types of data sources and perform preprocessing to obtain unified type data, wherein the unified type data includes input images and unified language text, and construct a unified open vocabulary detection model for language-aware selective query fusion, wherein the model includes an image encoder, a text encoder, a transformer encoder and a transformer decoder, wherein the transformer decoder includes a language-aware query selection module, a language-aware query fusion module, a category mapping layer and a target box regression layer; A text and image encoding module, used for inputting an input image into an image encoder to obtain an image feature vector, and inputting a unified language text into a text encoder to obtain a text feature vector; A refined image feature vector acquisition module, used to obtain a refined image feature vector according to the image feature vector; A query selection processing module, used for inputting the refined image feature vector and the text feature vector into the language-aware query selection module to obtain a dynamically selected content query feature vector and a statically learnable content query feature vector; A query fusion processing module, used for inputting the refined image feature vector, the dynamically selected content query feature vector and the static learnable content query feature vector into the language-aware query fusion module to obtain a fused content query feature vector; The model prediction module is used to input the fused content query feature vector into the category mapping layer and the target box regression layer respectively to obtain the corresponding predicted category and target box predicted coordinates; A model training module is used to construct a total loss function based on the predicted category and the predicted coordinates of the target box, train the model, and obtain a trained language-aware, selective query-fused unified open vocabulary detection model; The classification result and target box coordinate acquisition module is used to obtain the data to be detected, input it into the trained language-aware selective query fusion unified open vocabulary detection model, and output the classification result and target box coordinates.