Open-vocabulary detection method and apparatus
Patent Information
- Application Number
- PCT/CN2025/146799
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2025-12-29
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025146799_01102026_PF_FP_ABST
Abstract
Description
An open vocabulary detection method and device
[0001] Related applications
[0002] This application claims priority to Chinese Patent Application No. 202510368940.4, filed on March 26, 2025, and incorporates the entire contents of the aforementioned patent application as part of this application. Technical Field
[0003] This disclosure relates to the field of computer vision technology, and in particular to an open vocabulary detection method and apparatus. Background Technology
[0004] In recent years, open vocabulary detection has become a research hotspot, aiming to identify and detect object categories not seen during training. Existing open vocabulary detection methods have significantly improved the generalization ability of models in open environments through multimodal learning and visual-language alignment. However, existing methods still suffer from problems such as reliance on explicit language cues and complex multimodal training, limiting their widespread application in practice. Most open vocabulary detection methods rely on predefined text descriptions or specific cues to guide the model in detection, which limits the model's reasoning ability in scenarios involving completely unknown categories or lacking effective cues.
[0005] Furthermore, existing open vocabulary detection methods typically require simultaneous training of image and text embedding spaces, involving a complex multimodal learning process. This increases the complexity and computational cost of model training and necessitates a large amount of labeled data, limiting their widespread adoption in applications where data is scarce or labeling costs are high. Although open vocabulary detection methods perform well on known and partially unknown categories, their generalization ability needs improvement when facing highly diverse and complex open set environments, especially when dealing with novel and unseen object categories, where detection accuracy and robustness are still not ideal.
[0006] This section is intended to provide background or context for the embodiments of this disclosure set forth in the claims. The description herein is not an admission that it is prior art simply because it is included in this section. Summary of the Invention
[0007] This disclosure provides an open vocabulary detection method and apparatus to address the dependence of traditional object detection methods on fixed category sets, thereby enabling accurate detection and localization of objects of unknown categories in images.
[0008] This disclosure provides an open vocabulary detection method, including:
[0009] Obtain the natural language prompts to be detected and the input image;
[0010] A pre-trained feature fusion model is used to jointly model the visual language of natural language prompts and input images to obtain a set of object names. The feature fusion model is obtained by fusing features based on historical images and historical questions to obtain visual language features, and then iteratively training based on the visual language features.
[0011] The pre-trained open vocabulary detection model detects target objects in the input image based on the input image and the set of object names, and obtains the detection results. The open vocabulary detection model is obtained by locating the predicted bounding boxes in historical images and iteratively training based on visual language features and the predicted bounding boxes.
[0012] In one embodiment, a pre-trained feature fusion model is used to perform joint visual-language modeling of natural language prompts and input images to obtain a set of object names, including:
[0013] Multi-level feature extraction is performed on the input image to obtain its visual features;
[0014] The natural language prompts are parsed using a large language model to generate the linguistic features of the natural language prompts;
[0015] Visual features and linguistic features are fused across modalities to obtain fused visual-linguistic features;
[0016] Visual language features are input into a large language model to generate a set of object names.
[0017] In one embodiment, multi-layer feature extraction is performed on the input image to obtain the visual features of the input image, including:
[0018] Cut the input image into image blocks of a preset step size;
[0019] Image feature sequences are generated by extracting features from image patches using a visual encoder.
[0020] In one embodiment, visual features and linguistic features are fused across modalities to obtain fused visual-linguistic features, including:
[0021] Visual features are compressed using a cross-attention mechanism to obtain compressed visual features;
[0022] Generate encoded visual features based on compressed visual features and 2D absolute position encoding;
[0023] The encoded visual features and linguistic features are fused together using a visual-language adapter to obtain visual-linguistic features.
[0024] In one embodiment, the detection result includes the predicted bounding box of the object and its corresponding predicted category label; the target object in the input image is detected by a pre-trained open vocabulary detection model based on the input image and a set of object names, and the detection result includes:
[0025] The input image is detected based on each object name in the object name set, and a predicted bounding box and its corresponding predicted category label are generated for each target object.
[0026] In one embodiment, the steps of pre-training a visual language model include:
[0027] The training dataset consists of multiple pairs of historical images, historical questions, and real answers.
[0028] Historical images and historical questions are input into a visual language model for feature extraction, resulting in visual and linguistic features.
[0029] Visual-linguistic features are generated by fusing visual and linguistic features through a cross-attention mechanism.
[0030] Visual language features are input into a large language model to generate predicted answers;
[0031] The prediction error between the predicted answer and the true answer is calculated using the first loss function.
[0032] The visual language model is iteratively trained by adjusting model parameters using the Adam optimizer based on prediction errors.
[0033] In one embodiment, the steps of pre-training an open vocabulary detection model include:
[0034] The target object is located on the input image by the target detection probe, and the predicted bounding box of the target object is obtained.
[0035] The visual language features and the predicted bounding boxes of the target objects are classified by a classifier to generate the predicted category labels of the target objects.
[0036] The classification error between the predicted class label and the true class label, and the localization error between the predicted bounding box and the true bounding box are determined by the second loss function.
[0037] The open vocabulary detection model is iteratively trained by adjusting model parameters based on classification and localization errors.
[0038] This disclosure also provides an open vocabulary detection device, including:
[0039] The data acquisition module is used to acquire the natural language prompts and input images to be detected;
[0040] The feature processing module is used to perform visual language joint modeling on natural language prompts and input images through a pre-trained feature fusion model to obtain a set of object names; the feature fusion model is obtained by performing feature fusion processing on historical images and historical questions to obtain visual language features, and then iteratively training based on the visual language features;
[0041] The detection module is used to detect target objects in the input image based on the input image and the set of object names using a pre-trained open vocabulary detection model, and obtain the detection results. The open vocabulary detection model is obtained by locating the predicted bounding boxes in historical images and iteratively training based on visual language features and the predicted bounding boxes.
[0042] This disclosure also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described open vocabulary detection method.
[0043] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described open vocabulary detection method.
[0044] This disclosure also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described open vocabulary detection method.
[0045] Based on the open vocabulary detection method and apparatus provided in this disclosure, by fusing visual features and natural language prompts, it effectively overcomes the dependence of traditional object detection models on fixed category sets, achieving object detection with greater generalization ability. Compared with traditional object detection methods, this disclosure can not only detect predefined categories, but also identify unknown categories based on natural language prompts, thereby significantly improving the flexibility and applicability of the detection model. Visual features of the input image are extracted through a feature fusion model, and combined with a large language model to parse the natural language prompts input by the user. A cross-modal attention mechanism is adopted to fuse visual and linguistic features, and the large language model understands the semantics of the input text to generate object names related to the content of the input image. This method can automatically generate object categories without relying on fixed category labels, effectively improving the model's adaptability in open vocabulary detection tasks. A dual loss optimization strategy, namely classification loss and localization loss, is adopted to ensure the accuracy and robustness of object detection, effectively improving the generalization ability of the detection model. By integrating a pre-trained feature fusion model and an open vocabulary detection model, high detection capability can be maintained even with insufficient data, effectively reducing the dependence on manually labeled data and lowering training costs. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0047] Figure 1 is a flowchart illustrating an open vocabulary detection method in one embodiment of this disclosure;
[0048] Figure 2 is a flowchart illustrating the open vocabulary detection method in another embodiment of this disclosure;
[0049] Figure 3 is a flowchart illustrating the open vocabulary detection method in another embodiment of this disclosure;
[0050] Figure 4 is a flowchart illustrating the open vocabulary detection method in another embodiment of this disclosure;
[0051] Figure 5 is a flowchart illustrating the open vocabulary detection method in another embodiment of this disclosure;
[0052] Figure 6 is a flowchart illustrating the open vocabulary detection method in another embodiment of this disclosure;
[0053] Figure 7 is a flowchart illustrating an open vocabulary detection method in another embodiment of this disclosure;
[0054] Figure 8 is a schematic diagram of the structure of an open vocabulary detection device in one embodiment of this disclosure;
[0055] Figure 9 is a schematic diagram of the open vocabulary detection device in another embodiment of this disclosure;
[0056] Figure 10 is a schematic diagram of the open vocabulary detection device in another embodiment of this disclosure;
[0057] Figure 11 is a schematic diagram of the open vocabulary detection device in another embodiment of this disclosure;
[0058] Figure 12 is a schematic diagram of the open vocabulary detection device in another embodiment of this disclosure;
[0059] Figure 13 is a schematic diagram of the physical structure of the electronic device provided in the embodiment of this disclosure;
[0060] Figure 14 is a flowchart illustrating the open vocabulary detection method according to a specific embodiment of this disclosure. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the embodiments of this disclosure will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of this disclosure and their descriptions are used to explain this disclosure, but are not intended to limit this disclosure.
[0062] The information collected in the technical solution disclosed herein is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and it does not violate public order and good morals. Corresponding operation interfaces are provided for users to choose to authorize or refuse.
[0063] Provide users with corresponding operation entry points, allowing them to choose to agree to or reject the automated decision results; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0064] To automate the identification and localization of multiple objects in images, this disclosure provides an Open-Vocabulary Detection method. By integrating a Visual Language Model (VLM) and an Open-Vocabulary Detection Model (GLIP), it achieves fully automated operation from image understanding to object detection. The Visual Language Model is responsible for generating a set of object names based on the input image and natural language prompts, while the Open-Vocabulary Detection Model uses these generated object names and input image information for accurate object identification and localization. This Open-Vocabulary Detection method not only improves the generalization ability of object detection but also significantly simplifies the user interaction process and enhances the system's adaptability and robustness in open-set environments.
[0065] Figure 1 is a flowchart illustrating the open vocabulary detection method in this embodiment of the present disclosure. Specifically, this method is applied to the server side. In practice, as shown in Figure 1, the open vocabulary detection method includes steps 101 to 103.
[0066] Step 101: Obtain the natural language prompts to be detected and the input image.
[0067] Step 102: Perform visual language joint modeling on natural language prompts and input images through a pre-trained feature fusion model to obtain a set of object names; the feature fusion model is obtained by performing feature fusion processing on historical images and historical questions to obtain visual language features, and then iteratively training based on the visual language features.
[0068] Step 103: Detect objects in the input image using a pre-trained open vocabulary detection model based on the input image and the set of object names, and obtain the detection results; the open vocabulary detection model is obtained by locating predicted bounding boxes in historical images and iteratively training based on visual language features and predicted bounding boxes.
[0069] As shown in Figure 1, in this embodiment, a pre-trained feature fusion model is used to jointly model the visual language of natural language prompts and the input image. The feature fusion model can automatically generate object names in the image based on simple natural language prompts, obtaining a set of object names without the need for predefined category labels. This reduces reliance on large-scale labeled data, lowers labeling costs, and improves model training efficiency. Through feature fusion and dynamic classification mechanisms, multi-scale feature extraction is achieved in complex scenes. The fusion of visual features and semantic information yields visual language features, improving the detection capability for open-set targets. The pre-trained open-vocabulary detection model, based on the object name set of the feature fusion model, detects objects in the input image, overcoming the category limitations of traditional target detection methods and improving the model's generalization ability.
[0070] In the embodiments of this disclosure, an input image to be detected can be acquired through an image acquisition device, and natural language prompts can be received through a human-computer interaction interface. This method can be applied to target detection in open scenes. For example, it can be applied to autonomous driving environmental perception systems to identify and locate obstacles in the road environment that are not predefined or not present in the training data in real time. As shown in Figure 1, each step is explained in detail below.
[0071] Step 101: Obtain the natural language prompts to be detected and the input image.
[0072] Existing open-vocabulary object detection methods typically rely on predefined category labels or complex language prompts (such as "red car") provided by the user to guide the model in object recognition. These methods have limitations when handling complex scenes or unknown object categories. In this disclosure, the server acquires the input image I to be detected and the natural language prompt P. The input image I and the natural language prompt P are then fed into a pre-trained feature fusion model. The feature fusion model can automatically identify and generate a set of names for all objects in the input image I using simple natural language prompts (such as "Please tell me what's in the picture"), thus achieving more flexible and intelligent object detection. This greatly simplifies the user's operation process and improves the system's usability and user experience. Users do not need professional computer vision knowledge to interact efficiently with the system through natural language prompts.
[0073] Step 102: Perform visual language joint modeling on natural language prompts and input images through a pre-trained feature fusion model to obtain a set of object names; the feature fusion model is obtained by performing feature fusion processing on historical images and historical questions to obtain visual language features, and then iteratively training based on the visual language features.
[0074] Specifically, the feature fusion model (trained based on the visual language model VLM) generates a set O = {o1, o2, ..., on} containing the names of all objects in the input image I after receiving the input image I and the natural language prompt P through joint visual language modeling.
[0075] The feature fusion model formula is as follows: O=VLM(I,P) (1)
[0076] Where O is the set of object names, VLM is the visual language model, I is the input image, P is the natural language prompt, and VLM(I,P) represents the set of object names O generated by the visual language model VLM under the conditions of input image I and natural language prompt P.
[0077] In some embodiments of this disclosure, as shown in FIG2, step 102 includes steps 201 to 204.
[0078] Step 201: Perform multi-layer feature extraction on the input image I to obtain the visual features of the input image I.
[0079] The feature fusion model can use the visual encoder in the Qwen-VL model (which employs a Vision Transformer (ViT) architecture) to extract high-level visual features from the input image I.
[0080] As shown in Figure 3, step 201 may include steps 301 to 302.
[0081] Step 301: Cut the input image I into image blocks with a preset step size.
[0082] Specifically, the input image I is divided into image blocks of 14 steps in length. The preset step length can be set according to the needs of those skilled in the art, and this disclosure is not limited thereto.
[0083] Step 302: Extract features from image patches using a visual encoder to generate an image feature sequence.
[0084] Specifically, a visual feature sequence, i.e., visual features F, is generated by extracting features from each 14-step image patch using a visual encoder (ViT). I The image feature sequence includes information such as object edge features, color features, texture features, and shape features of the image patch.
[0085] The formula corresponding to the visual transformation is as follows: F I =ViT(I) (2)
[0086] Among them, F I F is a sequence of image features generated by a visual encoder. I ∈R L×d L is the length of the image feature sequence, d is the dimension of each image feature vector, and R identifies the real space.
[0087] In one embodiment, R L×d This represents an L×d dimensional matrix (or tensor) indicating that the features of these images are real-valued and stored according to a data structure of L rows and d columns. The length L of the image feature sequence corresponds to the number of image patches whose features are contained in the input image I. The dimension d of each image feature vector indicates how many values are used to represent the features of each image patch.
[0088] Step 202: Parse the natural language prompt P using a large language model to generate the linguistic features F of the natural language prompt P. p .
[0089] Specifically, the large language model uses the Qwen-7B model from the Qwen-VL model to parse the input natural language prompt P and obtain the language features F. p .
[0090] In one embodiment, the natural language prompt P undergoes text preprocessing. This text preprocessing includes tokenization and word embedding.
[0091] Specifically, the natural language prompt P is decomposed into minimal tokens through tokenization. For example, if the input natural language prompt P is "Please tell me what objects are in the image", the tokenized result is ["please", "tell", "I", "image", "in", "have", "what", "object"]. Each minimal token in the natural language prompt P is converted into a high-dimensional word vector through a word embedding model. The word embedding model can be Word2Vec or BERT, etc., and this disclosure is not limited to this.
[0092] Word vectors capture the syntactic structure and semantic information of each minimized word unit and are input into the Qwen-7B model for processing. A self-attention mechanism is used to establish relationships between each word vector. Then, through multi-layer stacking and context modeling, language features F corresponding to the natural language cues P are generated. p .
[0093] In one embodiment, context modeling involves continuously updating the language features F of each word vector through multiple Transformer layers. p This ensures that each word vector not only contains the features of the current word vector but also incorporates contextual information. The output of each Transformer layer is passed to the next Transformer layer as input, progressively generating context-rich language features F. p .
[0094] The Qwen-7B model is trained using a large-scale pre-training corpus, enabling it to accurately parse natural language cues P and generate linguistic features F. p .
[0095] Step 203: Visual feature F I With language features F p Cross-modal fusion is performed to obtain the fused visual language features F. VL .
[0096] Specifically, the visual features F are processed through the Position-aware Vision-Language Adapter in the Qwen-VL model. I and language features F p The feature vectors are fused together to obtain a fixed-length feature vector.
[0097] The feature fusion formula for the Qwen-VL model is as follows: F VL =Adapter(F I ,F p (3)
[0098] Among them, F VL The fused visual language features have a fixed length; F VL ∈R 256 R denotes the real space; F I For visual features, F p It is a linguistic feature.
[0099] In one embodiment, the fused visual language features F VL It contains comprehensive dimensional information of the input image I and natural language prompts P, which is used to further input into a large language model (such as the Qwen-7B model) to generate a set of object names O.
[0100] As shown in Figure 4, step 203 includes steps 401 to 403.
[0101] Step 401: Apply visual features F through a cross-attention mechanismI Compression is performed to obtain the compressed visual features F′ I .
[0102] Specifically, the visual language adapter employs a cross-attention mechanism to enable visual features F I and language features F p Interactions are used to reduce information loss. Based on visual feature F... I and language features F p The compressed visual features F′ are obtained by calculating attention weights and then fusing them in a weighted manner. I .
[0103] In one embodiment, the attention weights are normalized using the Softmax function to ensure that the sum of the weights of the attention distribution is 1.
[0104] Step 402: Based on the compressed visual features F′ I and 2D absolute position encoding P 2D Generate encoded visual features F″ I .
[0105] Specifically, P 2D It is a 2D absolute position code generated based on the (x, y) coordinates in a two-dimensional real space.
[0106] The 2D absolute position encoding formula is as follows: F″=F′+P 2D (4)
[0107] Where F″ is the encoded visual feature, F′ is the compressed visual feature, and P 2D It is a 2D absolute position encoding.
[0108] In one embodiment, 2D absolute position encoding can preserve the spatial information of the input image I, reduce information loss, and even if visual features are compressed, the feature fusion model can still clearly show the position information of different visual features in different regions of the input image I.
[0109] Step 403: Combine the encoded visual features F″ and linguistic features F p Visual language features F are obtained by fusion using a visual language adapter. VL .
[0110] Specifically, the encoded visual features F″ and language features F″ are combined using a visual language adapter. p To integrate.
[0111] The feature fusion formula is as follows: F VL =Adapter(F″,F p (5)
[0112] Among them, F VL The fused visual language features have a fixed length; F VL ∈R 256 R represents the real space; F″ represents the encoded visual feature, F p It is a linguistic feature.
[0113] Step 204: Transfer visual language features F VL Input into the large language model to generate a set of object names O.
[0114] Specifically, the large language model is based on visual language features F VL The formula for generating the set of object names O is as follows: O = QLM(F VL (6)
[0115] Where O is the set of generated object names, F VL The fused visual language features are represented by QLM, which is a large language model of the Qwen-VL model.
[0116] In this embodiment, multi-layer feature extraction is performed on the input image to obtain its visual features. A large language model is used to parse natural language prompts, generating linguistic features for the prompts. The visual and linguistic features are then fused across modally to obtain fused visual-linguistic features. Through these techniques, the feature fusion model can efficiently process the input image and natural language prompts, automatically generating a list of names covering all major objects in the input image, providing accurate detection targets for subsequent open vocabulary detection models. In practical applications, multi-modal fusion via a cross-attention mechanism enables the Qwen-VL model to handle more complex visual scenes and linguistic prompts, and avoids the limitations caused by predefined object categories in traditional detection methods when performing subsequent open vocabulary detection. Since traditional open vocabulary detection methods typically require a large amount of multimodal labeled data for training, this disclosure reduces the dependence on large-scale labeled data through a pre-trained feature fusion model, lowering training costs and enhancing the model's generalization ability and adaptability, making it suitable for application scenarios with scarce data or high labeling costs.
[0117] To improve the generalization ability of the Qwen-VL model, the model was fine-tuned during training based on labeled visual question answering (VQA) data.
[0118] As shown in Figure 5, the training steps of the pre-trained visual language model include steps 501 to 506.
[0119] Step 501: Use multiple sets of historical image I′, historical question Q and real answer A as the training dataset.
[0120] Specifically, during the fine-tuning process, a training dataset is obtained by taking a sample pair of historical image I′, historical question Q, and real answer A from the visual question answering dataset, i.e., {(I′, Q, A)}.
[0121] Step 502: Input the historical image I′ and the historical question Q into the visual language model for feature extraction to obtain the visual features F. I and language features F p .
[0122] Specifically, in each round of model training, the historical image I′ and the historical question Q are input into the Qwen-VL model (i.e., the visual language model) to generate the corresponding predicted answer A′.
[0123] In one embodiment, the input data is preprocessed before model training, including image normalization and text tokenization. Historical images I′ are normalized by adjusting image size, normalization, and data augmentation to adapt the input format to the visual encoder. Historical questions Q are tokenized, converting them into a form understandable by the large language model, thus adapting the text input format to the Qwen-7B model.
[0124] Visual features F are obtained by high-level feature extraction from the preprocessed historical image I′ using the visual encoder in the Qwen-VL model. I The linguistic features F are obtained by analyzing the preprocessed historical question Q using the Qwen-7B model within the Qwen-VL model. p .
[0125] Step 503: Apply visual features F through a cross-attention mechanism I and language features F p Perform feature fusion to generate visual language features F VL .
[0126] Specifically, visual features F are processed through a cross-attention mechanism. I and language features F p Information fusion is performed to generate visual language features F VL That is, it integrates the feature vectors of historical image I′ and historical question Q, enabling the Qwen-VL model to understand the relationship between historical image I′ and historical question Q.
[0127] Step 504: Transfer visual language features F VL The input is fed into a large language model to generate the predicted answer A′.
[0128] Specifically, visual language features F VL The input is fed into the Qwen-7B language model to generate a set of object names O, which is the predicted answer A′.
[0129] Step 505: Calculate the prediction error between the predicted answer A′ and the true answer A using the first loss function.
[0130] Specifically, cross-entropy loss is used to minimize the prediction error of the Qwen-VL model.
[0131] The formula for calculating the first loss function is:
[0132] Among them, P(A) i |I′,Q) represents the Qwen-VL model's prediction of answer A given historical image I′ and historical question Q. i The probability, where N is the number of samples input to the model, L VQA This represents the prediction error of the Qwen-VL model.
[0133] In one embodiment, the loss function maximizes the conditional probability P(A) of the correct answer. i |I′,Q), thereby minimizing the error between the predicted answer A′ of the Qwen-VL model and the true answer A, and thus optimizing the model parameters of the Qwen-VL model.
[0134] Step 506: Based on the prediction error, adjust the model parameters using the Adam optimizer and iteratively train the visual language model.
[0135] Specifically, the prediction error of the Qwen-VL model calculated based on the first loss function is used to automatically adjust the model parameters through the Adam optimizer.
[0136] In one embodiment, the Adam optimizer determines which model parameters need adjustment by calculating the gradient (i.e., rate of change) of the first loss function with respect to each model parameter. The Adam optimizer updates the model parameters that need adjustment using the gradient descent algorithm to minimize the prediction error; and employs adaptive learning rate scheduling and early stopping to prevent overfitting.
[0137] By continuously adjusting the model parameters and iteratively training until the first loss function converges or the preset number of model training rounds is reached, the trained Qwen-VL model is obtained.
[0138] Step 103: Using a pre-trained open vocabulary detection model, detect the target object i in the input image I based on the input image I and the object name set O, obtaining the detection result. The detection result is the predicted bounding box B of the target object i. i and predicted category label C i .
[0139] Step 103 specifically includes: based on each object name O in the object name set O i Detect the input image I and generate the predicted bounding box B for each target object i. i and its corresponding predicted category label C i .
[0140] Specifically, the open vocabulary detection model employs the GLIP model. The input image I and the set of object names O are input into the pre-trained GLIP model, and the object detection head detects each object name O in the set of object names O. i Detection is performed on the input image I, generating the location information (bounding box) of each target object i in the input image I and its corresponding predicted class label C. i .
[0141] As shown in Figure 6, the training steps of the pre-trained open vocabulary detection model include steps 601 to 604.
[0142] Step 601: Locate the target object i on the input image I using the target detection probe to obtain the predicted bounding box of the target object i.
[0143] Before locating objects in the input image I, feature fusion processing is required.
[0144] Specifically, multiple sets of input images I and their corresponding input language text Text are obtained. The input language text Text can be a text description of the input image I or a prompt indicating the detected target object.
[0145] The input image I is encoded using a visual encoder (such as the Vision Transformer) of the GLIP model to obtain the visual features F of the input image I. I Among them, visual feature F I It contains visual information from the input image I.
[0146] The input language text Text is processed by the text encoder of the GLIP model to generate text features F. t .
[0147] Visual feature F I and text features Ft Visual language features F are obtained through deep fusion using a cross-modal attention mechanism. VL .
[0148] The cross-modal fusion formula for the GLIP model is as follows:
[0149] Among them, F VL For visual language features, F I For visual features, F t It is a text feature, W q For the Query weight matrix, W k Let W be the key weight matrix. v This is the Value weight matrix.
[0150] In one embodiment, the Query weight matrix is used to weight the visual features F I Transformed into a query representation; the key weight matrix is used to weight the text features F t Transform into a key representation; the value weight matrix is used to represent the text features F. t This is transformed into a value representation. In the attention mechanism, the correlation between different features can be determined by calculating the dot product between the query and the key and normalizing it using the softmax function. The attention score obtained from the dot product operation is used as a weighted value, ultimately yielding the fused visual-language feature F. VL .
[0151] In one embodiment, the visual encoder and text encoder of the GLIP model can dynamically adjust the semantic weights of visual features based on the input language text, thereby improving the GLIP model's attention to the target region of the input image.
[0152] Specifically, when locating target object i, the target detection probe is used to locate target object i on the input image I, and the predicted bounding box coordinates B of each target object i are generated. i = (x1, y1, x2, y2). Where x1, y1, x2, y2 are the x and y coordinates of the predicted bounding box.
[0153] The formula for object detection is as follows: B i =DetHead(F VL ,i) (9)
[0154] Among them, B i Let DetHead be the predicted bounding box coordinates of target object i, and F be the target detection function. VL For visual language features, i represents the target object.
[0155] In one embodiment, the target detection probe reads visual language features F VL And based on the above target detection formula, the specific position of the target object i is calculated, that is, the predicted bounding box coordinates B. i .
[0156] Step 602: Classify the visual language features F using a classifier VL The predicted bounding box of target object i is used for classification, and the predicted class label C of target object i is generated. i .
[0157] Specifically, the class label C of the target object i is predicted using a classifier. i C i =Classifier(B i (10)
[0158] Among them, B i Let C be the predicted bounding box coordinates of target object i. i The predicted category label for target object i.
[0159] In one embodiment, after the target detection probe predicts the bounding box, the classifier uses the Patch extraction method to extract visual language features F. VL Extract the bounding box B that is predicted i The visual features of the corresponding region are analyzed. This is achieved by combining text features F through a cross-attention mechanism. t The visual language features of a specific region after fusion are obtained. Based on these features, a classifier performs target classification using a fully connected layer (FC) and a softmax function, obtaining the predicted class label C of target object i. i The Patch extraction method is mainly used for Transformer structures (such as ViT). Those skilled in the art can also use other extraction methods to extract features from specific regions, and this disclosure is not limited thereto.
[0160] Step 603: Determine the classification error L between the predicted class label and the true class label using the second loss function. cls and the positioning error L between the predicted bounding box and the true bounding box loc .
[0161] Specifically, a multi-task loss function is used, and the specific calculation formula is as follows:
[0162] in, C is the true category label, and C is the predicted category label. B is the ground truth bounding box, and A is the predicted bounding box. For classification loss function, This is the location loss function.
[0163] In one embodiment, the classification loss function is used to calculate the predicted class label C and the true class label. The error between the predicted bounding box B and the ground truth bounding box. The localization loss function is used to calculate the error between the predicted bounding box B and the ground truth bounding box. The error between them.
[0164] The classification loss function can be the cross-entropy loss function, which ensures that the predicted class label C output by the GLIP model is as close as possible to the true class label.
[0165] Step 604: Based on classification error L cls and positioning error L loc Adjust the model parameters and iteratively train the open vocabulary detection model.
[0166] Specifically, the classification error L of the GLIP model calculated based on the cross-entropy loss function. cls and positioning error L loc The Adam optimizer automatically adjusts model parameters.
[0167] In one embodiment, the Adam optimizer calculates the classification error L separately. cls Gradient of model parameters and localization error L loc The gradient of the model parameters determines which model parameters need adjustment. The Adam optimizer uses the gradient descent algorithm to update the model parameters that need adjustment, improving the classification error L. cls and positioning error L loc The learning rate is kept as small as possible; and adaptive learning rate scheduling and early stopping are employed to prevent overfitting. Model parameters may include: attention weights of the Transformer layer, text encoder weights, fully connected layer weights, and bounding box regression layer weights for the object detection probe, etc.
[0168] By continuously adjusting the model parameters and performing iterative training, until the classification error L is reached... cls and positioning error L loc Training stops when the model converges or reaches the preset number of training rounds, resulting in a trained GLIP model.
[0169] In this embodiment, through multimodal feature fusion and dynamic classification mechanisms, the open vocabulary detection model can accurately locate each target object and classify it according to the object name, ensuring high accuracy and reliability of the detection results. This enhances the recognition ability for complex scenes and diverse objects, improving the robustness and accuracy of detection, especially in images with small objects and high-density object distributions. Through joint visual-language modeling and a dynamic classifier, it can efficiently handle various object detection tasks, adapting to different object categories and complex scenes. A dual-loss optimization strategy, namely classification loss and localization loss, is adopted to ensure the accuracy and robustness of object detection. The classification loss calculates the error between the predicted class label and the true class label based on cross-entropy, ensuring that the model can correctly match the target category. The localization loss optimizes the target bounding box through regression loss and Intersection over Union (IoU) loss, enabling the model to accurately locate target objects in complex scenes. The dual-loss function optimization strategy effectively improves the generalization ability of the detection model, giving it high detection accuracy in different scenarios. By integrating a feature fusion model and an open vocabulary detection model, the entire process from image understanding to object detection is automated. It accurately identifies and locates multiple objects in input images without human intervention, and achieves real-time object detection and localization, meeting the requirements of high availability and low latency applications such as autonomous driving, real-time monitoring, and smart manufacturing. Automated object recognition not only improves detection efficiency and accuracy but also reduces errors and workload caused by human operation, making it suitable for large-scale image processing and real-time monitoring applications.
[0170] Figure 14 is a flowchart illustrating the open vocabulary detection method according to a specific embodiment of this disclosure. The open vocabulary detection method of this disclosure will be further described below with reference to Figure 14:
[0171] Suppose the input image I is a city street scene image containing various objects, such as cars, pedestrians, buildings, and trees. The natural language prompt P provided by the user is "Please tell me what's in the picture."
[0172] S1: Receives the input image I and natural language prompts P, and performs image preprocessing on the input image I. Image preprocessing includes operations such as resizing, normalization, and data augmentation.
[0173] Specifically, the input image I is resized to a uniform size (e.g., 224x224 pixels) to match the input requirements of the feature fusion model and the open vocabulary detection model. The pixel values of the input image I are normalized, scaling them to the range of [0,1] or [-1,1] to improve the stability of model training and inference. Data augmentation processing is performed on the input image I through cropping, rotation, and flipping to increase the diversity of training data and improve the model's generalization ability and robustness. Through these image preprocessing operations, high-quality images can be generated as model input, ensuring efficient model operation and high-accuracy detection.
[0174] S2: Input the preprocessed input image I and the natural language prompt P into the pre-trained feature fusion model (i.e., visual language model) to obtain the set of object names O.
[0175] In one embodiment, in object name O i If model anomalies or uncertainties occur during the generation process, the system will automatically adjust the model parameters or prompt the user to re-enter them to ensure the accuracy and reliability of the detection results.
[0176] In this embodiment, the set of object names O is O = {car, person, building, tree}, and the object names O i These correspond to the cars, pedestrians, buildings, and trees in the input image I, respectively.
[0177] S3: Input the set of object names O generated by the feature fusion model and the input image I into the open vocabulary detection model.
[0178] Specifically, the object name set O and the input image I are passed to the open vocabulary detection model through an interface to ensure that the open vocabulary detection model can accurately obtain the target information. Information synchronization processing is performed on the input image I and the object name set O to avoid information delays or misalignments, ensuring that the open vocabulary detection model can perform detection based on the latest object name set O.
[0179] In one embodiment, the set of object names O provides specific detection targets for the open vocabulary detection model, guiding the open vocabulary detection model to find the corresponding target object i in the input image I.
[0180] S4: Use an open vocabulary detection model to sequentially analyze each object name O in the object name set O. i Detection is performed on the input image I to generate the predicted bounding box B. i and predicted category label C i .
[0181] Specifically, for the object name "car", the object detection probe identifies the location of the car in the input image I and generates the corresponding predicted bounding box B1 and predicted class label C1. For "person", the object detection probe identifies the location of the pedestrian in the input image I and generates the corresponding predicted bounding box B2 and predicted class label C2. For "building", the object detection probe identifies the location of the building in the input image I and generates the corresponding predicted bounding box B3 and predicted class label C3. For "tree", the object detection probe identifies the location of the tree in the input image I and generates the corresponding predicted bounding box B4 and predicted class label C4. Through multimodal feature fusion and dynamic classification mechanisms, the object detection probe can accurately locate the position of target objects, effectively detect the edges and contours of large buildings, and accurately identify trees in complex environmental backgrounds. Through contextual modeling mechanisms, it can distinguish different pedestrians.
[0182] S5: Display the detection results as image annotations on the input image I, or list the predicted category label C for each target object i in text form. i .
[0183] Specifically, the target object "car" is located between coordinates (100, 150) and (200, 250), and its category label is "car"; the target object "person" is located between coordinates (300, 400) and (350, 450), and its category label is "person"; the target object "building" is located between coordinates (50, 50) and (400, 300), and its category label is "building"; the target object "tree" is located between coordinates (500, 100) and (550, 200), and its category label is "tree". The detection results are visually displayed using bounding boxes and a text list on the input image I, allowing users to quickly understand the distribution and specific location information of objects in the input image I.
[0184] In one embodiment, the detection results are exported as data files in various formats so that users can perform further analysis and applications.
[0185] This disclosure also provides an open vocabulary detection device, as shown in the following embodiment. Since the principle behind this device's problem-solving is similar to that of the open vocabulary detection method, its implementation can be found in the implementation of the open vocabulary detection method; repeated details will not be elaborated further.
[0186] As shown in Figure 7, the open vocabulary detection device 700 includes a data acquisition module 701, a feature processing module 702, and a detection module 703.
[0187] The data acquisition module 701 is used to acquire the natural language prompts and input images to be detected.
[0188] The feature processing module 702 is used to perform visual language joint modeling on natural language prompts and input images through a pre-trained feature fusion model to obtain a set of object names; the feature fusion model is obtained by performing feature fusion processing on historical images and historical questions to obtain visual language features, and then iteratively training based on the visual language features.
[0189] The detection module 703 is used to detect target objects in the input image based on the input image and the set of object names using a pre-trained open vocabulary detection model, and obtain the detection results. The open vocabulary detection model is obtained by locating the predicted bounding box in historical images and iteratively training based on visual language features and the predicted bounding box.
[0190] As shown in Figure 8, the feature processing module 702 includes a first feature extraction unit 801, a semantic parsing unit 802, a first feature fusion unit 803, and a name generation unit 804.
[0191] The first feature extraction unit 801 is used to perform multi-layer feature extraction on the input image to obtain the visual features of the input image.
[0192] The semantic parsing unit 802 is used to parse natural language prompts through a large language model and generate the linguistic features of the natural language prompts.
[0193] The first feature fusion unit 803 is used to perform cross-modal fusion of visual features and linguistic features to obtain fused visual-linguistic features.
[0194] The name generation unit 804 is used to input visual language features into a large language model to generate a set of object names.
[0195] As shown in Figure 9, the first feature extraction unit 801 includes an image segmentation subunit 901 and an image extraction subunit 902.
[0196] The image segmentation subunit 901 is used to cut the input image into image blocks with a preset step size.
[0197] The image extraction subunit 902 is used to extract features from image blocks using a visual encoder to generate an image feature sequence.
[0198] As shown in Figure 10, the first feature fusion unit 803 includes a feature compression subunit 1001, a feature encoding subunit 1002, and a feature fusion subunit 1003.
[0199] The feature compression subunit 1001 is used to compress visual features through a cross-attention mechanism to obtain compressed visual features.
[0200] The feature encoding subunit 1002 is used to generate encoded visual features based on the compressed visual features and 2D absolute position encoding.
[0201] The feature fusion subunit 1003 is used to fuse the encoded visual features and language features through a visual language adapter to obtain visual language features.
[0202] The detection module 703 is specifically used to detect the input image based on each object name in the object name set, and generate a predicted bounding box for each target object and its corresponding predicted category label.
[0203] As shown in Figure 11, the open vocabulary detection device 700 also includes a first model training module 704. The first model training module 704 includes a dataset acquisition unit 1101, a second feature extraction unit 1102, a second feature fusion unit 1103, an answer generation unit 1104, a first error determination unit 1105, and a first parameter adjustment unit 1106.
[0204] The dataset acquisition unit 1101 is used to use multiple sets of historical images, historical questions and real answer sample pairs as training datasets.
[0205] The second feature extraction unit 1102 is used to input historical images and historical questions into the visual language model for feature extraction, thereby obtaining visual features and language features.
[0206] The second feature fusion unit 1103 is used to generate visual-linguistic features by fusing visual and linguistic features through a cross-attention mechanism.
[0207] The answer generation unit 1104 is used to input visual language features into a large language model to generate predicted answers.
[0208] The first error determination unit 1105 is used to calculate the prediction error between the predicted answer and the true answer using a first loss function.
[0209] The first parameter adjustment unit 1106 is used to iteratively train the visual language model by adjusting the model parameters through the Adam optimizer based on the prediction error.
[0210] As shown in Figure 12, the open vocabulary detection device 700 also includes a second model training module 705. The second model training module 705 includes an object localization unit 1201, a classification unit 1202, a second error determination unit 1203, and a second parameter adjustment unit 1204.
[0211] The object localization unit 1201 is used to locate the target object on the input image through the target detection probe and obtain the predicted bounding box of the target object.
[0212] The classification unit 1202 is used to classify visual language features and predicted bounding boxes of target objects through a classifier, and generate predicted category labels for the target objects.
[0213] The second error determination unit 1203 is used to determine the classification error between the predicted class label and the true class label and the localization error between the predicted bounding box and the true bounding box through the second loss function.
[0214] The second parameter adjustment unit 1204 is used to adjust the model parameters based on classification error and localization error to iteratively train the open vocabulary detection model.
[0215] Figure 13 is a schematic diagram of the physical structure of an electronic device provided in an embodiment of this disclosure. As shown in Figure 13, the electronic device 130 includes a processor 1301, a memory 1302, and a bus 1303.
[0216] The processor 1301 and the memory 1302 communicate with each other via the bus 1303.
[0217] The processor 1301 is used to call program instructions in the memory 1302 to execute the methods provided in the above-described method embodiments.
[0218] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described open vocabulary detection method.
[0219] This disclosure also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described open vocabulary detection method.
[0220] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0221] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams.
[0222] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0223] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0224] The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of this disclosure. It should be understood that the above descriptions are merely specific embodiments of this disclosure and are not intended to limit the scope of protection of this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An open vocabulary detection method, characterized in that, include: Obtain the natural language prompts to be detected and the input image; A set of object names is obtained by jointly modeling the natural language prompts and input images using a pre-trained feature fusion model. The feature fusion model is obtained by performing feature fusion processing based on historical images and historical questions to obtain visual language features, and then iteratively training based on the visual language features. The target objects in the input image are detected by a pre-trained open vocabulary detection model based on the input image and the set of object names, and the detection results are obtained. The open vocabulary detection model is obtained by locating the historical image to obtain the predicted bounding box, and iteratively training based on the visual language features and the predicted bounding box.
2. The method according to claim 1, characterized in that, The method involves using a pre-trained feature fusion model to perform joint visual-language modeling on the natural language prompts and the input image, resulting in a set of object names, including: Multi-layer feature extraction is performed on the input image to obtain the visual features of the input image; The natural language prompts are parsed using a large language model to generate the linguistic features of the natural language prompts; The visual features and the language features are fused across modalities to obtain fused visual-language features; The visual language features are input into the large language model to generate the set of object names.
3. The method according to claim 2, characterized in that, The step of performing multi-layer feature extraction on the input image to obtain the visual features of the input image includes: The input image is cut into image blocks of a preset step size; The image feature sequence is generated by extracting features from the image blocks using a visual encoder.
4. The method according to claim 2, characterized in that, The step of fusing the visual features and the linguistic features across modalities to obtain fused visual-linguistic features includes: The visual features are compressed using a cross-attention mechanism to obtain compressed visual features; Generate encoded visual features based on compressed visual features and 2D absolute position encoding; The encoded visual features and language features are fused through a visual-language adapter to obtain the visual-language features.
5. The method according to claim 1, characterized in that, The detection results include the predicted bounding boxes of the objects and their corresponding predicted category labels; the detection of target objects in the input image using a pre-trained open vocabulary detection model based on the input image and the set of object names yields detection results, including: The input image is detected based on each object name in the set of object names, and a predicted bounding box and its corresponding predicted category label are generated for each target object.
6. The method according to claim 1, characterized in that, The steps for pre-training a visual language model include: The training dataset consists of multiple pairs of historical images, historical questions, and real answers. The historical images and historical questions are input into the visual language model for feature extraction to obtain visual features and language features. Visual-linguistic features are generated by fusing the visual features and the linguistic features through a cross-attention mechanism. The visual language features are input into a large language model to generate a predicted answer; The prediction error between the predicted answer and the true answer is calculated using a first loss function. Based on the prediction error, the model parameters are adjusted using the Adam optimizer to iteratively train the visual language model.
7. The method according to claim 2, characterized in that, The steps for pre-training the open vocabulary detection model include: The target object is located on the input image by the target detection probe, and the predicted bounding box of the target object is obtained. The visual language features and the predicted bounding boxes of the target objects are classified by a classifier to generate predicted category labels for the target objects. The classification error between the predicted class label and the true class label, and the localization error between the predicted bounding box and the true bounding box are determined by the second loss function. The open vocabulary detection model is iteratively trained by adjusting the model parameters based on the classification error and localization error.
8. An open vocabulary detection device, characterized in that, include: The data acquisition module is used to acquire the natural language prompts and input images to be detected; The feature processing module is used to perform visual language joint modeling on the natural language prompts and input images through a pre-trained feature fusion model to obtain a set of object names; the feature fusion model is obtained by performing feature fusion processing on historical images and historical questions to obtain visual language features, and then iteratively training based on the visual language features; The detection module is used to detect target objects in the input image based on the input image and the set of object names using a pre-trained open vocabulary detection model, and obtain detection results; The open vocabulary detection model is obtained by locating the historical image to obtain the predicted bounding box, and then iteratively training it based on the visual language features and the predicted bounding box.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.