Open set target detection method based on fusion text coding

By integrating text encoding technology in the object detection model, extracting and interacting images and text features, the problem of insufficient information on unknown categories of object detection is solved, and a higher precision object detection in open world scenarios is achieved.

CN120014621AActive Publication Date: 2025-05-16NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510279578.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-16
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

Traditional object detection models have difficulty providing detailed information when facing unknown categories of targets, resulting in the inability to effectively identify emerging unknown targets in open world scenarios.

Method used

The open-set object detection method based on fusion text encoding is adopted. By obtaining text template information, user text information and original images, an open-set object detection model is established, feature extraction and interaction are performed, and multimodal query features with text perception capabilities are output to achieve detailed detection of unknown categories of targets.

Benefits of technology

This method can provide rich detection information in open world scenarios, improve the accuracy and performance of object detection, and adapt to the detection needs of unknown categories of targets in the real world.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014621A_ABST
    Figure CN120014621A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, solves the technical problem that open set detection cannot be realized when a new unknown target appears in a detection image, and particularly relates to an open set target detection method based on fusion text information. Comprising the following steps: acquiring text template information preset according to a detection scene, user text information input by a user and corresponding to a detection requirement, and an original image; and establishing an open set target detection model for fusing the text mode for the target sample with the unknown category to realize open set detection. Rich detection information is provided for target detection for solving the problem of newly appearing unknown targets in detection images, finer-grained understanding of the images is achieved, the precision of target detection in the open world is improved, the limitation that target detection in a traditional method needs target limitation can be overcome, unknown category detection in the open world is achieved, and the detection efficiency is improved. And the method better adapts to real-world scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to an open set target detection method based on fusion text coding. Background Art

[0002] With the continuous improvement of industrial intelligence, the industrial production process has generated a large demand for target detection technology. Product defect detection, production workshop target positioning and tracking, etc. all use a large number of target detection technologies. Target detection technology is playing an increasingly important role in the industrial field. The research object of target detection is mainly two-dimensional images. The predicted bounding box of the target to be detected in the image is given and correctly classified to represent the detected object.

[0003] Traditional target detection models mainly focus on the recognition of objects of predefined known categories, which limits their usability in real-world applications. For example, in actual industrial scenarios, we often encounter some unknown, unlabeled defects, anomalies or new products. The emergence of these unknown categories has brought challenges to traditional closed-set detection. As a technology that can identify samples of unknown categories, open-set detection technology provides a new idea for solving this problem. Unlike traditional closed-set detection, open-set detection can not only accurately classify targets of known categories, but also provide information about samples that do not belong to any known category as much as possible. Some existing methods also try to detect targets of unknown categories, but can only provide simple information, such as only classifying them into unknown categories. Therefore, there is still a problem of being unable to recognize newly emerging unknown targets. Summary of the invention

[0004] The purpose of the present invention is to address the performance deficiencies of the above-mentioned prior art and propose an open set target detection method based on fusion of text information, which provides rich detection information for target detection to solve the problem of newly appearing unknown targets in the detection image, achieves a more fine-grained understanding of the image, and improves the accuracy of target detection in the open world.

[0005] To solve the above technical problems, the present invention provides the following technical solution: an open set target detection method based on fusion text encoding, the method comprising the following steps: S1, obtaining text template information preset according to the detection scenario, user text information corresponding to the detection requirements input by the user, and the original image; S2. establishing an open set target detection model for fusing text modalities for target samples with unknown categories to achieve open set detection; S3, based on the open set target detection model, feature extraction is performed on the text template information, user text information, and original image to obtain original text features and multi-scale image features; S4, interacting and enhancing the original text features and multi-scale image features to obtain multimodal query features with text perception and updated text features; S5. Decode the multimodal query feature to obtain the target area in the original image, including coordinate information and category information of the target sample; S6. Establish a contrastive focus classification loss for matching the target region with the updated text features and assigning labels ; S7, based on contrastive focus classification loss Establishing an overall loss function for updating the open-set object detection model , and set hyperparameters to obtain an open set object detection model for integrating text modality.

[0006] Furthermore, the open set object detection model includes a text feature extractor for extracting features from the text template information and the user text information to obtain original text features; A multi-scale image feature extractor is provided for extracting features from the original image to obtain multi-scale features, and a text perception module is provided for interactively outputting target sample coordinate information and category information based on the original text features and the multi-scale image features.

[0007] Furthermore, the text-aware module includes a text-aware encoder and a text-aware decoder with text feature perception capabilities. The text-aware encoder is used to realize the interaction between multi-scale image features and original text features, and output multimodal query features with text feature perception capabilities. The text-aware decoder is used to decode and output the coordinate information and category information of the target sample based on the input multimodal query features.

[0008] Furthermore, the method for constructing the text-aware encoder and text-aware decoder having text feature perception capability includes: The language-aware query selection module selects object features by evaluating the similarity between multi-scale image features and original text features, specifically: Computing multi-scale image features and original text features The similarity is then used to select the most relevant proposal embedding and target embedding , the most relevant proposal embedding Used to initialize the reference anchor point, the selected target embedding For subsequent language perception fusion, the expression is: ; In the formula, represents the Kronecker product; Indicates the first goals; represents transpose; The language-aware query fusion module fuses language features and object features while retaining the original semantics of the content query, and finally outputs a multimodal query feature with text-aware capabilities, expressed as: ; in, Indicates the layer number of the language-aware query fusion module; Represents the output of each layer in the language-aware query module; Represent the query, key, and value of the attention layer respectively; represents the attention layer, which can be divided into self-attention layer or cross-attention layer according to the input; It means that the feedforward layer consists of two layers of perceptrons and performs nonlinear transformation operations; is an activation function using a gating mechanism.

[0009] Furthermore, the contrast focus classification loss The expression is: ; In the formula, Output the probability that the target area belongs to the corresponding text feature for the open set object detection model; is the adjustment factor.

[0010] Furthermore, the overall loss function Include Loss and Loss, namely: ; Where Y is the real label data; X is the coordinate information output by the text-aware decoder; It is the intersection of the real box and the predicted box; is the minimum enclosed area between the real box and the predicted box; It is the area obtained by performing the operation on the real box and the predicted box; Overall loss function The expression is: ; In the formula, They represent the weight factors of the corresponding sub-loss functions respectively.

[0011] Furthermore, the training and hyperparameter settings of the open set object detection model are: The weight decay is The adaptive momentum estimation optimizer is used, the total batch size is 128, and the base learning rate of the text feature extractor, multi-scale image feature extractor, and text-aware decoder is The learning rate of the text-aware encoder is 0.1 times the base learning rate and trained for 24 epochs using a step learning rate schedule, where the learning rate is reduced to 0.1 and 0.01 times the base learning rate at the 16th and 22nd epochs, respectively.

[0012] By means of the above technical solution, the present invention provides an open set target detection method based on fusion text encoding, which has at least the following beneficial effects: 1. By combining language modality, the present invention can overcome the limitation of traditional target detection methods that require limited targets, realize open-world unknown category detection, and better adapt to real-world scenarios. This provides rich detection information for target detection to solve the problem of unknown targets newly appearing in the detection image, realizes a more fine-grained understanding of the image, and improves the accuracy of target detection in the open world.

[0013] 2. This invention aims to improve the performance of open set detection in images by integrating text modalities to provide richer information for open set detection, so that each region is divided into new categories in the perceptual semantic space, and the alignment of image targets and text information is achieved. This further improves the performance of open set target detection and provides richer detection information for industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 is a flow chart of the open set target detection method in the present invention; Figure 2 It is a network structure diagram of the open set target detection model in the present invention; Figure 3 This is a network structure diagram of the text perception module in the present invention. DETAILED DESCRIPTION

[0015] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods, so that the implementation process of how the present application uses technical means to solve technical problems and achieve technical effects can be fully understood and implemented accordingly.

[0016] This embodiment uses the current pre-trained text feature extractor to obtain text template information and original text features in user text information, which can inject rich semantic information into the open set object detection model; at the same time, it uses the existing pre-trained image feature extractor to obtain multi-scale image features of the original image.

[0017] On this basis, a text-aware encoder with text perception is constructed to realize the interaction between original text features and multi-scale image features and perform feature updates, output multimodal query features with text perception and updated text features, and then send them together to a text-aware decoder with text perception to decode them to obtain open set detection targets and obtain the target area in the original image; by establishing contrast focus classification loss Match and align the target area and text features to provide more fine-grained label information for the detected target. Finally, the open set target detection model is updated together with other losses to complete the training and obtain the final open set target detection model that integrates the text modality.

[0018] Please refer to Figure 1-Figure 3 This embodiment proposes an open set target detection method based on fusion text encoding. By combining language modality, it can overcome the limitation of traditional target detection methods that require limited targets, realize open world unknown category detection, and better adapt to real world scenarios. Figure 1 As shown, the implementation process of this method is divided into the following steps: S1. Obtain the text template information preset according to the detection scenario, the user text information corresponding to the detection requirements input by the user, and the original image. Among them, the preset text template information is a further supplement and enrichment of the semantics of the deployment scenario, which is generally an annotation of the scene image and can be obtained from a predefined knowledge base. The user text information is the text information input by the user according to the detection requirements, or the text information further supplemented.

[0019] S2. Establish an open set target detection model for fusing text modalities for target samples with unknown categories to achieve open set detection. Figure 2 As shown, the open set target detection model includes a text feature extractor for extracting features from text template information and user text information to obtain original text features, and a multiscale image feature extractor for extracting features from original images to obtain multiscale features, and a text perception module for interactively outputting target sample coordinate information and category information based on original text features and multiscale image features. In this embodiment, the text feature extractor uses BERT or other types of text feature extractors. The multiscale image feature extractor uses ResNet or other types of image multiscale feature extractors.

[0020] The text-aware module includes a text-aware encoder and a text-aware decoder with text feature perception capabilities. The text-aware encoder is used to realize the interaction between multi-scale image features and original text features, and output multimodal query features with text feature perception capabilities. The text-aware decoder is used to decode and output the coordinate information and category information of the target sample based on the input multimodal query features.

[0021] like Figure 3 As shown, the method for constructing a text-aware encoder and a text-aware decoder with text feature perception capability includes: ; Computing multi-scale image features and original text features The similarity is then used to select the most relevant proposal embedding and target embedding , the most relevant proposal embedding Used to initialize the reference anchor point, the selected target embedding Used for subsequent language perception fusion. represents the Kronecker product; Indicates the first A goal.

[0022] The language-aware query fusion module fuses language features and object features while retaining the original semantics of the content query. This module is also an important part of the text-aware decoder layer. Each layer contains self-attention and cross-attention operation sublayers, and finally outputs multimodal query features with text-aware capabilities, expressed as: ; in, Indicates the layer number of the language-aware query fusion module; Represents the output of each layer in the language-aware query module; Represent the query, key, and value of the attention layer respectively; Represents the attention layer. It can be divided into self-attention layer or cross-attention layer according to the input. There are three types of inputs, corresponding to query, key, and value. Each attention layer has different settings according to its position. , the q used is usually the input of the previous layer , k, v are usually the input of the previous layer , multi-scale image features And target embedding ; It means that the feedforward layer consists of two layers of perceptrons and performs nonlinear transformation operations; It is an activation function using a gate mechanism. The specific calculation formula is as follows: ; in, Represents input; are the model parameters that need to be learned; and Represents the activation function operation.

[0023] This embodiment constructs a text-aware encoder with text-aware capabilities, which can realize the interaction and update of multi-scale image features and original text features, and obtain multimodal query features. It uses a text-aware decoder to obtain the location and classification of targets in the image, thereby improving the open set detection capability of the model.

[0024] S3. Based on the open set target detection model, feature extraction is performed on the text template information, user text information and original image to obtain original text features and multi-scale image features. In this embodiment, the acquisition of original text features and multi-scale image features is completed by the text feature extractor and the multi-scale image feature extractor in the open set target detection model. Among them, there are many types of models that support text feature extraction, and BERT type models are preferred to provide good original text features. There are many models that support multi-scale feature extraction of images, such as ResNet native type and related variant models, or Swin-transformer type models, which can provide good multi-scale image features.

[0025] This embodiment uses a text feature extractor and a multi-scale image feature extractor to extract image and text features respectively, providing a good feature basis for the open set target detection model and injecting rich semantic information, so that the present invention has a good feature extraction function and ensures the open set detection accuracy of the model.

[0026] S4. The original text features and the multi-scale image features are interacted and feature enhanced to obtain multi-modal query features with text perception and updated text features. This embodiment constructs a text perception module with text feature perception, which includes a text perception encoder and a text perception decoder with text feature perception. By feeding the multi-scale image features and the original text features into the text perception encoder, the interaction and feature enhancement between the multi-scale image features and the original text features are realized, and the multi-modal query features with text perception and updated text features are output.

[0027] S5. Decode the multimodal query feature to obtain the target area in the original image, including the coordinate information and category information of the target sample; input the multimodal query feature with text perception capability into the text-aware decoder to decode and output the coordinate information and category information of the target sample.

[0028] S6. Establish a contrastive focus classification loss for matching the target region with the updated text features and assigning labels ; In this embodiment, the contrast focus classification loss The expression is: ; In the formula, Output the probability that the target area belongs to the corresponding text feature for the open set object detection model; It is a regulating factor used to dynamically reduce the weight of easily distinguishable samples during training and focus on difficult samples.

[0029] S7. Establish an overall loss function for updating the open set object detection model based on contrastive focus classification loss , and set hyperparameters to obtain an open set object detection model for integrating text modality. This embodiment uses the overall loss function Perform gradient calculation and update to complete the training of the open set target detection model, where the overall loss function Include Loss and loss.

[0030] Specifically, this embodiment calculates the loss based on the coordinate information of the target sample output by the text-aware decoder and the real coordinates and loss ,Right now: ; Where Y is the real label data; X is the coordinate information output by the text-aware decoder; It is the intersection of the real box and the predicted box; is the minimum enclosed area between the real box and the predicted box; It is the area obtained by performing the operation on the real box and the predicted box; Therefore, the overall loss function The expression is: ; In the formula, They represent the weight factors of the corresponding sub-loss functions respectively.

[0031] This embodiment establishes multiple losses for training, including the positioning loss Loss, Classification Loss loss, and contrastive focus classification loss for image-text alignment To optimize the open set target detection model, thereby enhancing the text perception ability of the open set target detection model and enriching the text semantics, so that the present invention can provide richer target information for target samples of unknown categories in the real world, thereby breaking the predefined closed set and realizing open set detection.

[0032] The training of the open set object detection model and the related hyperparameter settings are as follows: In order to maintain the simplicity of the open set target detection model, this embodiment adopts the weight decay The adaptive momentum estimation optimizer is used, the total batch size is 128, and the base learning rate of the text feature extractor, multi-scale image feature extractor, and text-aware decoder is , the learning rate of the text-aware encoder is 0.1 times the base learning rate, which is specifically set to , trained for 24 epochs, using a step learning rate schedule, where the learning rate was reduced to 0.1 and 0.01 times the base learning rate at the 16th and 22nd epochs, respectively. After the open set object detection model training is completed, the training model weights are saved and loaded, and then the image is input to obtain the open set detection results.

[0033] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, so the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0034] The above implementation methods have been described in detail. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. An open set object detection method based on fusion text encoding, characterized in that: The method comprises the following steps: S1, obtaining text template information preset according to the detection scenario, user text information corresponding to the detection requirements input by the user, and the original image; S2. establishing an open set target detection model for fusing text modalities for target samples with unknown categories to achieve open set detection; S3, based on the open set target detection model, feature extraction is performed on the text template information, user text information, and original image to obtain original text features and multi-scale image features; S4, interacting and enhancing the original text features and multi-scale image features to obtain multimodal query features with text perception and updated text features; S5. Decode the multimodal query feature to obtain the target area in the original image, including coordinate information and category information of the target sample; S6. Establish a contrastive focus classification loss for matching the target region with the updated text features and assigning labels ; S7, Contrastive Focus Classification Loss Establishing an overall loss function for updating the open-set object detection model , and set hyperparameters to obtain an open set object detection model for integrating text modality.

2. The open set target detection method according to claim 1, characterized in that: The open set object detection model includes a text feature extractor for extracting features from text template information and user text information to obtain original text features; A multi-scale image feature extractor is provided for extracting features from the original image to obtain multi-scale features, and a text perception module is provided for interactively outputting target sample coordinate information and category information based on the original text features and the multi-scale image features.

3. The open set target detection method according to claim 2, characterized in that: The text perception module includes a text perception encoder and a text perception decoder with text feature perception capabilities. The text perception encoder is used to realize the interaction between multi-scale image features and original text features, and output multimodal query features with text feature perception capabilities. The text perception decoder is used to decode and output the coordinate information and category information of the target sample based on the input multimodal query features.

4. The open set target detection method according to claim 3, characterized in that: The method for constructing the text-aware encoder and text-aware decoder having text feature perception capability includes: The language-aware query selection module selects object features by evaluating the similarity between multi-scale image features and original text features, specifically: Computing multi-scale image features and original text features The similarity is then used to select the most relevant proposal embedding and target embedding , the most relevant proposal embedding Used to initialize the reference anchor point, the selected target embedding For subsequent language perception fusion, the expression is: ; In the formula, represents the Kronecker product; Indicates the first goals; represents transpose; The language-aware query fusion module fuses language features and object features while retaining the original semantics of the content query, and finally outputs a multimodal query feature with text-aware capabilities, expressed as: ; in, Indicates the layer number of the language-aware query fusion module; Represents the output of each layer in the language-aware query module; Represent the query, key, and value of the attention layer respectively; represents the attention layer, which can be divided into self-attention layer or cross-attention layer according to the input; It means that the feedforward layer consists of two layers of perceptrons and performs nonlinear transformation operations; is an activation function using a gate mechanism.

5. The open set target detection method according to claim 1, characterized in that: The contrastive focal classification loss The expression is: ; In the formula, Output the probability that the target area belongs to the corresponding text feature for the open set object detection model; is the adjustment factor.

6. The open set target detection method according to claim 1, characterized in that: The overall loss function Include Loss and Loss, namely: ; Where Y is the real label data; X is the coordinate information output by the text-aware decoder; It is the intersection of the real box and the predicted box; is the minimum enclosed area between the real box and the predicted box; It is the area obtained by performing the operation on the real box and the predicted box; Overall loss function The expression is: ; In the formula, They represent the weight factors of the corresponding sub-loss functions respectively.

7. The open set target detection method according to claim 1, characterized in that: The training and hyperparameter settings of the open set object detection model are: The weight decay is The adaptive momentum estimation optimizer is used, the total batch size is 128, and the base learning rate of the text feature extractor, multi-scale image feature extractor, and text-aware decoder is The learning rate of the text-aware encoder is 0.1 times the base learning rate and trained for 24 epochs using a step learning rate schedule, where the learning rate is reduced to 0.1 and 0.01 times the base learning rate at the 16th and 22nd epochs, respectively.

Citation Information

Patent Citations

  • Lung sound recognition method based on variational auto-encoder and feature reconstruction

    CN118016107A

  • Target detection method and device, electronic equipment and program product

    CN118823316A

  • Task agnostic open set prototype for small sample open set identification

    CN119213445A