Self-evolution target detection method guided by text information based on prototype matching
Through the prototype matching technology of multimodal pre-training model and cross-modal attention mechanism, the problem of insufficient self-evolution ability of text information-guided object detection in complex scenarios is solved, and high-performance object detection is achieved.
Patent Information
- Application Number
- CN202510767136.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing text information-guided object detection methods lack self-evolution and generalization capabilities in complex scenarios, and cannot effectively match visual and language features, resulting in a decrease in detection accuracy, and the model cannot cope with unseen targets in real scenarios.
The multimodal pre-trained model and the cross-modal attention mechanism are used to extract local target information of images and texts through prototype matching technology, and perform cross-modal feature interactive calculations to achieve alignment and matching of visual and text features. The CLIP model is used for pre-training to enhance the model's autoevolution ability.
It realizes high-performance detection of unseen targets in complex scenarios, improves the model's self-evolution ability and detection accuracy, solves the problem of interference between redundant information on detection, and ensures the alignment and matching of key information.
Smart Images

Figure CN120339595A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text information-guided object detection, and specifically proposes a text information object detection method that respectively extracts key information from visual and language modality inputs, introduces the capabilities of multi-modal pre-trained models, and performs prototype matching between modalities. The entire system uses the image-text encoder of the vision-language pre-trained model to encode the input pictures and text information, maps the information of different modalities into the same vector space, and performs prototype matching using the key information in the complex input, enabling the model to have an accurate object detection ability with self-evolution. Background Art
[0002] The task of text information-guided object detection is a key technology in the multi-modal field. The task is similar to the re-identification task in the multi-modal vision-language retrieval field, aiming to find the objects in a given picture that match the description according to the text information. Specifically, the input of the task is the text describing the shape and features of the object and a picture with a complex scene containing many objects. The model will perform object detection in the image according to the text information, find the target object that matches the text description, and output its position coordinates in the image. It is required to comprehensively understand complex language semantic information and visual scene information, and mine the relevance between text semantics and various target objects in various pictures. It can be applied to a very large number of real task scenarios, not only for current application scenarios such as monitoring and security, but also can be deployed in robots for civilian use (accompanying, medical) and military use (battlefield robots). Therefore, it has attracted a lot of research and exploration in the academic and industrial fields in recent years. However, most of the existing technologies at present can only perform well in object detection on the test set of the training data set. For example, given a picture containing different targets (airplanes, ships, people) that various models have not seen, the existing methods at this time cannot understand the text and accurately detect the targets. However, in real scenarios, it is often necessary to detect various targets in complex scenarios, which poses higher requirements for the knowledge generalization ability and self-evolution ability of the model.
[0003] Most of the current text information guided target detection methods and technologies use image and text encoder modules or unified encoder modules to encode the input images and texts to obtain feature vectors. After the feature fusion calculation of the self-designed neural network, the model directly outputs the location coordinates of the target. The detection results can be output quickly because only the forward matrix calculation of the neural network is performed. However, there are also major defects. This type of method does not fully match and align the visual and language features, resulting in reduced detection accuracy. In addition, these models are only trained on public data sets, and the visual text knowledge they obtain is very limited. The model's ability to evolve and generalize autonomously is poor, and it cannot cope with complex and changeable real scenes. For example, when there is no data on "rainbow clothes" in the training data, the model cannot search for the correct target. Both existing problems lead to poor performance of the current algorithm. Summary of the invention
[0004] In order to overcome the shortcomings of the prior art, the present invention provides a self-evolving target detection method guided by text information based on prototype matching, which can dynamically obtain key information in different images and text inputs, vectorize them using a multimodal pre-trained model, perform prototype matching, and explore the relationship between visual targets and text descriptions to achieve accurate text information target detection.
[0005] The main modules of the technical solution of the present invention are as follows: the system includes three parts, the first part is the process of continuing the local and target information discovery for the input image and text respectively, the second part is the multimodal feature extraction and association process based on the multimodal pre-training model and the cross-modal attention mechanism, and the third part is the prototype matching process of the information between the modalities. In the first part, the target detector and the semantic extractor are used to extract the local and target information in the image and text respectively, simplify the complex input, which is conducive to the subsequent prototype matching and target detection, and obtain various candidate targets in the input image and the target and attribute pronoun information in the input text respectively. In the second part, a multimodal large model and a cross-modal guided attention mechanism are used to extract the visual and text feature vectors and interactively calculate the cross-modal information relationship. In the third part, the feature vector prototypes of the two modalities are dynamically matched, and the final target is comprehensively selected based on the matching results.
[0006] The technical solution adopted by the present invention to solve the technical problem includes the following steps:
[0007] Step 1: An input image First, we pass a target detection model of a standard target detection task to obtain the target image set in the detected image, which is recorded as ,in represents the target image set, Indicates that the targets, and a total of candidate targets are detected; this step is to extract candidate target objects from the input image in advance, where the target detection model uses any type of detection model, including convolutional neural network models and Transformer series models;
[0008] Step 2: For an input text , first, use the Natural Language Toolkit (NLTK) to build a semantic extractor to extract target-related information from the text information, and then output the target-related information; specifically, the working process of the semantic extractor is as Figure 1 shown;
[0009] Step 3: Input the obtained target objects in the visual modality and the target-related word information in the language modality into the image encoder and the text encoder respectively to obtain the feature maps of each target image and the feature vectors of each target-related word;
[0010] Step 4: First, straighten the obtained target image feature maps, and send them together with the feature vectors of the target-related words into the cross-modal attention layer for cross-modal feature interaction; among them, the word features and the image features are respectively subjected to self-attention calculation, then cross-modal attention calculation, and finally calculated through a linear network and LayerNorm;
[0011] Step 5: After obtaining the target images and related word features for prototype matching, use each target-related word feature as a prototype because the target-related words contain all the semantic information for the target description; then, restore the straightened target image features in Step 4 to feature maps through dimensional changes of the feature vectors;
[0012] Step 6: Use average pooling calculation on the target-related word feature set to obtain the aggregated text feature representation , and the target image feature with the highest score in Step 5 is denoted as ;
[0013] Step 7: In the text information-guided target detection stage, input an image and a text. After Steps 1 - 5, sort according to the final scores of each target object to obtain the target objects predicted by the model that match the text.
[0014] In the above Step 2, first, perform text tokenization on the input text to obtain a word set; then, convert the words in the word set into word tokens, and use part-of-speech parsing to perform part-of-speech tagging on the obtained words; finally, find the nouns and adjectives among the words through a matching method, denoted as , where, represents the set of target-related words in the output, represents the th word, and there are a total of words.
[0015] In step 3, in order to better map the visual and language modalities to the same vector space for subsequent matching and alignment, the multi-modal pre-training model CLIP is used as the model structure of the image and text encoders and its parameters are used in the encoder stage; the CLIP model has been pre-trained on a large number of image-text pairs and has strong multi-modal matching ability. The formula is as follows:
[0016] (1)
[0017] (2)
[0018] where, represents the image encoder, represents the text encoder, represents the set of target image feature maps, is the feature map of the th target, and there are a total of ones, represents the set of target-related word features, is the feature of the th word, and there are a total of ones.
[0019] In step 4, in the self-attention layer, the target image features and the feature vectors of the target-related words are respectively used as the query, key, and value of the attention mechanism for calculation. The image features are used as the "query" and "key" of the attention mechanism, and the word features are used as the "value";
[0020] In the cross-modal attention calculation, first, the target image features are used as the query, and the feature vectors of the target-related words are used as the key and value for calculation. Then, the feature vectors of the target-related words are used as the query, and the target image features are used as the key and value for calculation; finally, after passing through the linear network and LayerNorm calculation, the target image and related word features for prototype matching are obtained. The formula is as follows:
[0021] (3)
[0022] where, and are respectively the query, key, and value in the attention mechanism, represents the linear network and LayerNorm calculation, represents the self-attention calculation, Denotes cross-modal attention calculation Denotes the target image feature map obtained through calculation; the image-related word features are also calculated through formula (3) to obtain the final target-related word features, i.e., the prototypes used subsequently .
[0023] In step 5, in the matching stage, for each prototype, step 5 calculates the similarity between the matching prototype and all target image features through cosine similarity; after the prototype matching is completed, a score is calculated for each target, and the score is the sum of the similarities between the target image features and the matching prototype; through score comparison, the target with the highest score is selected as the prediction of the model, and average pooling calculation is used to aggregate the selected target image feature map and target-related word features into one target image feature and one target-related text feature
[0024] (4)
[0025] Among them Denotes cosine similarity calculation Denotes the th feature vector in the th target image feature map, and there are vectors in total Denotes the final score of the
[0026] In step 6, in order to optimize the network, the widely used Mean Square Error Loss (MSE) is adopted for cross-modal alignment, aiming to minimize the difference between the obtained target image features and target-related text features. The specific formula is as follows
[0027] (5)
[0028] Among them is the dimension size of the feature vector Denotes the value of the th dimension of the target image feature with the highest score is the value of the th dimension of the aggregated text feature representation. The values of each dimension in the text modality and the image modality are calculated through variance and then summed and averaged to obtain the mean square error loss
[0029] An electronic device includes: one or more processors; a memory; one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method described above
[0030] A computer-readable storage medium stores program code, and the program code can be called by a processor to execute the method as described above.
[0031] The beneficial effect of the present invention is that by introducing and generalizing cross-modal knowledge in the vision-language pre-training model and performing prototype matching, the knowledge in the original pre-training model is applied to the text-guided object detection task, enabling the model to have the ability of self-evolution when facing objects not seen in the dataset. At the same time, using an object detector and a semantic extractor to perform localized preprocessing on the input image and text also solves the problem that there is a large amount of redundant information in the input text and image, which is not conducive to object detection, enabling the key object image information and text information related to the object to be fully aligned and matched, achieving high-performance object detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a process diagram of the semantic extractor of the present invention.
[0033] Figure 2 It is a structural diagram of an object detection model guided by text information.
[0034] Figure 3 It is an example diagram of the result of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The present invention will be further described below with reference to the drawings and embodiments.
[0036] The present invention provides a self-evolving object detection method guided by text information based on prototype matching. The whole method is as Figure 2 shown, and the specific process is as follows:
[0037] Step 1, preprocessing of the target image;
[0038] Input the candidate target image and calculate the size of the target image. If it is less than 10K, it is directly discarded and does not participate in the subsequent training. The images that meet the conditions are adjusted to using bilinear interpolation to meet the requirements of the object detector for the image size.
[0039] Step 2, image data augmentation;
[0040] Randomly perform data augmentation methods such as perspective transformation, central cropping, and horizontal flipping on the same target image. The X-axis rotation angle of the perspective transformation takes any value in, the Y-axis rotation angle takes any value in, and the Z-axis rotation angle takes any value in. After central cropping, the image is scaled to keep its size.
[0041] Step 3: The semantic extractor extracts the text information related to the target;
[0042] For the original text description "An Asian girl, wearing a pink shirt, is eating at the table", it is marked as , first introduce the Natural Language Toolkit (NLTK) to perform text tokenization on the input text to obtain a set of words. Then, convert these words into word tokens and use part-of-speech parsing to perform part-of-speech tagging on the obtained words. Finally, find the nouns and adjectives among these words through a matching method and output them.
[0043] Step 4: The target detector detects the target in the image;
[0044] Input the preprocessed and data-augmented image into the target detector for target detection. In this implementation, the YOLO-v7 target detection model is used as the target detector. YOLO-v7 is the latest released real-time target detector based on the convolutional neural network architecture. It has higher accuracy compared to the YOLO series, RCNN series, and Transformer series models, and its detection speed can reach real-time. Therefore, YOLO-v7 is used in the embodiments of the present invention to obtain several target object images in the input image.
[0045] Step 5: Modal data feature extraction;
[0046] After obtaining the target object images detected by the target detector and the target-related words obtained by the semantic extractor, pass the target object images through the image encoder of the vision-language pre-trained model to obtain several image feature maps. Input the target-related words into the text encoder of the vision-language pre-trained model to obtain the text feature vectors corresponding to the target-related words.
[0047] Step 6: Attention mechanism calculation;
[0048] After obtaining the two-modal features through the feature encoder, input them into the cross-modal attention layer for attention mechanism calculation. First, straighten the obtained target image feature maps. Then, perform self-attention calculations on the features of the two modalities respectively. The target image features and the target-related word features are used as the query, key, and value of their respective self-attention calculations. Then, perform cross-modal attention calculation. The target image features and the target-related word features are used as the query of the cross-modal attention calculation, and the other feature is used as the key and value. Finally, through a linear network and LayerNorm calculation, obtain the target-related word features and the restored target image feature maps.
[0049] Step 7: Cross-modal prototype matching;
[0050] After obtaining the target image feature map and target-related word features through cross-modal attention calculation, start matching with all candidate target images using the target-related word features as prototypes. The matching calculation uses cosine similarity. The specific operation is as follows: for a target image feature map, use all prototypes to perform cosine similarity matching with each feature in the feature map, and then sum up the obtained similarity scores to get the matching score of the target image and the input text. After all candidate target image feature maps are subjected to the above prototype matching calculation, sort according to the scores, and select the object with the highest score as the detection result of the method.
[0051] Step 8: Model training;
[0052] The entire training process is end-to-end training. The model is trained using mean squared error loss, the model optimizer uses Adam optimizer, and the batch size of the input data is set to 64. A total of 80 epochs are trained, and the overall learning rate is set to , and the learning rate drops by 10 times after 50 epochs. The warm-up learning rate strategy is adopted in the first 10 epochs.
[0053] Step 9: Model application;
[0054] After the above training process, multiple models can be obtained. Select the optimal model (with the smallest loss function value) for application. At this time, data augmentation is not required for image data processing here. Only need to adjust the image to size and normalize it to be used as the input of the model. The parameters of the entire network model are fixed. Just input the image data and text to be detected and perform forward propagation calculation. In actual experiments, some obtained final detection results are as Figure 3 shown. The texts of the four examples are "the first suitcase on the left", "a sandwich with red vegetables", "a girl sitting in a box with her sister", and "the man on the far left" respectively. The self-evolving object detection method guided by text information based on prototype matching can perform good object detection on examples with relatively complex and redundant information.
Claims
1. A self-evolving object detection method guided by text information based on prototype matching, characterized in that It includes the following steps: Step 1: An input image First, it passes through an object detection model for a standard object detection task to obtain a set of target images in the detected image, denoted as where represents the set of target images, represents the th target, and a total of candidate targets are detected; Step 2: For a piece of input text , first, use the Natural Language Toolkit (NLTK) to build a semantic extractor to extract the target-related information in the text information, and then output the target-related information; Step 3: Input the obtained target object in the visual modality and the target-related word information in the language modality into the image encoder and the text encoder respectively to obtain the feature maps of each target image and the feature vectors of each target-related word; Step 4: First, straighten the obtained target image feature maps, and then send them together with the feature vectors of the target-related words into the cross-modal attention layer for cross-modal feature interaction; wherein, the word features and the image features are respectively subjected to self-attention calculation, then cross-modal attention calculation, and finally calculated through a linear network and LayerNorm; Step 5: After obtaining the target image and related word features for prototype matching, use each target-related word feature as a prototype because the target-related words contain all semantic information for target description; then, restore the straightened target image features in Step 4 to feature maps through dimensional change of the feature vectors; Step 6: Use average pooling on the target-related word feature set to calculate the aggregated text feature representation , and the target image feature with the highest score in Step 5 is represented as ; Step 7: In the text information-guided target detection stage, input an image and a text. After Steps 1-Step 5, sort according to the final scores of each target object to obtain the target objects matching the text predicted by the model.
2. The text information-guided self-evolving target detection method based on prototype matching according to Claim 1, characterized in that: In step 2, first, perform text tokenization on the input text to obtain a set of words; then, convert the words in the set of words into word tokens and use part-of-speech parsing to tag the obtained words with part-of-speech; finally, find the nouns and adjectives among the words through a matching method, expressed as , where represents the set of target-related words output,[[]] represents the th word, and there are a total of words.
3. The text information-guided self-evolving target detection method based on prototype matching according to Claim 1, characterized in that: In Step 3, in order to better map the visual and language modalities to the same vector space for subsequent matching and alignment, use the multi-modal pre-trained model CLIP as the model structure of the image and text encoders and use its parameters in the encoder stage; the CLIP model has been pre-trained on a large number of image-text pairs and has strong multi-modal matching ability. The formula is as follows: ; ; Among them, represents an image encoder, represents a text encoder, represents a set of target image feature maps, is the feature map of the th target, with a total of ones, represents a set of target-related word features, is the feature of the th word, with a total of ones.
4. The text information-guided self-evolving target detection method based on prototype matching according to Claim 1, characterized in that: In Step 4, in the self-attention layer, the target image features and the feature vectors of the target-related words are respectively used as the query, key, and value of the attention mechanism for calculation. The image features are used as the "query" and "key" of the attention mechanism, and the word features are used as the "value"; In the cross-modal attention calculation, first use the target image features as the query, and the feature vectors of the target-related words as the key and value for calculation, and then use the feature vectors of the target-related words as the query, and the target image features as the key and value for calculation; finally, calculate through a linear network and LayerNorm to obtain the target image and related word features for prototype matching. The formula is as follows: ; Among them, and are the query, key, and value in the attention mechanism respectively, represents the linear network and LayerNorm calculations, represents the self-attention calculation, represents the cross-modal attention calculation, represents the target image feature map obtained through calculation; the image-related word features are also calculated by formula (3) to obtain the final target-related word features, that is, the prototypes used later .
5. The text information-guided self-evolving target detection method based on prototype matching according to Claim 1, characterized in that: In step 5, in the matching stage, for each prototype, step 5 calculates the similarity between the matching prototype and all target image features through cosine similarity; after the prototype matching is completed, a score is calculated for each target, and the score is the sum of the similarities between the target image features and the matching prototype; through score comparison, the target with the highest score is selected as the prediction of the model, and average pooling calculation is used to aggregate the selected target image feature map and the target-related word features into a target image feature and a target-related text feature; ; Among them, represents cosine similarity calculation, represents the th feature vector in the th target image feature map, and there are vectors in total. represents the final score of the th target image.
6. The text information-guided self-evolving object detection method based on prototype matching according to claim 1, wherein: In step 6, in order to optimize the network, the widely used mean square error loss is adopted for cross-modal alignment, aiming to minimize the difference between the obtained target image feature and the target-related text feature. The specific formula is as follows: ; Among them, is the dimension size of the feature vector, represents the value of the -th dimension of the target image feature with the highest score, is the value of the -th dimension of the aggregated text feature representation. The value of each dimension in the text modality and the image modality is calculated for variance and then summed and averaged to obtain the mean squared error loss.
7. An electronic device, characterized in that, Including: One or more processors; A memory; One or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1-6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code, and the program code can be called by the processor to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Text-target retrieval method based on dynamic self-evolution information extraction and alignment
CN116645694A
Multi-modal representation learning method based on text guide image block screening
CN117421591A
Information guide target searching method based on cross-modal self-evolution knowledge generalization
CN118170938A