Finger understanding method and system based on pixel-vocabulary association modeling
By embedding a feature enhancement module based on pixel-word association in the convolution encoder, the correlation between visual pixels and text vocabulary is solved, and the problem of lack of specific visual features for text in the prior art is significantly improved.
Patent Information
- Application Number
- CN202510594784.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The prior art lacks specific visual features for text in the reference comprehension task and fails to clearly align visual and linguistic features of local pixels and text words, making it difficult to highlight image areas related to specific semantics during recognition.
A reference representation understanding method based on pixel-vocabulary correlation modeling is proposed. By embedding a feature enhancement module based on pixel-vocabulary correlation in a convolutional encoder, the correlation between visual pixels and text vocabulary is calculated, the text features are embedded in pixel features for enhancement, and the features are integrated through a multi-stage cascade decoder and a cross-attention mechanism to output the final detection result.
It significantly improves the accuracy of object detection, can better focus on the image areas related to language, and improves the language guidance effect during visual feature extraction.
Smart Images

Figure CN120107570A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a method and system for understanding referential expressions based on pixel-vocabulary association modeling. Background Art
[0002] As an important research direction in the field of human-computer interaction, referential expression understanding has received widespread attention in recent years. With the continuous development of multimodal reasoning technology, the practical value of this task has become increasingly prominent. The core of referential expression understanding lies in the effective combination of natural language description and visual information, and locating the target object in the image by analyzing text instructions. This not only involves the coordinated processing of language and visual information, but also plays an important role in downstream tasks such as visual question answering, image description, and visual language navigation.
[0003] In the prior art, the referential expression understanding task is usually handled based on the target detection framework. The two-stage method relies on pre-generated region candidate boxes or anchor boxes, but it often seems powerless when faced with complex text or flexible targets. In addition, the one-stage method extracts features through independent visual and text encoders, and then fuses and predicts them in the multimodal decoder. Although these methods improve processing efficiency, due to the relative independence of visual and text features, they may face overfitting problems when locating multiple objects with different text descriptions in the same image, resulting in unsatisfactory detection performance. With the deepening of research on the referential expression understanding task, the academic community has gradually realized the limitations of relying solely on independent branches to process visual and text features.
[0004] For referential expression understanding tasks, existing methods often adopt a two-branch architecture, that is, using independent visual encoders and text encoders to extract features of specific modalities. Therefore, the visual encoder relies on prior knowledge to extract potential foreground features during training. However, since the same image is usually associated with multiple objects related to unique text expressions, the independent visual feature extraction process limits the capabilities of the visual encoder. It can only train broad foreground features, but lacks specific visual features for text. Therefore, it is difficult for the visual encoder to focus on the areas most relevant to the text description, which leads to significant differences between the extracted visual features and the visual features required in multimodal reasoning. Therefore, the independent visual encoder cannot fully meet the requirements of referential expression understanding tasks.
[0005] Another category addresses the challenges through multiple language-guided visual encoder structures. These methods can be roughly divided into two categories: one is feature-based methods, which directly process intermediate visual features based on language information, or incorporate language information into visual representations through cross-attention mechanisms; the other is structure-based methods, which dynamically modify the parameters or architecture of the visual encoder based on language features to extract language-related visual features. Although these methods have achieved significant performance improvements in visual feature enhancement, most methods still require complex designs, such as query-aware attention modules, dynamic weight generators, etc., which increases the difficulty of model training. In addition, the above methods fail to explicitly align the visual and linguistic features of local pixels and text words, which may lead to the neglect of descriptive vocabulary for specific target objects, thereby indiscriminately highlighting other image areas related to similar semantic concepts. Summary of the invention
[0006] The present invention proposes a referential expression understanding method and system based on pixel-lexical association modeling to solve the technical problem that the prior art lacks specific visual features for text or fails to clearly align the visual and language features of local pixels with text words, resulting in difficulty in highlighting image areas related to specific semantics during recognition.
[0007] In order to solve the above technical problems, the present invention provides a method for understanding referential expressions based on pixel-vocabulary association modeling, comprising the following steps: Step S1: Obtain text description and corresponding original image; Step S2: extracting text features of the text description; extracting pixel features of the original image through a convolutional encoder, and embedding a feature enhancement module based on pixel-word association after several convolutional layers of the convolutional encoder; the feature enhancement module based on pixel-word association calculates the correlation between visual pixels and text vocabulary, embeds the text features into the pixel features to enhance the pixel features; Step S3: The enhanced visual features and the original text features are input into a multi-stage cascade decoder, the features are integrated through a cross-attention mechanism, and a feed-forward neural network is used to output the final detection results.
[0008] Preferably, in step S2, the method for calculating the correlation between visual pixels and text words by the feature enhancement module based on pixel-word association includes: mapping pixel features and text features to a unified dimensional space for alignment through linear mapping, and calculating the correlation at the pixel-word level; generating word-level attention weights for different words in the input expression; and operating based on attention pooling to convert the pixel-word level correlation into the pixel-sentence level correlation through the word-level attention weights.
[0009] Preferably, the vocabulary-level attention weight The expression is: ; In the formula, represents the nonlinear projection function; Represents the first i vocabulary; N l Indicates the number of words.
[0010] Preferably, the expression of the pixel-sentence level correlation is: ; In the formula, it represents the The association of words to pixels; is the sigmoid function.
[0011] Preferably, in step S2, a ResNet network is used to extract pixel features.
[0012] Preferably, a feature enhancement module is embedded after the third convolutional layer, the fourth convolutional layer and the fifth convolutional layer of the ResNet network.
[0013] Preferably, a Transformer encoder is stacked after the ResNet network.
[0014] Preferably, a feature enhancement module is embedded before the linear projection of the Transformer encoder.
[0015] The present invention also provides a referential expression understanding system based on pixel-vocabulary association modeling, which is applicable to the above-mentioned referential expression understanding method based on pixel-vocabulary association modeling, and includes a text feature extraction module, a pixel feature extraction module, a pixel-word association enhancement module and a decoding module; The text feature extraction module is used to extract text features; The pixel feature extraction module is used to extract pixel features; The pixel-word association enhancement module is embedded after several convolutional layers in the pixel feature extraction module, and is used to obtain the correlation between pixels and text words, and embed the text features into the pixel features for enhancement based on the correlation; The decoding module is used to integrate the enhanced visual features and the original text features through a cross-attention mechanism, and output the final detection results using a forward neural network.
[0016] Preferably, the cross-layer regularization loss For each pixel-word association enhancement module generated Apply distribution constraints for training: the true labels Convert to a binary mask of the same size as the input image , where the bounding box , that is, the pixel values in the true label are set to 1, and the remaining pixel values are set to 0; adjust the mask The size of , so that it corresponds to different pixel-word association enhancement modules Keep the same spatial resolution; finally, perform spatial normalization to obtain the flattened and , mapped into a probability distribution through the Softmax function and Then, define the loss for and The KL divergence between is calculated as: ; ; ; In the formula, and Respectively and The elements, is the temperature parameter that controls the distribution shape in the Softmax function, represents the total number of feature enhancement modules, It is the product of width and height in the spatial resolution of the feature map.
[0017] The beneficial effects of the present invention include at least: the present invention proposes a feature enhancement module based on pixel-word association, which improves the language guidance effect in the visual feature extraction process by explicitly calculating the pixel-word correlation between visual features and language features, and can better focus on image areas related to language, thereby significantly improving the accuracy of target detection.
[0018] The present invention optimizes the text parameter update design of the convolution module by integrating the feature enhancement module into the existing pixel feature extraction network, namely the convolution encoder. This integration method not only reduces the number of parameters, but also avoids the complexity of designing additional modules for different architectures. It can effectively focus on the text reference area, thereby significantly improving the detection effect of the target object. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a schematic diagram of the method flow of Example 1 of the present invention; Figure 2 This is a schematic diagram of the structure of a feature enhancement module according to Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the structure of a multi-stage cascade decoder according to Embodiment 1 of the present invention; Figure 4 Schematic diagram of the detection effect of the method of Example 1 of the present invention and the existing method; Figure 5 This is a schematic diagram of the improved structure of the ResNet network of Example 2 of the present invention; Figure 6 This is a schematic diagram of the improved structure of the Transformer network of Example 3 of the present invention; Figure 7 This is a schematic diagram of the overall network structure of Example 3 of the present invention. DETAILED DESCRIPTION
[0020] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.
[0021] Example 1 like Figure 1 As shown, an embodiment of the present invention provides a method for understanding referential expressions based on pixel-vocabulary association modeling, comprising the following steps: Step S1: Get the text description and the corresponding original image.
[0022] Specifically, in the embodiment of the present invention, a text description and its corresponding original image are obtained, and a pre-processed text-image pair is formed through data enhancement or data standardization.
[0023] First, basic data augmentation or normalization is performed to generate preprocessed image and text pairs.
[0024] The image is a 3-channel color image obtained by a common image acquisition device. The expression text is used to describe the object to be detected, including the object's attributes, location, and category. The target box is used to indicate the coordinates of the detected object in the image, expressed as ,in represents the coordinates of the center of the object, and Represents the width and height of the object respectively.
[0025] In terms of image data enhancement, commonly used techniques include random scaling, cropping, and flipping. The range of random scaling can fluctuate based on 640×640, and the size can be 480, 576, and 640, etc. It is recommended that the probability of random cropping be controlled below 0.5. The processing in this embodiment mainly involves setting special word vectors at the beginning and end of the text for marking, as well as cropping and completing the text to ensure that the text sequence has a fixed length.
[0026] Step S2: extracting text features of the text description; extracting pixel features of the original image through a convolutional encoder, and embedding a feature enhancement module based on pixel-word association after several convolutional layers of the convolutional encoder; the feature enhancement module based on pixel-word association calculates the correlation between visual pixels and text vocabulary, and embeds text features into pixel features to enhance the pixel features.
[0027] In this embodiment, the input text is encoded into discrete word vectors through a pre-trained word segmenter and input into a text encoder to extract text features.
[0028] Specifically, use the word segmentation tool to segment the text and convert it into word vectors for subsequent calculation and analysis; add specified marker word vectors at the beginning and end of the text to mark the start and end of the text expression; truncate the text and fill it with specified word vectors to ensure that all text input lengths remain consistent, so as to achieve standardization of text expressions of different lengths.
[0029] The text encoder usually uses a pre-trained text encoding model, such as Bert or CLIP. Such models are usually composed of several layers of Transformer encoding layers and are pre-trained using self-supervision or masked data modeling to extract and integrate the global information of the text. In order to achieve subsequent multimodal fusion and input it into the decoder for prediction, a linear layer is usually used to reduce the dimension of Bert's output so that the text features match the dimensions of the visual features. I will not go into details here.
[0030] The processed image is input to a convolutional encoder and visual features are extracted layer by layer. The convolutional encoder is usually a convolutional neural network. Common convolutional encoders include Resnet and ConvNext, etc., which are not limited in this embodiment. These networks effectively extract visual features in images through a multi-layer structure.
[0031] In the process of extracting visual features, in order to accurately locate the target object referred to by the natural language description in the image, the extracted visual features should focus on the local area related to the text description. To this end, this embodiment provides a feature enhancement module based on pixel-word association, the structure of which is as follows: Figure 2As shown, by setting the feature enhancement module after any convolution layer of the convolution encoder, the pixel features extracted by the convolution layer are processed, and the correlation map between the visual features and the text features is calculated to enhance the visual features. The specific operations are as follows.
[0032] Given the flattened visual features and the text features of the input sentence ,in Indicates the height of the image. Indicates width, The number of channels representing visual features, Indicates the length of the sentence, The number of channels representing text features. First, the visual features and language features are mapped to a unified dimensional space through linear mapping to align and calculate the pixel-word level association. The specific form is: ; in, and are the linear projection matrices of visual and language features respectively, is the unified dimension after projection, U It represents the pixel-word level association between visual features and text features. It is shared in the feature enhancement modules based on pixel-word association in different layers to ensure that the text features are projected into the same feature space.
[0033] Use attention-based pooling to generate weights to obtain text for visual features Pixel-level correlation information.
[0034] Considering that different words in a text representation have different contributions to the target they refer to, this embodiment generates word-level attention weights in the following way for different words in the input expression: , to express the importance of each word: ; in, represents a nonlinear projection function that transforms language features The The vocabulary is mapped to the corresponding attention weight . This function consists of two linear layers and a Non-linear activation layer.
[0035] Through pixel-to-word level association and attention weights , the calculation process of calculating the pixel-sentence level association information is as follows: ; Here, express No. i List, is the sigmoid function, which is responsible for normalizing each element value to Range. Calculation results Aggregates the correlation between each pixel and all words. is considered as a spatial weight mask and applied to to enhance features.
[0036] The final feature calibration is expressed as: ; in represents a broadcast element-wise multiplication operation, so that the calibrated features More attention is paid to language-related pixels, thereby improving the effectiveness of the feature enhancement module in visual localization tasks.
[0037] Step S3: The enhanced visual features and the original text features are input into a multi-stage cascade decoder, the features are integrated through a cross-attention mechanism, and a feed-forward neural network is used to output the final detection result, i.e., the predicted target box coordinates.
[0038] Specifically, this embodiment adopts a multi-stage cascade decoder, whose structure is as follows: Figure 3 As shown in Figure 1, the decoder uses a learnable token and a cross-attention mechanism to iteratively aggregate and extract useful information from visual and text features. The design of the cascaded decoder removes the computational overhead required for further feature enhancement, thereby improving the efficiency of the overall framework.
[0039] The cross-attention mechanism specifically uses the learnable token as the query and cross-attention with the text and visual features to aggregate the target-related features, and the forward neural network used to output the results refers to a fully connected network with several layers stacked. The number of channels in the hidden layer can be set as needed, usually set to 256, and the number of output channels in the output layer is 4, corresponding to the relative position of the final target coordinates in the image. .
[0040] like Figure 4 As shown, it is a schematic diagram comparing the detection results of the method of this embodiment and the existing method, wherein the red box is the true label and the green box is the prediction result. It can be seen that compared with the existing method, the method of this embodiment has achieved better prediction results in different scenarios. Even if different methods have located the position of the object, the method of this embodiment still achieves a more accurate detection effect, such as the prediction of barbecue in the second row in the figure.
[0041] Example 2 In this embodiment, based on the embodiment 1, the convolution encoder in step S2 is replaced by a ResNet network, and after further selecting the third convolution layer, the fourth convolution layer and the fifth convolution layer of the ResNet network, a feature enhancement module is added, and its structure is as follows: Figure 5 As shown. Specifically expressed as: ; Afterwards, the enhanced features will be adjusted to The size of the feature map is input to the next layer. , and The value of may vary.
[0042] This operation is designed to highlight language-related visual information and ensure that the enhanced features match the dimensions of subsequent layers.
[0043] Example 3 This embodiment is based on Embodiment 2, and a Transformer encoder is stacked after adding a ResNet network, and a feature enhancement module is added before the linear projection of the query and key of the Transformer encoder. The structure is as follows: Figure 6 As shown, the calculation process is as follows: ; ; in, represents a multi-head self-attention layer, It is the abbreviation of query vector query, key vector key and value vector value in the attention layer input, and the enhanced output It will be passed to the subsequent normalization layer and feed-forward network.
[0044] The area associated with the text is enhanced by adjusting the attention score. This design ensures that the enhanced output can more effectively reflect the relationship between the image and the text, and pass it to the subsequent layer normalization and feedforward network. The integration of feature enhancement modules based on pixel-word association not only improves the feature expression ability of each layer, but also enhances the model's ability to understand and process image and text information.
[0045] The overall network structure after integrating Examples 1 to 3 is as follows Figure 7 It is shown but not intended to be a limitation of the present invention.
[0046] Example 4 The present invention also provides a referential expression understanding system based on pixel-vocabulary association modeling, which is applicable to the above-mentioned referential expression understanding method based on pixel-vocabulary association modeling, and includes a text feature extraction module, a pixel feature extraction module, a pixel-word association enhancement module and a decoding module; A text feature extraction module, used to extract text features; A pixel feature extraction module, used to extract pixel features; The pixel-word association enhancement module is embedded after several convolutional layers in the pixel feature extraction module to obtain the correlation between pixels and text words, and embeds text features into pixel features for enhancement based on the correlation; The decoding module is used to integrate the enhanced visual features and original text features through the cross-attention mechanism, and use the forward neural network to output the final detection results.
[0047] This embodiment trains the system by minimizing the gap between the predicted box and the labeled box. At this stage, the model is trained by minimizing the IoU loss and L1 loss between the predicted box and the labeled box, and an additional cross-layer consistency loss is used. To constrain the pixel-word association distribution of the feature enhancement module.
[0048] First, the coordinate box constraint aims to directly compare the difference between the predicted bounding box and the true bounding box. This process is achieved by calculating the difference between the coordinate position, width and height of the model output and the true annotation, thereby ensuring that the model can accurately locate the target object. Second, cross-layer consistency aims to ensure that the visual feature matrices generated by different decoder layers are consistent. Maintain consistency when processing the same input. Since the visual features extracted by different encoder layers There are significant differences in receptive field and information extraction degree. The returned matrix will also be different. To solve this problem, a loss function is introduced for cross-layer consistency , by imposing distribution constraints on the matrix returned by each layer to ensure consistency when processing similar spatial regions. Combining these two constraints, the model can not only improve the accuracy of predictions, but also maintain consistent attention to the target area at different levels, thereby achieving better image understanding and object detection capabilities.
[0049] Specifically, the true label Convert to a binary mask of the same size as the input image , where the true label Set the pixel values inside to 1 and the rest to 0; adjust the mask The size of , so that it corresponds to different pixel-word association enhancement modules Keep the same spatial resolution; finally, perform spatial normalization to obtain the flattened and , mapped into a probability distribution through the Softmax function and Then, define the loss for and The KL divergence between is calculated as: ; ; ; in, and Represents vectors and The elements, It is the temperature parameter that controls the distribution shape in the Softmax function and is generally set to 0.2. represents the total number of feature enhancement modules based on pixel-word association, It is the product of width and height in the spatial resolution of the feature map. The process will minimize and The distribution difference between The large values are concentrated in designated target area. This process ensures The area cited by the text can be effectively enhanced. At the same time, through the above operation, the true value of the annotation The information in can also be obtained through binary masking Efficiently passed to shallow layers of the visual encoder.
[0050] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. Only the preferred embodiments of the present invention are expressed. The description is more specific and detailed, but it cannot be understood as limiting the scope of the present invention. As long as there is no contradiction in the combination of these technical features, they should be considered as within the scope of this specification.
[0051] It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these modifications and improvements all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.
Claims
1. A method for understanding referential expressions based on pixel-lexical association modeling, characterized in that: The following steps are involved: Step S1: Obtain text description and corresponding original image; Step S2: extracting text features of the text description; extracting pixel features of the original image through a convolutional encoder, and embedding a feature enhancement module based on pixel-word association after several convolutional layers of the convolutional encoder; the feature enhancement module based on pixel-word association calculates the correlation between visual pixels and text vocabulary, embeds the text features into the pixel features to enhance the pixel features; Step S3: The enhanced visual features and the original text features are input into a multi-stage cascade decoder, the features are integrated through a cross-attention mechanism, and a feed-forward neural network is used to output the final detection results.
2. The method for understanding referential expressions based on pixel-vocabulary association modeling according to claim 1, characterized in that: In step S2, the method for calculating the correlation between visual pixels and text words by the feature enhancement module based on pixel-word association includes: mapping pixel features and text features to a unified dimensional space for alignment through linear mapping, and calculating the correlation at the pixel-word level; generating word-level attention weights for different words in the input expression; and operating based on attention pooling to convert the pixel-word level correlation into the pixel-sentence level correlation through the word-level attention weights.
3. The method for understanding referential expressions based on pixel-vocabulary association modeling according to claim 2, characterized in that: The word-level attention weights The expression is: ; In the formula, represents the nonlinear projection function; Represents the first i vocabulary; N l Indicates the number of words.
4. The method for understanding referential expressions based on pixel-vocabulary association modeling according to claim 3, characterized in that: The expression of the pixel-sentence level correlation is: ; In the formula, Indicates The association of words to pixels; is the sigmoid function; Indicates the number of words in the text expression.
5. The method for understanding referential expressions based on pixel-vocabulary association modeling according to claim 1, characterized in that: In step S2, the ResNet network is used to extract pixel features.
6. The method for understanding referential expressions based on pixel-vocabulary association modeling according to claim 5, characterized in that: A feature enhancement module is embedded after the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer of the ResNet network.
7. The method for understanding referential expressions based on pixel-vocabulary association modeling according to claim 6, characterized in that: A Transformer encoder is stacked after the ResNet network.
8. The method for understanding referential expressions based on pixel-vocabulary association modeling according to claim 7, characterized in that: A feature enhancement module is embedded before the linear projection of the Transformer encoder.
9. A referential expression understanding system based on pixel-word association modeling, applicable to the referential expression understanding method based on pixel-word association modeling according to any one of claims 1 to 8, characterized in that: It includes a text feature extraction module, a pixel feature extraction module, a pixel-word association enhancement module and a decoding module; The text feature extraction module is used to extract text features; The pixel feature extraction module is used to extract pixel features; The pixel-word association enhancement module is embedded after several convolutional layers in the pixel feature extraction module, and is used to obtain the correlation between pixels and text words, and embed the text features into the pixel features for enhancement based on the correlation; The decoding module is used to integrate the enhanced visual features and the original text features through a cross-attention mechanism, and output the final detection results using a forward neural network.
10. The referential expression understanding system based on pixel-vocabulary association modeling according to claim 9, characterized in that: Through cross-layer regularization loss For each pixel-word association enhancement module generated Apply distribution constraints for training: the true labels Convert to a binary mask of the same size as the input image , where the true label Set the pixel values inside to 1 and the rest to 0; adjust the mask The size of , so that it corresponds to different pixel-word association enhancement modules Keep the same spatial resolution; finally, perform spatial normalization to obtain the flattened and , mapped into a probability distribution through the Softmax function and Then, define the loss for and The KL divergence between is calculated as: ; ; ; In the formula, and Respectively and The elements, is the temperature parameter that controls the distribution shape in the Softmax function, represents the total number of pixel-word association-based feature enhancement modules in the embodiment, It is the product of width and height in the spatial resolution of the feature map.
Citation Information
Patent Citations
Image analysis method and system based on multi-modal information
CN116994069A
Priori knowledge guidance Transform method and system for remote sensing image description
CN119339228A
Self-adaptive indicator understanding method and system and storage medium
CN119378682A
Image description generation method and device based on multi-level interaction of visual features and text features
CN119741582A
Open-vocabulary object detection based on frozen vision and language models
WO2024006340A1