Image-text retrieval method based on text semantic guidance and adaptive feature aggregation

By adopting a method based on text semantic guidance and adaptive feature aggregation in cross-modal retrieval, the problem of heterogeneous gap between image and text features is solved, and more efficient and accurate graphic and text retrieval performance is achieved.

CN120030208AActive Publication Date: 2025-05-23YUNNAN NORMAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510500473.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-23
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The heterogeneous gap between image and text features and the problem of low retrieval accuracy among existing cross-modal search methods.

Method used

Using the graphic and text search method based on text semantic guidance and adaptive feature aggregation, the image feature enhancement module and the text semantic guidance feature purification module are used to highlight key fine-grained features and filter out unnecessary redundant information, so as to realize compact feature representation and efficient retrieval of images and text.

Benefits of technology

The accuracy of cross-modal retrieval was significantly improved. Through experiments on the Flickr 30k and MS COCO datasets, the maximum improvement of the RSUM indicator reached 22.2%, verifying the effectiveness and advancedness of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030208A_ABST
    Figure CN120030208A_ABST
Patent Text Reader

Abstract

The invention relates to an image-text retrieval method based on text semantic guidance and adaptive feature aggregation, and belongs to the related fields of computer vision, image processing, natural language processing and the like. According to the method, the image blocks most related to text description in the image are identified, the redundancy of the image blocks is reduced by adopting a hardening score method, key fine-grained features are highlighted, and unnecessary redundant information is filtered, so that image feature representation is more compact, and effective feature purification is realized. Secondly, on the basis of the purified image features, an adaptive aggregation strategy is introduced before image text matching, the most representative features are selected on each dimension of the single-mode features for aggregation, and more efficient and accurate image-text retrieval is achieved. According to the method, by optimizing redundancy filtering and cross-modal alignment of image features, the problems of semantic gaps between different modals and low accuracy in a current traditional retrieval method are effectively solved, and the actual demand of a user for cross-modal image-text retrieval is better met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a picture and text retrieval method based on text semantic guidance and adaptive feature aggregation, and belongs to related fields such as computer vision, image processing, and natural language processing. Background Art

[0002] In recent years, multimodal data such as images, videos, voices, and texts have been widely used in social media, digital libraries, medical imaging, and other fields, greatly enriching the expression of information. However, with the increasing complexity of information, traditional single-modal data retrieval methods are difficult to meet the needs of the current Internet. Therefore, researchers have begun to explore the possibility of using multimodal data for retrieval tasks. How to mine and utilize the relationship between these multimodal data has become an important challenge in the field of information technology. Cross-modal retrieval is an interdisciplinary research hotspot that integrates vision and language fields and is closely related to many tasks, such as image captioning and visual question answering. Its core lies in breaking down the barriers between different modal data and achieving effective association and retrieval between them. In this field, cross-modal retrieval between images and texts is the most common research direction, including two subtasks: text-to-image retrieval and image-to-text retrieval.

[0003] At present, most cross-modal retrieval methods still adhere to the detector-based image feature extraction paradigm, which makes the inconsistency with the text feature extraction model architecture aggravate the semantic distribution gap between modalities and lead to the problem of inaccurate alignment in the cross-modal alignment process. In addition, image data often contains rich background details and redundant information, while text descriptions tend to focus on key information points. This difference in information density and focus makes it complicated to directly align the two modal features. In order to truly reflect the real world, images contain a lot of details and background information, which makes their information redundancy higher than that of text. This information, like impurities, not only increases the complexity of model training, but also leads to a significant increase in resource consumption. In order to solve the above problems, a text-image retrieval method based on text semantics guidance and adaptive feature aggregation is proposed. The method first captures the key information in the image more accurately through the image feature enhancement module, thereby narrowing the difference in feature representation between image and text. Subsequently, the text semantics are used to guide the identification of the image blocks that are most relevant to the text description, and the hardening method is used to reduce block redundancy and reduce unnecessary computational costs. Finally, in the feature aggregation stage, after obtaining the most discriminative features in each dimension of the image and text through the adaptive aggregation network, the image and text are mapped to the same feature space for similarity calculation. The method proposed in the present invention can significantly improve the accuracy of cross-modal retrieval and has certain practical application value. Summary of the invention

[0004] The purpose of the present invention is to provide a picture-text retrieval method based on text semantic guidance and adaptive feature aggregation, aiming to solve the technical problems of heterogeneous gap between image and text features and low retrieval accuracy, so as to achieve the goal of efficient retrieval in large-scale cross-modal retrieval.

[0005] To achieve the above-mentioned purpose, the specific technical scheme of the present invention is: a method for image-text retrieval based on text semantic guidance and adaptive feature aggregation, for all images including the query image or all texts including the query text in the image-text library, the initial feature vectors of the image and the feature vectors of the text are extracted respectively, a new image feature representation is generated through feature enhancement and purification, and the features of the image and text are aggregated respectively through adaptive selection, and then the cosine similarity is used to calculate the distance between the image and the text, and the hinge-based triple ranking loss is used to optimize the model, and finally the image-text retrieval task is completed.

[0006] Based on the Transformer architecture, the present invention emphasizes key fine-grained features and filters out unnecessary redundant information through an image feature enhancement module and a text semantics-guided feature purification module. This process enhances the compactness of image feature representation and achieves effective feature purification. Then, the model adaptively selects the most representative features in each dimension for aggregation, thereby achieving more efficient and accurate retrieval.

[0007] The specific steps are as follows:

[0008] Step 1: Extract features of images and text;

[0009] Step 2: Enhance the extracted image features;

[0010] Step 3: Calculate the significance scores of the enhanced image features within the image modality and between the image and text modalities to obtain the purified image features;

[0011] Step 4: Based on the purified image features, an adaptive aggregation feature method is introduced to obtain a global feature vector containing key information in each modality of image and text;

[0012] Step 5: Based on the obtained global feature vector, calculate the similarity between image features and text features to complete image and text retrieval.

[0013] The specific steps of Step 1 are:

[0014] Step 1.1: Extract the original features of the image as , where N represents the number of image blocks in the image, and d is the dimension of the feature vector. Represented as the feature vector of the learned image block, R is the set of image block features;

[0015] Step 1.2: Extract the original features of the text as , where l represents the number of words in the sentence, Represented as the learned feature vector of each word in the text, T is a set of fine-grained features of the text.

[0016] The specific steps of Step 2 are:

[0017] The enhancement module for image semantic features enables the model to focus more on the key information in the image, thereby significantly improving the expressiveness of image features. The following formula is used to enhance image features:

[0018]

[0019]

[0020] in, is the generated attention weight, mlp means sending the features after maxpooling and avgpooling into the multi-layer perceptron. It means limiting the output to the range of (0, 1). is the enhanced image block feature, which is then added to the original feature R to obtain the image block embedding .

[0021] The specific steps of Step 3 are:

[0022] The process of image feature purification combines two strategies: text-supervised image and image self-supervision. That is, the saliency scores between image-text modalities and within image modalities are calculated respectively, aiming to screen out the most concerned image blocks and obtain the purified image features.

[0023] Step 3.1: Interact the feature representation of the image patch with the embedding vector of the text sequence, and generate a text-related saliency score for each image patch through cross attention. :

[0024]

[0025] in, is the fine-grained feature sequence of the image, is the global feature of the text, is the feature dimension, and the softmax function performs normalization, so all output values ​​are non-negative and sum to 1;

[0026] Step 3.2: Based on the interaction between the elements inside the image, the self-attention mechanism is used to reveal the salient features inside the image and calculate the saliency score of each image block relative to other image blocks. :

[0027]

[0028] in, It is the global feature of the image;

[0029] Step 3.3: By introducing weight parameters The saliency scores between image-text modalities and within image modalities are weighted and fused to obtain a comprehensive saliency score. ;

[0030] Step 3.4: Convert the obtained comprehensive significance score into a binary decision matrix , define a selection ratio γ, determine the top score among N image blocks The image blocks will be selected, and the corresponding decision matrix elements are 1, while the remaining image blocks will not be selected, and the corresponding decision matrix elements are 0; Step 3.5: Binary decision matrix , acts as a mask on the original feature R, by retaining all The corresponding feature rows get the final purified image features ,in Represented as the purified image block feature vector.

[0031] The specific steps of Step 4 are:

[0032] The following formula is used to aggregate features:

[0033]

[0034]

[0035] in, and are the feature values ​​of image and text on dimension j, is a function that selects the largest eigenvalue in each dimension, It is the final global feature representation of the image or text. Represents the feature value on each feature dimension.

[0036] Specifically, the above formula performs a maximum selection operation on the local features of each dimension, selects the most significant value of each dimension from all image blocks or text word features, and forms the final global feature representation. This aggregation method is not a simple average pooling or maximum pooling process, but a dynamic selection operation at the dimension level based on the input features. , adaptively extracting the most significant and representative eigenvalues ​​in each dimension, ensuring that the key information in each dimension can be retained even when the features are unevenly distributed among different dimensions.

[0037] The specific steps of Step 5 are:

[0038] Step 5.1: Based on the global feature embedding of the image and text, the cosine similarity is used to calculate the similarity value between the global features of the image and text; in a batch, the larger the similarity value, the more the image and text match, so as to query the corresponding image or text.

[0039] Step 5.2: The hinge-based triple ranking loss and mean square error loss are combined for training to improve the robustness of the proposed method;

[0040] Step 5.2.1: Using the calculated image-text similarity values, the distance between matching image-text pairs in the common embedding space is reduced through triple ranking loss, thereby increasing the distance between mismatching image-text pairs.

[0041]

[0042] in, is the alignment loss for cross-modal retrieval, represents a marginal parameter, is a similarity function, , V and T represent the global feature embedding of image and text respectively, and B represents the mini-batch data. (V, T) is the matching image-text pair in the mini-batch data. Definition as well as As an example of the hardest negative image-text pair in a mini-batch.

[0043] Step 5.2.2: The selection ratio γ is used as the predefined value for training, and the obtained binary decision matrix is ​​used to train the model through the mean square error loss;

[0044]

[0045] in, To select the block ratio;

[0046] Step 5.3: Combine the alignment loss of cross-modal retrieval with the loss of the selection block ratio to obtain the final loss function.

[0047] The beneficial effects of the present invention are as follows: The present invention provides an image-text retrieval method based on text semantic guidance and adaptive feature aggregation. By means of an image feature enhancement module and a text semantic-guided feature purification module, key fine-grained features are highlighted and unnecessary redundant information is filtered out, making the image feature representation more compact and effectively improving the performance of image-text retrieval. Experimental results prove that the method proposed by the present invention shows good performance on two datasets, Flickr 30k and MS COCO. Compared with the baseline method, the maximum improvement in the RSUM index reaches 22.2%, fully verifying the effectiveness and advancement of the method of the present invention. Brief Description of the Drawings

[0048] Figure 1 is a flowchart of the image-text retrieval method proposed by the present invention. Detailed Embodiments

[0049] The present invention will be further described below in conjunction with the drawings and detailed embodiments.

[0050] Embodiment 1: As Figure 1 shown, an image-text retrieval method based on text semantic guidance and adaptive feature aggregation. In this embodiment, a multi-modal database consisting of one image corresponding to five text descriptions is taken as an example. Each image or each text sentence is used as a query object respectively, and the retrieval is completed by calculating the similarity between the query image and the text in the database. The specific process includes:

[0051] Step1: Extract the features of the image and the text.

[0052] Step 1.1: Extract the original features of the image as , where N represents the number of image patches in the picture, d is the dimension of the feature vector, represents the feature vector of the learned image patch, and R is the set of image patch features. For the given input image I, it is fed into the Swin Transformer composed of multiple self-attention layers, and then through a fully connected layer, a 512-dimensional initial image feature vector is obtained;

[0053] Step 1.2: Extract the original features of the text as , where l represents the number of words in the sentence, Denoted as the feature vectors for each word of the learned text, T is the set of text fine-grained features. Given a sentence T, the words in the sentence are processed by a pre-trained BERT model, and then through a fully connected layer, a 512-dimensional word embedding vector is obtained.

[0054] Step2: Enhance the extracted image features.

[0055] Through fine feature processing and enhancement, the model can focus more on the key information in the image, thus significantly improving the expression ability of image features. Max pooling and average pooling operations are respectively performed on the input features to capture significant feature information and overall feature information. The combination of these two pooling operations can comprehensively and effectively summarize the important characteristics of the feature channels. Then, a multi-layer perceptron and a Sigmoid activation function are introduced to obtain attention weights, realizing the adaptive weighting of the feature channels. The following formula is used for the enhancement of image features:

[0056]

[0057]

[0058] Among them, is the generated attention weight, mlp represents sending the features after maxpooling and avgpooling into the multi-layer perceptron, represents restricting the output to the range (0, 1), is the enhanced image patch feature, and then after adding it to the original feature R, the image patch embedding is obtained . Using the idea of the residual network, the image patch embedding is obtained by combining with the initial feature as .

[0059] Step3: Calculate the saliency scores within the image modality and between the image-text modalities for the enhanced image features to obtain the purified image features.

[0060] Step 3.1: Interact the feature representation of the image patch with the embedding vector of the text sequence to learn the image patch most relevant to the semantic concept of the text. Through cross-attention calculation, a text-related attention score is generated for each image patch according to the following formula . The patches with high scores are usually closely related to the key information in the text and are important clues for understanding the overall context.

[0061]

[0062] Among them, is the fine-grained feature sequence of the image, is the global feature of the text, is the feature dimension, and the softmax function performs normalization so that all output values ​​are non-negative and sum to 1.

[0063] Step 3.2: Focus on the interactions between elements within the image. Without relying on external information, identify key blocks only through the interactions within the image. Use the self-attention mechanism to reveal the salient features within the image, and calculate the saliency score of each block relative to other blocks according to the following formula: .

[0064]

[0065] in, It is the global feature of the image.

[0066] Step 3.3: By introducing weight parameters The saliency scores between image-text modalities and within image modalities are weighted and fused to obtain a comprehensive saliency score. .

[0067] Step 3.4: Convert the obtained comprehensive significance score into a binary decision matrix , define a selection ratio γ, determine the top score among N image blocks The image blocks will be selected, and the corresponding decision matrix elements are 1, while the remaining image blocks will not be selected, and the corresponding decision matrix elements are 0; Step 3.5: Binary decision matrix , acts as a mask on the original feature R, by retaining all The corresponding feature rows get the final purified image features ,in Represented as the purified image block feature vector.

[0068] Step 4: Based on the purified image features, an adaptive aggregation feature method is introduced to obtain a global feature vector that contains the key information in each modality of the image and text.

[0069] Through the above steps, we obtain an enhanced and purified image feature sequence. Then, we use the feature aggregation module to integrate these fine-grained features into a global feature vector that can fully characterize the image and text. This method first analyzes the filtered features dimension by dimension, and uses the max-pooling operation on the dimension in the following formula to independently extract the most significant and representative feature values ​​from each dimension. This ensures that even when the features are unevenly distributed between different dimensions, the key information on each dimension can be accurately captured.

[0070]

[0071]

[0072] in, and are the feature values ​​of image and text on dimension j, is a function that selects the largest eigenvalue in each dimension, It is the final global feature representation of the image or text. Represents the feature value on each feature dimension.

[0073] Step 5: Based on the obtained global feature vector, calculate the similarity between image features and text features to complete image and text retrieval.

[0074] Step 5.1: Based on the global feature embedding of the image and text, the similarity value between the global features of the image and text is obtained by using cosine similarity calculation;

[0075] Step 5.2: Use online hard negative mining to optimize the hinge-based triple ranking loss to optimize the model. This is done by reducing the distance between matching image-text pairs in the common embedding space, thereby increasing the distance between mismatching image-text pairs.

[0076]

[0077] in, is the alignment loss for cross-modal retrieval, represents a marginal parameter, is a similarity function, , V and T represent the global feature embedding of image and text respectively, and B represents the mini-batch data. (V, T) is the matching image-text pair in the mini-batch data. Definition as well as As an example of the hardest negative image-text pair in a mini-batch.

[0078] Step 5.3: For filtering redundant blocks of image blocks, the model is trained using mean square error loss.

[0079]

[0080] in, To select the block ratio.

[0081] Step 5.4: Alignment loss for cross-modal retrieval Loss with selected block ratio Combined to get the final loss function .

[0082] In order to verify the effect of the present invention, the inventor selected two most widely used benchmark datasets in the field of image and text retrieval, Filckr30k and MS-COCO, for experiments. These two datasets consist of paired images and texts. The Flickr30K dataset has 31,783 images, and each image corresponds to 5 texts. According to the general partition setting, the dataset was divided into three parts during the experiment. The training set contains 29,783 images, the validation set contains 1,000 images, and the test set contains 1,000 images. The MSCOCO dataset has 123,287 images. According to the general partition, 113,287 images are used for training, 5,000 images are used for validation, and 5,000 images are used for testing. Each image corresponds to 5 texts, and the average length of the text is 8.7. For fair comparison, according to previous standards, the experimental results of the MSCOCO dataset 5K and 1K are provided respectively, where the 5K result is the average result of 5 1K cross-validations. All experiments in this paper were conducted on Nvidia A100 GPUs. The experimental environment used Python 3.8.0, Pytorch version 2.3.0, and CUDA version 12.4. For image input, the Swin transformer was used as the visual encoder (the size of an image block is 32*32 pixels, and the number is 12*12). For text input, the BERT model was pre-trained on the mask language task of English sentences. For the transformer encoders of the two modes, the feature dimensions were unified as The model is optimized using the AdamM optimizer and trained for 30 epochs, with batch sizes of 128 and 256 for Flickr30k and MSCOCO, respectively. In the loss function, the default margin parameter , weight parameter , select the ratio of the image block .

[0083] Table 1 Comparison of image-text retrieval performance on the flickr30k test set

[0084] Table 1 shows the quantitative results of the existing methods and the proposed method on the flickr30K test set. The proposed method TGAA outperforms other methods in all evaluation indicators. The best results are highlighted in bold. Specifically, the proposed method is 29.8% higher than the baseline model EAFG in RSUM (from 510.9 to 540.7), and is also completely higher than the performance of other methods in the R@K indicator.

[0085] Table 2 Comparison of image-text retrieval on the MSCOCO 5-fold 1k test set

[0086] Table 3 Comparison of image-text retrieval on the MSCOCO 5k test set

[0087] Tables 2 & 3 are the quantitative evaluation results of the existing methods and the method proposed in the present invention on the MSCOCO test set. Table 2 shows the retrieval performance on MSCOCO 5-fold 1k. Compared with the baseline model EAFG, the method proposed in the present invention achieves a significant 15% improvement (from 526.1 to 541.1) in the comprehensive evaluation index RSUM, and is also superior to other methods in the R@K indicator. Table 3 shows the retrieval performance on MSCOCO5k. Compared with all other methods, the method of the present invention also maintains a leading position. Compared with the baseline model, a significant improvement of 29.2% (from 433.2 to 462.4) was achieved. The results show that in the two retrieval tasks, the method proposed in the present invention has better retrieval performance than other methods.

[0088] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.

Claims

1. A method for image-text retrieval based on text semantics guidance and adaptive feature aggregation, characterized in that: The specific steps of the method are as follows: Step 1: Extract features of images and text; Step 2: Enhance the extracted image features; Step 3: Calculate the significance scores of the enhanced image features within the image modality and between the image and text modalities to obtain the purified image features; Step 4: Based on the purified image features, an adaptive aggregation feature method is introduced to obtain a global feature vector containing key information in each modality of image and text; Step 5: Based on the obtained global feature vector, calculate the similarity between image features and text features to complete image and text retrieval.

2. A method for image-text retrieval based on text semantic guidance and adaptive feature aggregation according to claim 1, characterized in that: The specific steps of Step 1 are: Step 1.1: Extract the original features of the image as , where N represents the number of image blocks in the image, and d is the dimension of the feature vector. Represented as the feature vector of the learned image block, R is the set of image block features; Step 1.2: Extract the original features of the text as , where l represents the number of words in the sentence, Represented as the learned feature vector of each word in the text, T is a set of fine-grained features of the text.

3. The image-text retrieval method based on text semantic guidance and adaptive feature aggregation according to claim 1, characterized in that: The specific steps of Step 2 are: The following formula is used to enhance image features: ; ; in, is the generated attention weight, mlp means sending the features after maxpooling and avgpooling into the multi-layer perceptron. It means limiting the output to the range of (0, 1). is the enhanced image block feature, which is then added to the original feature R to obtain the image block embedding .

4. The method for image-text retrieval based on text semantic guidance and adaptive feature aggregation according to claim 1, characterized in that: The specific steps of Step 3 are: Step 3.1: Interact the feature representation of the image patch with the embedding vector of the text sequence, and generate a text-related saliency score for each image patch through cross attention. ; Step 3.2: Based on the interaction between the elements inside the image, the self-attention mechanism is used to reveal the salient features inside the image and calculate the saliency score of each image block relative to other image blocks. ; Step 3.3: By introducing weight parameters The saliency scores between image-text modalities and within image modalities are weighted and fused to obtain a comprehensive saliency score. ; Step 3.4: Convert the obtained comprehensive significance score into a binary decision matrix , define a selection ratio γ, determine the top score among N image blocks The image blocks will be selected, and the corresponding decision matrix elements are 1, while the remaining image blocks will not be selected, and the corresponding decision matrix elements are 0; Step 3.5: Binary decision matrix , acts as a mask on the original feature R, by retaining all The corresponding feature rows get the final purified image features ,in Represented as the purified image block feature vector.

5. The method for image-text retrieval based on text semantic guidance and adaptive feature aggregation according to claim 1, characterized in that: The specific steps of Step 4 are: The following formula is used to aggregate features: ; ; in, and are the feature values ​​of image and text on dimension j, is a function that selects the largest eigenvalue in each dimension, It is the final global feature representation of the image or text. Represents the feature value on each feature dimension.

6. The method for image-text retrieval based on text semantic guidance and adaptive feature aggregation according to claim 1, characterized in that: The specific steps of Step 5 are: Step 5.1: Based on the global feature embedding of the image and text, the similarity value between the global features of the image and text is obtained by using cosine similarity calculation; Step 5.2: Use hinge-based triple ranking loss and mean square error loss for training; Step 5.2.1: Using the calculated image-text similarity values, the distance between matching image-text pairs in the common embedding space is reduced through triple ranking loss, thereby increasing the distance between mismatching image-text pairs. Step 5.2.2: The selection ratio γ is used as the predefined value for training, and the obtained binary decision matrix is ​​used to train the model through the mean square error loss; Step 5.3: Combine the alignment loss of cross-modal retrieval and the loss of selected block ratio to obtain the final loss function.

Citation Information

Patent Citations

  • Learning beautiful and ugly visual attributes

    US20150055854A1