A Graphic-Text Retrieval Method Based on Text Semantic Guidance and Adaptive Feature Aggregation

By adopting a method based on text semantic guidance and adaptive feature aggregation in cross-modal retrieval, the problem of poor cross-modal alignment accuracy caused by inconsistent architecture of image and text feature extraction model is solved, and efficient and accurate graphic and text retrieval performance is achieved.

CN120030208BActive Publication Date: 2025-06-20YUNNAN NORMAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510500473.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-06-20
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing cross-modal retrieval methods have different semantic distributions between modals due to the inconsistent architecture of image and text feature extraction models, which in turn affects the accuracy of cross-modal alignment, increasing model training complexity and resource consumption.

Method used

The graphic and text search method based on text semantic guidance and adaptive feature aggregation is adopted, and the image feature enhancement module and text semantic guidance feature purification module are used to highlight key fine-grained features and filter redundant information to achieve compactness and effective purification of image feature representation. Then, through the adaptive aggregation network, the most discriminant features in each dimension are selected to perform feature mapping and similarity calculations for image and text.

Benefits of technology

It significantly improves the accuracy of cross-modal retrieval, reduces calculation costs, improves the training efficiency and performance of the model, and has good practical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030208B_ABST
    Figure CN120030208B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for image-text retrieval based on text semantic guidance and adaptive feature aggregation, belonging to related fields such as computer vision, image processing, and natural language processing. This method identifies the image patches in an image that are most relevant to the text description, uses a hardening method to reduce the redundancy of the image patches, highlights key fine-grained features and filters out unnecessary redundant information, making the image feature representation more compact and achieving effective feature purification. Secondly, based on the purified image features, an adaptive aggregation strategy is introduced before image-text matching, and the most representative features are selected for aggregation in each dimension of the unimodal features, achieving more efficient and accurate image-text retrieval. By optimizing the redundancy filtering and cross-modal alignment of image features, the present invention effectively solves the problems of semantic gap and low accuracy between different modalities in current traditional retrieval methods, and better meets the actual needs of users for cross-modal image-text retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a graphic and text retrieval method based on text semantic guidance and adaptive feature aggregation, belonging to related fields such as computer vision, image processing, and natural language processing. Background Art

[0002] In recent years, multi-modal data such as pictures, videos, voices, and texts widely exist in fields such as social media, digital libraries, and medical images, greatly enriching the form of information expression. However, with the continuous increase in information complexity, traditional single-modal data retrieval methods are difficult to meet the current Internet requirements. Therefore, researchers have begun to explore the possibility of using multi-modal data for retrieval tasks. How to mine and utilize the relationships between these multi-modal data has become an important challenge in the field of information technology. Cross-modal retrieval is a cross-research hotspot that integrates the fields of vision and language and has a close relationship with many tasks, such as image captioning and visual question answering. Its core lies in breaking the barriers between different modal data and realizing effective association and retrieval between them. In this field, cross-modal retrieval between images and texts is the most common research direction, including two sub-tasks: text-to-image retrieval and image-to-text retrieval.

[0003] Currently, most cross-modal retrieval methods still adhere to the detector-based image feature extraction paradigm, which exacerbates the semantic distribution gap between modalities due to the inconsistency with the text feature extraction model architecture, resulting in inaccurate alignment problems during cross-modal alignment. In addition, image data often contains rich background details and redundant information, while text descriptions tend to focus on key information points. This difference in information density and focus of attention makes it complex to directly align the features of the two modalities. Moreover, in order to truly reflect the real world, images contain a large amount of details and background information, resulting in a higher information redundancy compared to texts. This information, like impurities, not only increases the complexity of model training but also leads to a significant increase in resource consumption. To solve the above problems, a graphic and text retrieval method based on text semantic guidance and adaptive feature aggregation is proposed. This method first more accurately captures the key information in the image through an image feature enhancement module, thereby narrowing the difference in feature representation between the image and the text. Subsequently, text semantic guidance is used to identify the image patches in the image that are most relevant to the text description, and the hard segmentation method is adopted to reduce patch redundancy and unnecessary computational costs. Finally, in the feature aggregation stage, after obtaining the most discriminative features in each dimension of the image and the text through an adaptive aggregation network, the image and the text are mapped to the same feature space for similarity calculation. The method proposed by the present invention can significantly improve the accuracy of cross-modal retrieval and has certain practical application value. Summary of the Invention

[0004] The object of the present invention is to provide a method for image-text retrieval based on text semantic guidance and adaptive feature aggregation, aiming to solve the technical problems of the heterogeneous gap between image and text features and the low retrieval accuracy, so as to achieve the goal of efficient retrieval in large-scale cross-modal retrieval.

[0005] To achieve the above object, the specific technical solution of the present invention is: a method for image-text retrieval based on text semantic guidance and adaptive feature aggregation. For all images including the query image or all texts including the query text in the image-text library, the initial feature vectors of the images and the feature vectors of the texts are respectively extracted, a new image feature representation is generated through feature enhancement and purification, and the features of the images and texts are respectively aggregated through adaptive selection. Then, the cosine similarity is used to calculate the distance between the image and text, and the hinge-based triplet ranking loss is used to optimize the model, and finally the image-text retrieval task is completed.

[0006] Based on the Transformer architecture, the present invention emphasizes key fine-grained features and filters out unnecessary redundant information through an image feature enhancement module and a feature purification module guided by text semantics. This process enhances the compactness of the image feature representation and achieves effective feature purification. Then, the model adaptively selects the most representative features in each dimension for aggregation, thus achieving more efficient and accurate retrieval.

[0007] The specific steps are as follows:

[0008] Step1: Extract the features of the images and texts;

[0009] Step2: Enhance the extracted image features;

[0010] Step3: Calculate the saliency scores within the image modality and between the image-text modalities for the enhanced image features to obtain the purified image features;

[0011] Step4: Based on the purified image features, introduce an adaptive aggregation feature method to obtain a global feature vector containing the key information within each modality of the image and text;

[0012] Step5: Based on the obtained global feature vector, calculate the similarity between the image features and the text features to complete the image-text retrieval.

[0013] The specific steps of the said Step1 are:

[0014] Step 1.1: Extract the original features of the image as , where N represents the number of image patches in the picture, d is the dimension of the feature vector, represents the feature vector of the learned image patch, and R is the set of image patch features;

[0015] Step 1.2: Extract the original features of the text as , where l represents the number of words in the sentence, represents the feature vector of each word in the learned text, and T is the set of text fine-grained features.

[0016] The specific steps of Step 2 are as follows:

[0017] The enhancement module for image semantic features enables the model to focus more on the key information in the image, thus significantly improving the expression ability of image features. The following formula is used for image feature enhancement:

[0018]

[0019]

[0020] Among them, is the generated attention weight, mlp means sending the features after maxpooling and avgpooling into a multi-layer perceptron, means restricting the output to the range (0, 1), is the enhanced image patch feature, and then it is added to the original feature R to obtain the image patch embedding .

[0021] The specific steps of Step 3 are as follows:

[0022] The process of image feature purification integrates two strategies: text-supervised image and image self-supervision, that is, calculating the saliency scores between image-text modalities and within the image modality respectively, aiming to screen out the most concerned image patches and obtain the purified image features.

[0023] Step 3.1: Interact the feature representation of the image patch with the embedding vector of the text sequence, and generate a text-related saliency score for each image patch through cross-attention :

[0024]

[0025] Among them, is the fine-grained feature sequence of the image, is the global feature of the text, is the feature dimension, and the softmax function is used for normalization processing, and all the output values are non-negative and the sum is 1;

[0026] Step 3.2: Based on the interactions between the elements within the image, reveal the salient features within the image through the self-attention mechanism, and calculate the saliency scores of each image patch relative to other image patches :

[0027]

[0028] Among them, is the global feature of the image;

[0029] Step 3.3: By introducing the weight parameter weightedly fuse the saliency scores between the image-text modalities and within the image modality to obtain a comprehensive saliency score ;

[0030] Step 3.4: Convert the obtained comprehensive saliency score into a binary decision matrix , define a selection ratio γ, which determines that among the N image patches, the top image patches with the highest scores will be selected. At this time, the corresponding elements in the decision matrix are 1, while the remaining image patches are not selected, and the corresponding elements in the decision matrix are 0;

[0031] Step 3.5: Use the binary decision matrix as a mask to act on the original feature R. By retaining all corresponding feature rows, the finally purified image features are obtained, where represents the purified image patch feature vector.

[0032] The specific steps of the said Step 4 are as follows:

[0033] Adopt the following formula for feature aggregation:

[0034]

[0035]

[0036] Among them, and are the eigenvalue of the image and the text on dimension j respectively, is the function to select the maximum eigenvalue on each dimension, is the final global feature representation of the image or the text, represents the eigenvalue on each feature dimension.

[0037] Specifically, the above formula selects the most significant value for each dimension from all image patches or text word features by performing a maximum selection operation on the local features of each dimension, thereby constructing the final global feature representation. This aggregation method does not simply perform average pooling or max pooling, but rather performs a dynamic selection operation at the dimension level according to the input features , adaptively extracting the most significant and representative feature values for each dimension, ensuring that even in the case of uneven feature distributions among different dimensions, the key information for each dimension can be retained.

[0038] The specific steps of Step 5 are as follows:

[0039] Step 5.1: Based on the obtained global feature embeddings of the image and text, use cosine similarity to calculate the similarity value between the global features of the image and text; in a batch, the larger the similarity value, the more matching the image and text are, thereby querying the corresponding image or text.

[0040] Step 5.2: Combine the hinge-based triplet ranking loss and mean squared error loss for training to improve the robustness of the proposed method;

[0041] Step 5.2.1: Using the calculated image-text similarity value, reduce the distance between the matching image-text pairs in the common embedding space through the triplet ranking loss, and further increase the distance between the non-matching image-text pairs;

[0042]

[0043] Among them, is the alignment loss for cross-modal retrieval, represents a margin parameter, is the similarity function, , V and T represent the global feature embeddings of the image and text respectively, and B represents the mini-batch data. (V, T) is the matching image-text pair in the mini-batch data. Define and as examples of the most difficult negative image-text pairs in a mini-batch data.

[0044] Step 5.2.2: Use the selection ratio γ as a predefined value for training, and use the obtained binary decision matrix to train the model through the mean squared error loss;

[0045]

[0046] Among them, is the selection block ratio;

[0047] Step 5.3: Combine the alignment loss of cross-modal retrieval with the loss of the selection block ratio to obtain the final loss function.

[0048] The beneficial effects of the present invention are as follows: The present invention provides an image-text retrieval method based on text semantic guidance and adaptive feature aggregation. By means of an image feature enhancement module and a text semantic-guided feature purification module, key fine-grained features are highlighted and unnecessary redundant information is filtered out, making the image feature representation more compact and effectively improving the performance of image-text retrieval. Experimental results prove that the method proposed by the present invention shows good performance on two datasets, Flickr 30k and MS COCO. Compared with the baseline method, the maximum improvement in the RSUM index reaches 22.2%, fully verifying the effectiveness and advancement of the method of the present invention. Description of the Drawings

[0049] Figure 1 is a flowchart of the image-text retrieval method proposed by the present invention. Detailed Embodiments

[0050] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0051] Embodiment 1: As Figure 1 shown, an image-text retrieval method based on text semantic guidance and adaptive feature aggregation. In this embodiment, a multi-modal database consisting of one image corresponding to five text descriptions is taken as an example. Each image or each text sentence is used as a query object respectively, and the retrieval is completed by calculating the similarity between the query image and the text in the database. The specific process includes:

[0052] Step1: Extract the features of the image and the text.

[0053] Step 1.1: Extract the original features of the image as , where N represents the number of image patches in the picture, d is the dimension of the feature vector, represents the feature vector of the learned image patch, and R is the set of image patch features. For the given input image I, it is fed into the Swin Transformer composed of multiple self-attention layers, and then through a fully connected layer, a 512-dimensional initial image feature vector is obtained;

[0054] Step 1.2: Extract the original features of the text as , where l represents the number of words in the sentence, It is represented as the feature vector of each word in the learned text, and T is the set of text fine-grained features. Given a sentence T, the words in the sentence are processed by a pre-trained BERT model, and then through a fully connected layer, a 512-dimensional word embedding vector is obtained.

[0055] Step2: Enhance the extracted image features.

[0056] Through fine feature processing and enhancement, the model can focus more on the key information in the image, thus significantly improving the expression ability of image features. Max-pooling and average-pooling operations are respectively performed on the input features to capture significant feature information and overall feature information. The combination of these two pooling operations can comprehensively and effectively summarize the important characteristics of the feature channels. Then, a multi-layer perceptron and a Sigmoid activation function are introduced to obtain attention weights, realizing adaptive weighting of the feature channels. The following formula is used for image feature enhancement:

[0057]

[0058]

[0059] Among them, are the generated attention weights, mlp represents sending the features after maxpooling and avgpooling into the multi-layer perceptron, represents restricting the output to the range (0, 1), are the enhanced image patch features, and then after adding them to the original feature R, the image patch embedding is obtained . Using the idea of the residual network, the image patch embedding is obtained by combining with the initial feature as .

[0060] Step3: Calculate the saliency scores within the image modality and between the image-text modalities for the enhanced image features to obtain the purified image features.

[0061] Step 3.1: Interact the feature representation of the image patch with the embedding vector of the text sequence to learn the image patch most relevant to the semantic concept of the text. Through cross-attention calculation, a text-related attention score is generated for each image patch according to the following formula . The patches with high scores are usually closely related to the key information in the text and are important clues for understanding the overall context.

[0062]

[0063] Among them, is the fine-grained feature sequence of the image, is the global feature of the text, is the feature dimension. The softmax function performs normalization processing, and all the output values are non-negative and their sum is 1.

[0064] Step 3.2: Pay attention to the interactions between elements within the image. Without relying on external information, identify the key blocks only through the mutual influence within the image. Reveal the significant features within the image through the self-attention mechanism, and calculate the significance score of each block relative to other blocks according to the following formula .

[0065]

[0066] where, is the global feature of the image.

[0067] Step 3.3: By introducing the weight parameter weightedly fuse the significance scores between the image-text modalities and within the image modality to obtain a comprehensive significance score .

[0068] Step 3.4: Convert the obtained comprehensive significance score into a binary decision matrix , define a selection ratio γ, which determines that among N image blocks, the top image blocks with the highest scores will be selected, and at this time the corresponding elements in the decision matrix are 1, while the remaining image blocks will not be selected, and at this time the corresponding elements in the decision matrix are 0;

[0069] Step 3.5: Use the binary decision matrix as a mask to act on the original feature R, and obtain the final purified image feature by retaining all the corresponding feature rows, where represents the purified image block feature vector.

[0070] Step4: Based on the purified image features, introduce an adaptive aggregation feature method to obtain a global feature vector that contains the key information within each modality of the image and text.

[0071] Through the above steps, an enhanced and purified image feature sequence is obtained. Subsequently, using the feature aggregation module, these fine-grained features are integrated into a global feature vector that can comprehensively represent the image and text. This method first performs a dimension-by-dimension analysis on the selected features, and uses the max-pooling operation on the dimension in the following formula to independently extract the most significant and representative feature values from each dimension. Ensure that even in the case of uneven feature distributions between different dimensions, the key information in each dimension can be accurately captured.

[0072]

[0073]

[0074] Among them, and are the eigenvalues of the image and text on dimension j respectively, is the function to select the maximum eigenvalue on each dimension, is the final global feature representation of the image or text, represents the eigenvalues on each feature dimension.

[0075] Step5: Based on the obtained global feature vector, calculate the similarity between the image feature and the text feature to complete the image-text retrieval.

[0076] Step 5.1: Based on the obtained global feature embeddings of the image and text, use the cosine similarity to calculate the similarity value between the global features of the image and text;

[0077] Step 5.2: Adopt online hard negative mining to optimize the model based on the hinge-based triplet ranking loss. By reducing the distance between the matching image-text pairs in the common embedding space, the distance between the non-matching image-text pairs is further enlarged.

[0078]

[0079] Among them, is the alignment loss for cross-modal retrieval, represents a margin parameter, is the similarity function, , V and T represent the global feature embeddings of the image and text respectively, and B represents the mini-batch data. (V, T) is the matching image-text pair in the mini-batch data. Define and as examples of the most difficult negative image-text pairs in a mini-batch data.

[0080] Step 5.3: For the filtering of redundant image patches, train the model through the mean squared error loss.

[0081]

[0082] Among them, is the selection block ratio.

[0083] Step 5.4: Combine the alignment loss of cross-modal retrieval with the loss of the selection block ratio to obtain the final loss function .

[0084] To verify the effectiveness of the present invention, the invention selected two of the most widely used benchmark datasets in the field of image-text retrieval, Filckr30k and MS-COCO, for experiments. These two datasets consist of paired images and texts. The Flickr30K dataset has 31,783 images, and each image corresponds to 5 texts. According to the general division setting, during the experiment, the dataset is divided into three parts. The training set contains 29,783 images, the validation set contains 1,000 images, and the test set contains 1,000 images. The MSCOCO dataset has 123,287 images. Similarly, according to the general division, 113,287 are used for training, 5,000 for validation, and 5,000 for testing. Each image corresponds to 5 texts, and the average length of the texts is 8.7. For fair comparison, according to previous standards, the experimental results of 5K and 1K of the MSCOCO dataset are provided respectively. The result of 5K is the average result of 5 times of 1K cross-validation. All experiments in this paper were carried out on Nvidia A100 GPUs. All Python in the experimental environment is 3.8.0, the Pytorch version is 2.3.0, and the CUDA version is 12.4. For image input, Swin transformer is used as the visual encoder (the size of an image patch is 32*32 pixels, and the number is 12*12). For text input, the BERT model is used to pre-train the masked language task of English sentences. For the transformer encoders of both modes, the feature dimensions are unified to . The model is optimized using the AdamM optimizer, trained for 30 epochs, and the batchsize sizes of Flickr30k and MSCOCO are set to 128 and 256 respectively. In the loss function, the default margin parameter , the weight parameter , and the ratio of selecting image patches .

[0085] Table 1 Comparison of image-text retrieval performance on the flickr30k test set

[0086]

[0087] Table 1 shows the quantitative results of the existing methods and the method proposed in the present invention on the flickr30K test set. The method TGAA proposed in the present invention is superior to other methods in all evaluation metrics. Among them, the optimal results are highlighted in bold. Specifically, the method proposed in this paper has a 29.8% higher value on RSUM than the baseline model EAFG (from 510.9 to 540.7), and also completely outperforms other methods in terms of the R@K metric.

[0088] Table 2 Comparison of Image-Text Retrieval on the MSCOCO 5-fold 1k Test Set

[0089]

[0090] Table 3 Comparison of Image-Text Retrieval on the MSCOCO 5k Test Set

[0091]

[0092] Tables 2 and 3 are the quantitative evaluation results of the existing methods and the method proposed in the present invention on the MSCOCO test set. Among them, Table 2 shows the retrieval performance on MSCOCO 5-fold 1k. Compared with the baseline model EAFG, the method proposed in the present invention has achieved a significant 15% improvement in the comprehensive evaluation index RSUM (from 526.1 to 541.1), and is also superior to other methods in the R@K index. Table 3 shows the retrieval performance on MSCOCO 5k. Compared with all other methods, the method of the present invention also remains leading. Among them, compared with the baseline model, a significant improvement of 29.2% has been achieved (from 433.2 to 462.4). The results show that the method proposed in the present invention has better retrieval performance than other methods in the two retrieval tasks.

[0093] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A method for image-text retrieval based on text semantics guidance and adaptive feature aggregation, characterized in that: The specific steps of the method are as follows: Step 1: Extract features of images and text; Step 2: Enhance the extracted image features; Step 3: Calculate the significance scores of the enhanced image features within the image modality and between the image and text modalities to obtain the purified image features; Step 4: Based on the purified image features, an adaptive aggregation feature method is introduced to obtain a global feature vector containing key information in each modality of image and text; Step 5: Based on the obtained global feature vector, calculate the similarity between image features and text features to complete image and text retrieval; The specific steps of Step 2 are: The following formula is used to enhance image features: ; ; in, is the generated attention weight, mlp means sending the features after maxpooling and avgpooling into the multi-layer perceptron. It means limiting the output to the range of (0, 1). is the enhanced image block feature, which is then added to the original feature R to obtain the image block embedding ; The specific steps of Step 3 are: Step 3.1: Interact the feature representation of the image patch with the embedding vector of the text sequence, and generate a text-related saliency score for each image patch through cross attention. ; Step 3.2: Based on the interaction between the elements inside the image, the self-attention mechanism is used to reveal the salient features inside the image and calculate the saliency score of each image block relative to other image blocks. ; Step 3.3: By introducing weight parameters The saliency scores between image-text modalities and within image modalities are weighted and fused to obtain a comprehensive saliency score. ; Step 3.4: Convert the obtained comprehensive significance score into a binary decision matrix , define a selection ratio γ, determine the top score among N image blocks The image blocks will be selected, and the corresponding decision matrix elements are 1, while the remaining image blocks will not be selected, and the corresponding decision matrix elements are 0; Step 3.5: Binary decision matrix , acts as a mask on the original feature R, by retaining all The corresponding feature rows get the final purified image features ,in Represented as the purified image block feature vector; The specific steps of Step 4 are: The following formula is used to aggregate features: ; ; in, and are the feature values ​​of image and text on dimension j, is a function that selects the largest eigenvalue in each dimension, It is the final global feature representation of the image or text. Represents the feature value on each feature dimension.

2. According to claim 1, a method for image-text retrieval based on text semantic guidance and adaptive feature aggregation is characterized in that: The specific steps of Step 1 are: Step 1.1: Extract the original features of the image as , where N represents the number of image blocks in the image, and d is the dimension of the feature vector. Represented as the feature vector of the learned image block, R is the set of image block features; Step 1.2: Extract the original features of the text as , where l represents the number of words in the sentence, Represented as the learned feature vector of each word in the text, T is a set of fine-grained features of the text.

3. The image-text retrieval method based on text semantic guidance and adaptive feature aggregation according to claim 1, characterized in that: The specific steps of Step 5 are: Step 5.1: Based on the global feature embedding of the image and text, the similarity value between the global features of the image and text is obtained by using cosine similarity calculation; Step 5.2: Use hinge-based triple ranking loss and mean square error loss for training; Step 5.2.1: Using the calculated image-text similarity values, the distance between matching image-text pairs in the common embedding space is reduced through the triple ranking loss, thereby increasing the distance between mismatching image-text pairs; Step 5.2.2: The selection ratio γ is used as the predefined value for training, and the obtained binary decision matrix is ​​used to train the model through the mean square error loss; Step 5.3: Combine the alignment loss of cross-modal retrieval and the loss of selected block ratio to obtain the final loss function.

Citation Information

Patent Citations

  • Learning beautiful and ugly visual attributes

    US20150055854A1