A cross-modal image text retrieval method based on deep learning
By designing a multimodal feature extraction module and cross-attention mechanism based on BERT and ResNet, combined with the OpenCLIP model and fast and slow model strategy, the shortcomings of the existing technology in cross-modal semantic alignment and challenging candidate set processing are solved, and efficient and accurate cross-modal retrieval effect is achieved.
Patent Information
- Application Number
- CN202411675311.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-11-19
- Filing Date
- 2024-11-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-11-21
AI Technical Summary
Existing cross-modal image text retrieval techniques perform poorly in fine-grained cross-modal semantic alignment and more challenging candidate set processing, and existing benchmarks are not sufficient to validate the real-modal fine-grained semantic understanding of cross-modal.
A multimodal feature extraction module based on BERT and ResNet is designed to acquire deep feature representations of text and images and to align features through cross attention mechanisms. At the same time, the OpenCLIP model was introduced to enhance the depth and accuracy of feature extraction, and a fast and slow model strategy was adopted to improve retrieval speed and accuracy.
It significantly improves the accuracy and efficiency of cross-modal retrieval, can more accurately identify and match the correlation between images and text, improves the speed and accuracy of information retrieval, and shows excellent performance in practical industrial applications.
Smart Images

Figure CN119311911B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multimodal feature alignment method for cross-modal retrieval of images and texts, and in particular to a cross-modal image-text retrieval method based on deep learning, and belongs to the technical fields of artificial intelligence, computer vision and natural language processing. Background Art
[0002] With the vigorous development of artificial intelligence and big data technology, multimodal data has become a key resource in various industries, especially in e-commerce, media, education and medical fields, where the combined use of image and text data is particularly important. An effective cross-modal image-text retrieval model is of great significance for improving information retrieval efficiency, enhancing data analysis capabilities, and promoting the development of intelligent applications. First, the importance of multimodal data is reflected in its ability to provide more comprehensive information. Multi-source and multimodal data can provide more accurate information and increase the robustness of information presentation and expression through mutual support, supplementation and correction. For example, in fields such as medical diagnosis and weather forecasting, multi-source data fusion and integration is an important basis for improving the quality of data analysis. Second, the necessity of cross-modal image-text retrieval software lies in its ability to improve user experience and work efficiency. When faced with massive data, users can quickly locate images related to the query text, or retrieve corresponding text information from the image, which greatly improves the speed and accuracy of information retrieval. However, at this stage, cross-modal image-text retrieval technology still faces some challenges and shortcomings:
[0003] First, the existing image and text data association retrieval technology is still in its infancy, mainly relying on single-modal retrieval methods, such as text-based search or image-based search, which limits the efficiency of comprehensive data utilization. Second, although large-scale multimodal pre-training models have brought significant performance improvements to image-text retrieval, there is still much room for improvement in fine-grained cross-modal semantic alignment, especially in actual industrial applications. Third, the performance of existing image-text retrieval models will drop significantly when faced with more challenging candidate sets, which indicates that current benchmarks may not be sufficient to verify the true model capabilities of cross-modal fine-grained semantic understanding.
[0004] Faced with these challenges, researchers have tried a variety of methods to improve cross-modal image text retrieval technology:
[0005] First, traditional statistical analysis techniques and deep learning-based techniques are widely used for feature extraction and content relevance measurement of cross-modal data. Second, hash coding methods improve retrieval speed and reduce storage space by converting multimedia data into binary codes. In particular, deep hashing methods achieve a good balance between retrieval accuracy and efficiency. Third, generative adversarial networks (GANs) are used to retrieve images by generating image data from text data, effectively reducing cross-modal differences. Fourth, cross-attention (CA) methods extract features by considering dense pairwise cross-modal interactions and tend to obtain high-accuracy retrieval results. Summary of the invention
[0006] In order to improve the accuracy and efficiency of cross-modal retrieval of images and texts, the key is to understand the semantic associations between different modal data, create a cross-modal knowledge sharing system that can adapt to diverse data features, and learn the invariant and dynamic feature representations between images and texts. Specifically, in order to obtain the invariant feature representations in cross-modal retrieval, in the feature extraction stage, a multimodal feature extraction module based on BERT and ResNet is designed to obtain the deep feature representations of text and images, and merge the text and image features from different data sources into a unified and high-dimensional feature representation. The core of this stage is to capture the semantic information of the text and the visual content of the image, laying the foundation for subsequent cross-modal alignment. After obtaining a stable cross-modal feature representation, the relevance scoring stage combines the dynamic association information of images and texts with the static semantic information by designing a feature alignment strategy based on the cross-attention mechanism to enhance cross-modal feature embedding. The core of this stage is to improve the accuracy of cross-modal retrieval through fine-grained feature alignment, using image-to-text attention (I2T Attention) and text-to-image attention (T2I Attention) to capture the most relevant parts of the image and the most relevant words in the text. In the optimization stage, a retrieval optimization module based on similarity scoring and ranking loss is designed. The OpenCLIP model is used to enhance the depth and transfer learning capabilities of feature extraction and capture the relationship between images and text from large-scale unlabeled data. Compared with traditional methods, OpenCLIP has significant advantages in semantic alignment. By pre-extracting OpenCLIP features from database images and combining dynamic scheduling of fast and slow models, the retrieval speed is improved while ensuring high accuracy. The core of this stage is to fine-tune the model parameters using data from query images or texts to ensure that the obtained model parameters can capture the unique characteristics of cross-modal data distribution, and optimize the model to maximize the similarity of related images and texts while minimizing the similarity of unrelated pairs. In this way, the model can achieve more accurate image and text matching in the feature space, improving the overall performance of cross-modal retrieval.
[0007] A cross-modal image text retrieval method based on deep learning includes a feature extraction stage, a relevance scoring stage and an optimization stage.
[0008] The feature extraction stage aims to extract deep feature representations from images and texts in order to achieve effective cross-modal semantic alignment in subsequent processing. In this stage, the text is encoded through the BERT model to obtain its rich semantic features, while Faster R-CNN and ResNet-101 are used to analyze the image and extract the visual features of the image to ensure that the features of the two modalities can be compared and aligned in a unified feature space.
[0009] The relevance scoring stage aims to evaluate and quantify the similarity between image features and text features and generate a relevance scoring matrix. In this stage, through the I2T Attention and T2I Attention mechanisms, the model can identify fine-grained associations between images and texts, capture the most relevant parts of the image and the most relevant words in the text, and calculate a similarity score for each image-text pair, providing a basis for the final retrieval results.
[0010] The optimization phase aims to fine-tune the model through similarity scoring and ranking loss to optimize the performance of cross-modal retrieval. In this phase, the model parameters are adjusted according to the data of the query image or text to ensure that the model can capture the unique characteristics of the cross-modal data distribution and maximize the similarity of related image and text pairs while minimizing the similarity of unrelated pairs. This process involves careful adjustment of the model to improve the accuracy and efficiency of retrieval, ensuring that the most relevant retrieval results can be returned quickly and accurately in practical applications.
[0011] As an improvement of the present invention, the feature extraction stage includes a text feature extraction module, an image feature extraction module and an image and text feature alignment module. The fully connected layer of ResNet-101 is modified to achieve the unification of the image and text dimensions (alignment process). The BERT model is directly used to replace the traditional i-GRU method to achieve context extraction. The BERT model is theoretically superior and can provide a deeper level of text feature representation, thereby enhancing the performance of cross-state retrieval.
[0012] The text feature extraction module aims to extract deep semantic features from the input text so as to achieve effective cross-modal semantic alignment in subsequent processing. In cross-modal data processing, there is often a lack of effective alignment mechanisms, resulting in weak correlation between image and text features. The present invention achieves fine-grained cross-modal alignment by introducing an image and text feature alignment module, so that the model can more accurately identify and match the correlation between images and texts, thereby improving the effect of cross-modal retrieval. First, the module converts the text into a format that the model can understand by loading the pre-trained BERT model and word segmenter. Specifically, the module includes two main parts: BertTokenizer and BertModel.
[0013] Among them, BertTokenizer is used to encode the input natural language text, including word segmentation, adding special tags (such as [CLS] and [SEP]), padding and other operations, so as to convert the text into an input format acceptable to the model. The calculation formula is as follows:
[0014] inputs=tokenizer(text,return tensors =pt,padding=True,truncation=True)
[0015] Among them, text represents the input text list, and inputs is the encoded PyTorch tensor containing the input text input ids 、attention mask And other information for subsequent model input.
[0016] Next, BertModel is used to receive the encoded text tensor and output the embedded representation of the text. In the forward propagation process of the model, there is no need to calculate the gradient, so torch.no_grad() is used to optimize performance. The calculation formula is as follows:
[0017] outputs = model(input ids =inputs[input ids ], attention mask =inputs[attention mask ])
[0018] Among them, input ids and attention mask They represent the input ID and attention mask of the encoded text respectively. These two variables are necessary for the model to perform self-attention calculation.
[0019] Finally, by extracting the output of the last layer of the BERT model as the embedded representation of the text, the present invention provides a flexible and powerful feature representation method. This representation not only contains rich semantic information, but also can adapt to different downstream tasks and application scenarios. Specifically, last_hidden_state represents the output of the last layer of the BERT model, which contains the hidden state of each token, and its dimension is (batch_size, seq_len, hidden_size), where batch_size represents the batch size, seq_len represents the sequence length, and hidden_size represents the dimension of the hidden layer.
[0020] The image feature extraction module aims to extract deep visual features from the input image in order to achieve effective cross-modal semantic alignment in subsequent processing. The module first performs object detection by loading the pre-trained Faster R-CNN model, identifies salient areas in the image, and extracts features of these areas. Specifically, the module consists of two main parts: Faster R-CNN and ResNet101.
[0021] Faster R-CNN is used to identify objects in an image and provide bounding boxes and confidence scores. The calculation formula is as follows:
[0022] detections=faster rcnn(images)
[0023] Among them, images represents the preprocessed image tensor, and detections contains the bounding boxes and confidence scores of the objects detected in each image. This step is the prerequisite for feature extraction because it determines which regions will be further analyzed. Next, in the feature extraction stage, the top k regions with the highest confidence are selected from the detection results to ensure that the most representative features are extracted. The calculation formula for this step is as follows:
[0024] k = min(k,len(boxes))
[0025]
[0026]
[0027] in, represents the index corresponding to the top k highest confidences, Represents the bounding boxes corresponding to these indices. This step ensures that the model only focuses on the most relevant parts of the image, thereby improving the efficiency and accuracy of feature extraction.
[0028] For each selected bounding box, the corresponding image portion is cropped and features are extracted using the ResNet101 model. The specific operations are as follows:
[0029] image parts = images[:,:,y min :y max ,x min :x max ]
[0030] Here, image parts Represents the portion of the image cropped according to the bounding box coordinates. Next, convolution features are extracted through ResNet101:
[0031] pooled feature =resnet101(image parts )
[0032] Among them, pooled feature represents the feature vector output from the ResNet101 model, typically with dimensions
[0033] (d), where d is the dimension of the feature (e.g., 768). This step extracts the deep visual features of the image and provides rich information for subsequent cross-modal semantic alignment.
[0034] Finally, all extracted features are stacked and transposed for easier subsequent processing:
[0035] features = torch.stack(conv features )).transpose(0,1)
[0036] Among them, conv features It is a list that stores all extracted features. features is the final output feature tensor with a shape of (k, d), which represents the features of the top k significant regions.
[0037] As an improvement of this experiment, the relevance scoring stage aims to quantify the similarity between image features and text features, and then generate a relevance scoring matrix to achieve fine-grained cross-modal alignment. This stage directly inherits the output of the feature extraction stage, namely the deep text semantic features and image visual features, as well as the preliminary alignment results between them. The relevance scoring stage is implemented through the following four sub-stages: pre-allocated attention module, calculation relevance score module, extraction module of shared features between text and image, and calculation relevance module.
[0038] The pre-allocated attention module calculates the cosine similarity between the image region features and the text features, and strengthens this similarity score by the amplification factor α to more clearly distinguish the similarity between different features. The calculation formula is as follows:
[0039] similarities ij =α·F.cosinesimilarity(imageregions i ,textwords j , dim=0)
[0040] Among them, imageregions i and textwords j Represent the embedding vectors of the i-th image region and the j-th text feature, respectively, and α is the preset magnification factor. The purpose of this step is to quickly screen out potential related pairs from a large number of possible image-text pairs, laying the foundation for subsequent fine-grained alignment. The Monte Carlo sampling method is used to deal with uncertainty, and the Gaussian distribution representation of the expected embedding matrix and the covariance embedding matrix is used to retain uncertainty.
[0041] The calculation relevance score module calculates a more detailed relevance score based on the pre-assigned attention score by using a weighted difference feature relevance measurement method. The calculation formula for this step is as follows:
[0042]
[0043] Among them, F scores,ij represents the relevance score between the i-th image region and the j-th text feature, and seq_len is the sequence length. This step calculates a comprehensive score to measure the relevance between each image region and all text features by considering the similarity difference between them. This calculation method can capture the subtle differences between image regions and text features, thereby improving the accuracy of the relevance score.
[0044] The module for extracting shared features between text and image extracts shared features between text and image according to the relevance score matrix F_scores and the text feature tensor text_words. The calculation formula is as follows:
[0045]
[0046] Among them, text_share_semantics_i represents the feature vector shared by the i-th image region and the text. The purpose of this step is to integrate the text features most relevant to each image region through weighted summation to form a new feature vector that can represent the shared information between the image and the text. This method can make full use of the correlation between text and image and extract richer cross-modal features.
[0047] The relevance calculation module calculates the similarity between the shared semantic features of the text and the image region features to obtain the final relevance score. The calculation formula is as follows:
[0048]
[0049] Here, relevance i represents the average relevance score between the ith text feature and all image regions, and box_num is the total number of image regions. This step calculates the similarity between the shared semantic features of the text and each image region to obtain a comprehensive score to measure the overall relevance between the text and the image. This calculation method can comprehensively evaluate the relevance between text and images and provide accurate relevance scores for cross-modal retrieval.
[0050] The optimization phase aims to minimize the difference between image and text features by fine-tuning the model parameters, thereby optimizing the performance of the cross-modal retrieval system. After the feature extraction and relevance scoring phases, the optimization phase, as a key step of the present invention, guides model learning by defining and calculating a loss function to ensure that the model can accurately capture and utilize the complex relationship between images and text. Specifically, the optimization phase includes the following key steps:
[0051] Definition and calculation of loss function:
[0052] In the optimization stage, as an improvement of the present invention, the present invention further adopts the OpenCLIP model to enhance the depth and accuracy of feature extraction. The model can convert images or texts into embedded vectors, which are then used to calculate the similarity between images and texts, providing a powerful method to evaluate the correlation between images and texts. In addition, the present invention also introduces a fast and slow model strategy, where the fast model is first used to quickly screen out potential related images or texts, while the slow model is used to fine-tune and optimize model parameters to improve the accuracy and relevance of retrieval. This strategy allows the system to improve the accuracy and relevance of retrieval while maintaining the retrieval speed. We first define a loss function class Loss, which inherits from PyTorch's nn.Module. The design of the loss function is based on the Triplet Loss principle, which guides model learning by calculating the distance difference between positive samples (related image-text pairs) and negative samples (irrelevant image-text pairs). The calculation formula is as follows:
[0053]
[0054] Among them, sin(a i ,p i ) represents the anchor point a i and positive sample p i The similarity between them, sin(a i ,n i ) represents the anchor point a i and negative samples n i The similarity between them, margin is a predefined boundary value used to control the distance between positive and negative samples.
[0055] Loss calculation and back propagation:
[0056] After constructing the similarity matrix in the relevance scoring stage, the loss is calculated using the compute_loss method. First, extract the elements on the diagonal (i.e., the similarity between each text and its corresponding image), then find the maximum value of each row and column (excluding the diagonal elements), and calculate the loss. The calculation formula is as follows:
[0057]
[0058] here, Represents the diagonal elements of the i-th row, row_max i and col_max i denote the maximum values of the i-th row and i-th column, respectively. In this way, we ensure that the similarity of each text with its corresponding image is at least a preset boundary value higher than the similarity with other images.
[0059] Finally, we backpropagate the loss function to update the model parameters. This step is achieved by calling PyTorch's backward method, which adjusts the model parameters according to the gradient of the loss function to minimize the value of the loss function.
[0060] In addition, the present invention introduces a fast and slow model strategy: the fast model quickly screens potential related images or texts, and the slow model further optimizes the matching details to improve the retrieval accuracy and relevance. Compared with the existing technology, this method significantly improves the accuracy and efficiency of cross-modal retrieval and provides a new technical path for the semantic matching of images and texts. The system can find the most relevant image based on the text, or vice versa, and has a wide range of application potential.
[0061] Beneficial effects: The present invention proposes a novel multi-stage processing framework for the problem of feature alignment and optimization in the field of cross-modal retrieval of images and texts. Through deep feature extraction, fine-grained relevance scoring and precise optimization strategies, the accuracy and efficiency of cross-modal retrieval are significantly improved. In order to obtain stable and efficient feature representation, the present invention develops a multimodal feature extractor integrating BERT and Faster R-CNN, and uses Gaussian embedding to fuse and align multi-type features. This design ensures the stability and consistency of cross-modal feature representation and overcomes the problem of insufficient feature expression ability of traditional methods. The OpenCLIP framework is introduced, and a contrastive learning strategy is adopted to capture the deep semantic relationship between images and natural languages from large-scale unlabeled data. This model not only has powerful transfer learning capabilities, but also can efficiently cope with the bidirectional retrieval requirements from text to image and image to text. Combining the fast model of OpenCLIP with the self-built slow model, the OpenCLIP features of all images in the database are pre-extracted at the same time, which greatly reduces the delay of online retrieval and significantly improves the real-time performance.
[0062] The present invention has conducted a large number of experiments on real-world image and text datasets. By comparing standard information retrieval evaluation indicators, such as R@K values, in the text-to-image retrieval task, the model of the present invention achieved R@1 of 13.00, R@5 of 40.80, and R@10 of 54.80. In the image-to-text retrieval task, the model of the present invention achieved R@1 of 55.00, R@5 of 75.00, and R@10 of 83.00. These results reflect the depth and accuracy of the present invention in understanding the relevance of cross-modal data. Compared with existing methods, such as CAMP
[222] and KCRw / o KGAt methods, the present invention improves R@1 of text-to-image retrieval by about 200% and 190%, R@5 by about 166% and 98%, and R@10 by about 98% and 62%, respectively. The R@1 of image-to-text retrieval is improved by about 98% and 39%, the R@5 is improved by about 48% and 41%, and the R@10 is improved by about 48% and 22%. The advanced retrieval performance of the present invention can be used for real-time prediction and meet the needs of large-scale multimodal data retrieval. The retrieval results are used for personalized content push in recommendation systems and relevance ranking in search engines, improving user experience and data retrieval accuracy.
[0063] In addition, the present invention not only far exceeds traditional methods in performance, but also demonstrates excellent generalization ability and adaptability, and is particularly suitable for multimodal data-intensive fields such as e-commerce, media, and education. In these scenarios, it is possible to efficiently retrieve the information most relevant to user needs from massive data, greatly improving information acquisition efficiency and user satisfaction. Through the technology of the present invention, image and text data can be effectively embedded in the same feature space to achieve deep semantic alignment, which is of great theoretical and practical significance for promoting the development of the field of artificial intelligence, especially in multimodal learning and cross-modal retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 This is a detailed framework diagram of a cross-modal retrieval system for images and texts according to the present invention. DETAILED DESCRIPTION
[0065] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0066] Embodiment 1: A cross-modal image text retrieval method based on deep learning includes a feature extraction stage, a relevance scoring stage and an optimization stage.
[0067] The feature extraction stage aims to extract deep feature representations from images and texts in order to achieve effective cross-modal semantic alignment in subsequent processing. In this stage, the text is encoded through the BERT model to obtain its rich semantic features, while Faster R-CNN and ResNet-101 are used to analyze the image and extract the visual features of the image to ensure that the features of the two modalities can be compared and aligned in a unified feature space.
[0068] The relevance scoring stage aims to evaluate and quantify the similarity between image features and text features and generate a relevance scoring matrix. In this stage, through the I2T Attention and T2I Attention mechanisms, the model can identify fine-grained associations between images and texts, capture the most relevant parts of the image and the most relevant words in the text, and calculate a similarity score for each image-text pair, providing a basis for the final retrieval results.
[0069] The optimization phase aims to fine-tune the model through similarity scoring and ranking loss to optimize the performance of cross-modal retrieval. In this phase, the model parameters are adjusted according to the data of the query image or text to ensure that the model can capture the unique characteristics of the cross-modal data distribution and maximize the similarity of related image and text pairs while minimizing the similarity of unrelated pairs. This process involves careful adjustment of the model to improve the accuracy and efficiency of retrieval, ensuring that the most relevant retrieval results can be returned quickly and accurately in practical applications.
[0070] As an improvement of the present invention, the feature extraction stage includes a text feature extraction module, an image feature extraction module and an image and text feature alignment module.
[0071] The text feature extraction module aims to extract deep semantic features from the input text so as to achieve effective cross-modal semantic alignment in subsequent processing. In cross-modal data processing, there is often a lack of effective alignment mechanisms, resulting in weak correlation between image and text features. By introducing an image and text feature alignment module, the present invention achieves fine-grained cross-modal alignment, so that the model can more accurately identify and match the correlation between images and texts, thereby improving the effect of cross-modal retrieval. First, the module converts the text into a format that the model can understand by loading the pre-trained BERT model and word segmenter. Specifically, the module includes two main parts: BertTokenizer and BertModel.
[0072] Among them, BertTokenizer is used to encode the input natural language text, including word segmentation, adding special tags (such as [CLS] and [SEP]), padding and other operations, so as to convert the text into an input format acceptable to the model. The calculation formula is as follows:
[0073] inputs=tokenizer(text,return tensors =pt,padding=True,truncation=True)
[0074] Among them, text represents the input text list, and inputs is the encoded PyTorch tensor containing the input text input ids 、attention mask And other information for subsequent model input.
[0075] Next, BertModel is used to receive the encoded text tensor and output the embedded representation of the text. In the forward propagation process of the model, there is no need to calculate the gradient, so torch.no_grad() is used to optimize performance. The calculation formula is as follows:
[0076] outputs = model(input ids =inputs[input ids ], attention mask =inputs[attention mask ])
[0077] Among them, input ids and attention mask They represent the input ID and attention mask of the encoded text respectively. These two variables are necessary for the model to perform self-attention calculation.
[0078] Finally, by extracting the output of the last layer of the BERT model as the embedded representation of the text, the present invention provides a flexible and powerful feature representation method. This representation not only contains rich semantic information, but also can adapt to different downstream tasks and application scenarios. Specifically, last_hidden_state represents the output of the last layer of the BERT model, which contains the hidden state of each token, and its dimension is (batch_size, seq_len, hidden_size), where batch_size represents the batch size, seq_len represents the sequence length, and hidden_size represents the dimension of the hidden layer.
[0079] The image feature extraction module aims to extract deep visual features from the input image in order to achieve effective cross-modal semantic alignment in subsequent processing. The module first performs object detection by loading the pre-trained Faster R-CNN model, identifies salient areas in the image, and extracts features of these areas. Specifically, the module consists of two main parts: Faster R-CNN and ResNet101.
[0080] Faster R-CNN is used to identify objects in an image and provide bounding boxes and confidence scores. The calculation formula is as follows:
[0081] detections=faster rcnn(images)
[0082] Among them, images represents the preprocessed image tensor, and detections contains the bounding boxes and confidence scores of the objects detected in each image. This step is the prerequisite for feature extraction because it determines which regions will be further analyzed. Next, in the feature extraction stage, the top k regions with the highest confidence are selected from the detection results to ensure that the most representative features are extracted. The calculation formula for this step is as follows:
[0083] k = min(k,len(boxes))
[0084]
[0085]
[0086] in, represents the index corresponding to the top k highest confidences, Represents the bounding boxes corresponding to these indices. This step ensures that the model only focuses on the most relevant parts of the image, thereby improving the efficiency and accuracy of feature extraction.
[0087] For each selected bounding box, the corresponding image portion is cropped and features are extracted using the ResNet101 model. The specific operations are as follows:
[0088] image parts = images[:,:,y min :y max ,x min :x max ]
[0089] Here, image parts Represents the portion of the image cropped according to the bounding box coordinates. Next, convolution features are extracted through ResNet101:
[0090] pooled feature =resnet101(image parts )
[0091] Among them, pooled feature Represents the feature vector output from the ResNet101 model, typically with dimension (d), where d is the dimension of the feature (e.g., 768). This step extracts the deep visual features of the image, providing rich information for subsequent cross-modal semantic alignment.
[0092] Finally, all extracted features are stacked and transposed for easier subsequent processing:
[0093] features = torch.stack(conv features )).transpose(0,1)
[0094] Among them, conv features It is a list that stores all extracted features. features is the final output feature tensor with a shape of (k, d), which represents the features of the top k significant regions.
[0095] As an improvement of this experiment, the relevance scoring stage aims to quantify the similarity between image features and text features, and then generate a relevance scoring matrix to achieve fine-grained cross-modal alignment. This stage directly inherits the output of the feature extraction stage, namely the deep text semantic features and image visual features, as well as the preliminary alignment results between them. The relevance scoring stage is implemented through the following four sub-stages: pre-allocated attention module, pre-allocated attention module, extraction of text and image shared features module, and calculation of relevance module.
[0096] The pre-allocated attention module calculates the cosine similarity between the image region features and the text features, and strengthens this similarity score by the amplification factor α to more clearly distinguish the similarity between different features. The calculation formula is as follows:
[0097] similarities ij =α·F.cosinesimilarity(imageregions i ,textwords j , dim=0)
[0098] Among them, imageregions i and textwords jRepresent the embedding vectors of the i-th image region and the j-th text feature, respectively, and α is the preset magnification factor. The purpose of this step is to quickly screen out potential related pairs from a large number of possible image-text pairs, laying the foundation for subsequent fine-grained alignment. The Monte Carlo sampling method is used to deal with uncertainty, and the Gaussian distribution representation of the expected embedding matrix and the covariance embedding matrix is used to retain uncertainty.
[0099] The calculation relevance score module calculates a more detailed relevance score based on the pre-assigned attention score by using a weighted difference feature relevance measurement method. The calculation formula for this step is as follows:
[0100]
[0101] Among them, F scores,ij represents the relevance score between the i-th image region and the j-th text feature, and seq_len is the sequence length. This step calculates a comprehensive score to measure the relevance between each image region and all text features by considering the similarity difference between them. This calculation method can capture the subtle differences between image regions and text features, thereby improving the accuracy of the relevance score.
[0102] The module for extracting shared features between text and image extracts shared features between text and image according to the relevance score matrix F_scores and the text feature tensor text_words. The calculation formula is as follows:
[0103]
[0104] Among them, text_share_semantics i Represents the feature vector shared by the i-th image region and the text. The purpose of this step is to integrate the text features most relevant to each image region by weighted summation to form a new feature vector that can represent the common information between the image and the text. This method can make full use of the correlation between text and image and extract richer cross-modal features.
[0105] The relevance calculation module calculates the similarity between the shared semantic features of the text and the image region features to obtain the final relevance score. The calculation formula is as follows:
[0106]
[0107] Here, relevance irepresents the average relevance score between the ith text feature and all image regions, and box_num is the total number of image regions. This step calculates the similarity between the shared semantic features of the text and each image region to obtain a comprehensive score to measure the overall relevance between the text and the image. This calculation method can comprehensively evaluate the relevance between text and images and provide accurate relevance scores for cross-modal retrieval.
[0108] The optimization phase aims to minimize the difference between image and text features by fine-tuning the model parameters, thereby optimizing the performance of the cross-modal retrieval system. After the feature extraction and relevance scoring phases, the optimization phase, as a key step of the present invention, guides model learning by defining and calculating a loss function to ensure that the model can accurately capture and utilize the complex relationship between images and text. Specifically, the optimization phase includes the following key steps:
[0109] Definition and calculation of loss function:
[0110] In the optimization stage, we first define a loss function class Loss, which inherits from PyTorch's nn.Module. The design of the loss function is based on the Triplet Loss principle, which guides model learning by calculating the distance difference between positive samples (related image-text pairs) and negative samples (irrelevant image-text pairs). The calculation formula is as follows:
[0111]
[0112] Among them, sin(a i ,p i ) represents anchor a i and positive sample p i The similarity between them, sim(a_i,n_i) represents the anchor point a i and negative samples n i The similarity between them, margin is a predefined boundary value used to control the distance between positive and negative samples.
[0113] Loss calculation and back propagation:
[0114] After constructing the similarity matrix in the relevance scoring stage, the loss is calculated using the compute_loss method. First, extract the elements on the diagonal (i.e., the similarity between each text and its corresponding image), then find the maximum value of each row and column (excluding the diagonal elements), and calculate the loss. The calculation formula is as follows:
[0115]
[0116] here, Represents the diagonal elements of row ii, row_maxi and col_max i denote the maximum values of the i-th row and i-th column, respectively. In this way, we ensure that the similarity of each text with its corresponding image is at least a preset boundary value higher than the similarity with other images.
[0117] Finally, we backpropagate the loss function to update the model parameters. This step is achieved by calling PyTorch's backward method, which adjusts the model parameters according to the gradient of the loss function to minimize the value of the loss function.
[0118] Embodiment 2: Learning invariant feature representation in a cross-modal retrieval system of images and texts includes the following steps:
[0119] Module 1: Feature Extraction. Encode the text, use the BERT model and word segmenter to convert the text into a format that the model can understand, and output the text embedding vector through the BERT encoder; perform object detection and feature extraction on the image, using the Faster R-CNN and ResNet101 models.
[0120] Module 2: Relevance Scoring. Calculate the cosine similarity between image region features and text features, and use the pre-assigned attention function and softmax normalization to obtain the attention scoring matrix; calculate the relevance score based on the attention scoring matrix, and measure the relevance between image region and text features through the feature relevance metric of weighted difference.
[0121] Module 3: Optimization. Define the loss function, based on the Triplet Loss principle, and guide model learning by calculating the distance difference between positive and negative samples; use the OpenCLIP model to further extract deep features of images and texts; use a fast and slow model strategy, with the fast model used to quickly screen potential related pairs and the slow model used to fine-tune and optimize model parameters.
[0122] This case study studies multi-task feature learning for cross-modal retrieval of images and texts, including text-level, image-level, and cross-modal-level tasks, and proposes a multimodal feature alignment system consisting of three stages. 1) In the feature extraction stage, deep visual and semantic features are extracted from images and texts, and effective cross-modal semantic alignment is achieved in subsequent processing. 2) In the relevance scoring stage, a relevance scoring matrix is generated by calculating the similarity scores between image features and text features to achieve fine-grained cross-modal alignment. 3) In the optimization stage, it aims to optimize the performance of the cross-modal retrieval system by minimizing the difference between image and text features, and learns model parameters shared across multiple datasets through adaptive parameter updates.
Claims
1. A cross-modal image text retrieval method based on deep learning, characterized by: The method comprises the following steps: Step 1: Feature extraction stage: extract deep features from images and texts to ensure semantic alignment of cross-modal data in a unified feature space; Step 2: Relevance scoring stage, the relevance between image and text is evaluated based on a fine-grained attention mechanism; Step 3: Optimization phase, optimize the retrieval performance through specific loss functions and model strategies. The feature extraction stage aims to extract deep feature representations from images and texts so as to achieve effective cross-modal semantic alignment in subsequent processing. The text is encoded through the BERT model to obtain its rich semantic features. At the same time, Faster R-CNN and ResNet-101 are used to analyze the image and extract the visual features of the image to ensure that the features of the two modalities can be compared and aligned in a unified feature space. The relevance scoring stage aims to evaluate and quantify the similarity between image features and text features and generate a relevance scoring matrix. Through the I2T Attention and T2I Attention mechanisms, the model can identify fine-grained associations between images and texts, capture the most relevant parts of the image and the most relevant words in the text, and calculate a similarity score for each image-text pair, providing a basis for the final retrieval results. The optimization stage aims to fine-tune the model through similarity scoring and ranking loss to optimize the performance of cross-modal retrieval. The model parameters are adjusted according to the data of the query image or text to ensure that the model can capture the unique characteristics of the cross-modal data distribution and maximize the similarity of related image and text pairs while minimizing the similarity of unrelated pairs; The optimization stage includes the definition and calculation of the loss function, the calculation and back propagation of the loss; The definition and calculation of the loss function is based on the Triplet Loss principle, which guides model learning by calculating the distance difference between positive samples and negative samples; The loss calculation and back propagation are performed by using the compute_loss method to calculate the loss L. The calculation formula of L is as follows: in, Represents the diagonal elements of the i-th row, row_max i and col_max i Represents the maximum value of the i-th row and i-th column respectively, and margin is a predefined boundary value used to control the distance between positive and negative samples; First, extract the elements on the diagonal, then find the maximum value of each row and column, calculate the loss, and finally backpropagate the loss function to update the model parameters; The optimization stage also includes a fast and slow model strategy, where the fast model is used to quickly screen out potential related images or texts, and the slow model is used to fine-tune and optimize model parameters to improve the accuracy and relevance of retrieval.
2. The cross-modal image text retrieval method based on deep learning according to claim 1, characterized in that: The feature extraction stage includes a text feature extraction module and an image feature extraction module; The text feature extraction module is designed to extract deep semantic features from the input text so as to achieve effective cross-modal semantic alignment in subsequent processing. It converts the text into a format that the model can understand by loading the pre-trained BERT model and word segmenter. It includes BertTokenizer and BertModel. BertTokenizer is used to encode the input natural language text, and BertModel is used to receive the encoded text tensor and output the embedded representation of the text. The image feature extraction module is designed to extract deep visual features from the input image in order to achieve effective cross-modal semantic alignment in subsequent processing. It performs object detection by loading the pre-trained Faster R-CNN model, identifies salient areas in the image, and extracts features of these areas. It includes Faster R-CNN and ResNet101. Faster R-CNN is used to identify objects in the image and provide bounding boxes and confidence scores, and ResNet101 is used to extract convolutional features.
3. The cross-modal image text retrieval method based on deep learning as claimed in claim 2, characterized in that: The text feature extraction module includes two main parts: BertTokenizer and BertModel. Among them, BertTokenizer is used to encode the input natural language text, including word segmentation, adding special tags, and padding operations, so as to convert the text into an input format acceptable to the model. The calculation formula is as follows: inputs=tokenizer(text,return tensors =pt,padding=True,truncation=True) where inputs is the encoded PyTorch tensor, which contains the input_ids and attention_mask information of the input text, and is used for subsequent model input. Tokenizer represents the tokenizer object used for text processing, which is part of the BERT model and is used to convert natural language text into a form that the model can understand. Text represents the input text list. Return_tensors indicates whether the tokenizer returns a PyTorch tensor. Padding is set to True to pad all text sequences to make them of the same length to meet the model input requirements. Next, BertModel is used to receive the encoded text tensor and output the embedded representation of the text. In the forward propagation process of the model, there is no need to calculate the gradient, so torch.no_grad() is used to optimize the performance. The calculation formula is as follows: outputs = model(input ids =inputs[input ids ],attention mask =inputs[attention mask ]) Where model represents an instance of the BERT model, which is responsible for receiving the processed text tensor and outputting the embedded representation of the text, and input ids and attention mask They represent the input ID and attention mask of the encoded text respectively, and outputs represents the output of the model. Finally, by extracting the output of the last layer of the BERT model as the embedded representation of the text, last_hidden_state represents the output of the last layer of the BERT model, which contains the hidden state of each token. Its dimension is (batch_size, seq_len, hidden_size), where batch_size represents the batch size, seq_len represents the sequence length, and hidden_size represents the dimension of the hidden layer. The image feature extraction module consists of two main parts: Faster R-CNN and ResNet101. Faster R-CNN is used to identify objects in an image and provide bounding boxes and confidence scores. The calculation formula is as follows: detections=faster rcnn(images) Among them, images represents the preprocessed image tensor, detections contains the bounding boxes and confidence scores of the objects detected in each image. In the feature extraction stage, the top k regions with the highest confidence are selected from the detection results to ensure that the most representative features are extracted. The calculation formula is as follows: k = min(k,len(boxes)) Among them, k represents the number of selected image regions, which is used in the feature extraction stage to ensure the extraction of the most representative features. Boxes represents the list of bounding boxes of objects detected from the image. Each bounding box is a tensor containing the location information of the object. Len(boxes) represents the length of the bounding box list, that is, the number of objects detected in the image. represents the index corresponding to the top k highest confidences, represents the bounding boxes corresponding to these indices, For each selected bounding box, the corresponding image portion is cropped and features are extracted using the ResNet101 model as follows: image parts =images[:,:,y min :y max ,x min :x max ] Here, image parts Represents the portion of the image cropped according to the bounding box coordinates. Next, convolution features are extracted through ResNet101: pooled feature =resnet101(image parts ) Among them, pooled feature represents the feature vector output from the ResNet101 model, typically with dimensions d, where d is the dimension of the feature, Finally, all extracted features are stacked and transposed for easier subsequent processing: features=torch.stack(conv features ).transpose(0,1) Among them, conv features It is a list that stores all extracted features. features is the final output feature tensor with a shape of (k, d), which represents the features of the top k significant regions.
4. The cross-modal image text retrieval method based on deep learning as claimed in claim 2, characterized in that: The relevance scoring stage includes a pre-allocation attention module, a relevance scoring module, a text and image shared feature extraction module, and a relevance calculation module. The pre-allocated attention module calculates the cosine similarity between the image region features and the text features, and strengthens the similarity score by the amplification factor α to more clearly distinguish the similarity between different features. The calculation formula is as follows: similarities ij =α·F.cosine similarity(image regions i ,text words j ,dim=0) Among them, similarities ij represents the similarity score between the i-th image region and the j-th text feature. This score is used to measure the proximity between the image and the text in the feature space and is the key output of the relevance scoring stage. imageregions i and text words j Represent the embedding vectors of the i-th image region and the j-th text feature respectively, α is the preset magnification factor, F.cosine similarity is the cosine similarity calculation function in the PyTorch library, which is used to calculate the cosine similarity between two vectors, dim=0 specifies that the cosine similarity is calculated along the 0th dimension, The module for calculating the relevance score calculates a more detailed relevance score based on the pre-assigned attention score by using a weighted difference feature relevance measurement method. The calculation formula is as follows: Among them, F scores,ij represents the correlation score between the i-th image region and the j-th text feature, sql len is the sequence length, i.e. the length of the text feature vector, represents the attention score between the i-th image region and the j-th text feature, represents the attention score between the i-th image region and the t-th text feature, The module for extracting shared features between text and image extracts shared features between text and image according to the relevance score matrix F_scores and the text feature tensor text_words. The calculation formula is as follows: Among them, text_share_semantics i Represents the feature vector shared by the i-th image region and the text, sql len represents the length of the text sequence, that is, the length of the text feature vector, F scores,ij Represents the correlation score between the i-th image region and the j-th text feature, which is used to measure the correlation between the image region and the text feature. text_words j The embedding vector representing the jth text feature contains the deep semantic information of the text. The purpose of this step is to integrate the most relevant text features of each image region by weighted summation to form a new feature vector that can represent the common information of the image and text. The relevance calculation module calculates the similarity between the shared semantic features of the text and the image region features to obtain the final relevance score. The calculation formula is as follows: Among them, relevance i represents the average relevance score between the i-th text feature and all image regions, box_num is the total number of image regions, F.cosine_similarity is the cosine similarity calculation function in the PyTorch library, which is used to calculate the cosine similarity between two feature vectors, text_share_semantics j Indicates the feature vector shared by the jth image region and the text. This vector integrates the text features most relevant to the image region by weighted summation and is used to represent the shared information between the image and the text. image_regions i Represents the feature vector of the i-th image region. It is a high-dimensional tensor that contains the deep visual information of the image region. This step calculates the similarity between the shared semantic features of the text and each image region to obtain a comprehensive score to measure the overall relevance between the text and the image.
5. The cross-modal image text retrieval method based on deep learning as claimed in claim 4, characterized in that: In the definition and calculation of the loss function, a loss function class Loss is first defined, which is inherited from PyTorch's nn.Module. The design of the loss function is based on the Triplet Loss principle, which guides model learning by calculating the distance difference between positive samples and negative samples. The calculation formula is as follows: Among them, sin(a i ,p i ) represents anchor a i and positive sample p i The similarity between them, sin(a i ,n i ) represents the anchor point a i and negative samples n i The similarity between Loss calculation and back propagation: After constructing the similarity matrix in the relevance scoring stage, the loss is calculated through the compute_loss method. First, the elements on the diagonal are extracted, that is, the similarity between each text and its corresponding image. Then, the maximum value of each row and column is found, the diagonal elements are excluded, and the loss is calculated.
6. The cross-modal image text retrieval method based on deep learning according to claim 5, characterized in that: The optimization stage further uses an OpenCLIP model to enhance the depth and accuracy of feature extraction, which is capable of converting images or text into embedding vectors, which are then used to calculate the similarity between images and text.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the cross-modal image text retrieval method based on deep learning as described in any one of claims 1 to 5 above.
8. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instruction is executed by a processor, the cross-modal image text retrieval method based on deep learning as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Video natural language text retrieval method based on space time sequence characteristics
CN113704546A
Visual question answering method and apparatus, electronic device and storage medium
WO2024164616A1
Cited By
Multi-modal data fusion modeling method and system based on multi-task learning
CN120744812A