Image Recipe Retrieval Method Based on Multi-Level Features and Attention Mechanism
By constructing an image recipe search model based on multi-level features and attention mechanisms, optimizing image features and combining triple loss, the problem of insufficient image representation in cross-modal retrieval is solved, and more efficient image and recipe matching is achieved.
Patent Information
- Application Number
- CN202310301992.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-24
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-03-24
AI Technical Summary
The prior art is difficult to effectively optimize picture representation in cross-modal retrieval, and the distribution distance between pictures and text is large, resulting in limited retrieval performance.
The image recipe search method based on multi-level features and attention mechanism is adopted. By constructing an image recipe search model, the channel and spatial attention mechanism are used to optimize image features, and combining context learning modules and triple loss, the matching accuracy of images and recipes is improved.
It significantly improves the performance of cross-modal retrieval, can better match images and recipes, and improves the accuracy and interpretability of the retrieval.
Smart Images

Figure CN116361497B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cross-modal retrieval, and in particular to an image recipe retrieval method based on multi-level features and attention mechanism. Background Art
[0002] In recent years, with the booming development of the Internet, the food industry has become one of the popular categories on social platforms, and a large number of high-quality blogs and popular notes have emerged in this industry. Therefore, there is a large amount of heterogeneous data on social platforms, namely food images and food preparation processes. For users, single-modal retrieval can no longer meet their needs, so cross-modal retrieval has emerged.
[0003] Cross-modal retrieval aims to retrieve relevant data of another modality with data of one modality. The related tasks of cross-modal retrieval are data feature extraction and relevance measurement of content between different modalities.
[0004] Food is an inseparable part of people's daily life, and food computing is also a research hotspot. Image recipe retrieval is a typical cross-modal retrieval problem and an important branch of food computing. Cross-modal retrieval is to learn a common embedding subspace and use text queries to retrieve relevant images, and vice versa. To represent the two modalities of text and pictures, previous studies encoded pictures through residual neural networks and encoded recipes through gated recurrent units or long short-term memory networks. To obtain an excellent joint embedding space, cosine loss, triplet loss, and triplet loss with hard example mining were used to train the model. However, there are the following problems:
[0005] (1) Difficulty in obtaining distinguishable picture representations;
[0006] (2) A large distance between the distributions of pictures and texts.
[0007] Picture representations and recipe representations, as the inputs for picture and recipe retrieval, are the basis of retrieval performance. Many previous methods have improved text representations, and these methods have greatly improved retrieval performance. However, they did not consider optimizing picture representations. Convolutional feature channels usually correspond to specific functions, such as texture extraction or intensity detection. Existing methods do not consider the importance of different features and do not use channel attention to flexibly adjust features. In food pictures, only the areas where food appears are the areas we need to focus on. Also, in recipes, some ingredients are not visible, and we need to pay more attention to the areas where ingredients can appear in the pictures and the corresponding text content. Summary of the Invention
[0008] The object of the present invention is to solve the above-mentioned existing technical problems.
[0009] To achieve the above object, the present invention provides an image recipe retrieval method based on multi-level features and an attention mechanism, and the specific steps are as follows:
[0010] Step S1: Collect food image data and recipe data;
[0011] Step S2: Construct an image recipe retrieval model based on multi-level features and a context-aware attention mechanism;
[0012] Step S3: Train the image recipe retrieval model in Step S2 with the food image data and recipe data in Step S1;
[0013] Step S4: Perform cross-modal retrieval on food images and recipes through the trained image recipe retrieval model.
[0014] Preferably, the recipe data includes the formula and the cooking method data of the food.
[0015] Preferably, Step S2 is specifically as follows:
[0016] Step S21: Extract the initial image features of the food image, and at the same time extract the initial text features of the formula and the cooking method;
[0017] Step S22: Pass the extracted initial image features through a channel attention mechanism and a spatial attention mechanism to obtain regional picture features, thereby obtaining a set of regional picture features;
[0018] Step S23: Use a self-attention mechanism for the initial text features of the formula and the cooking method to obtain word feature representations with higher importance than the initial text features, thereby obtaining a set of formula features and a set of cooking method features;
[0019] Use the obtained set of regional picture features and the combined feature set of the formula features and the cooking method features as the input of the context learning module to obtain the fine-grained relationship between the regional pictures and the words, thereby obtaining the context attention loss;
[0020] Step S25: Overlap the regional picture features and the initial image features obtained in Step S21 and input them into the first FC layer to obtain the first image feature representation V, splice the formula and cooking method features in the set of formula features and the set of cooking method features together and input them into the first FC layer to obtain the first recipe representation R;
[0021] Step S26: Calculate the triplet loss and the maximum mean error through the first image feature representation V and the first recipe representation R;
[0022] Step S27: The first image feature representation V and the first recipe representation R are respectively passed through their respective second FC layers to obtain the second image feature representation V' and the second recipe representation R'. The translation consistency loss is calculated using the second image feature representation V' and the second recipe representation R'.
[0023] Step S28: The total loss of the image-recipe retrieval model includes the context attention loss, the triplet loss, the maximum mean error, and the translation consistency loss.
[0024] Preferably, step S21 is specifically as follows:
[0025] Step S211: Use the ResNet50 network pre-trained on ImageNet with the last avgpool and fc layers to extract features from the food image to obtain the initial image features.
[0026] Step S212: For each recipe, first use the word2vec model to obtain a feature representation, and then use Bi-LSTM to obtain the initial text features of the recipe.
[0027] Step S213: Use a two-stage LSTM model to represent the cooking method. That is, first, each sentence in the cooking method is represented as a skip-instructions vector, and then the sequence of skip-instructions vectors is trained by LSTM to extract the initial text features of the cooking method.
[0028] Preferably, step S22 is specifically as follows:
[0029] Step S221: Use average pooling and max pooling to aggregate spatial information to generate two different spatial context description features. First, compress the spatial context description features, then restore the spatial context description features, and use element-wise summation to merge the output feature vectors.
[0030] Step S222: Connect the two spatial context description features along the channel axis to obtain the spatial attention map by convolving and fusing the feature maps with a standard convolutional layer.
[0031] Step S223: The spatial attention map passes through the channel attention module and the spatial attention module to obtain the optimized regional image feature H.
[0032] Preferably, step S24 is specifically as follows:
[0033] Step S241: Pay attention to the words related to each image region in each sentence, and calculate the cosine similarity matrix for each region-word pair in the image-recipe pair. The formula for the cosine similarity matrix is as follows:
[0034]
[0035] where h i is an element in the set of regional image features, and e j is an element in the set of recipe features and the set of cooking method features. And when j ∈ [1, m], e j is an element in the set of recipe features, and when j ∈ [m + 1, m + n], e j is an element in the set of cooking method features;
[0036] Step S242: Calculate the weighted word representation through the cosine similarity matrix, and the calculation formula is as follows:
[0037]
[0038]
[0039] where α ij is the attention weight, is the sentence vector relative to the i-th food image region;
[0040] Step S243: Given the sentence context, determine the importance of each image region by calculating the cosine similarity. The formula for calculating the cosine similarity is as follows:
[0041]
[0042] Step S244: Calculate the similarity between the regional image and the word through LogSumExp pooling to obtain the fine-grained relationship between the regional image and the word. The calculation formula is as follows:
[0043]
[0044] where k is the number of regions into which the picture is divided, λ is a parameter that determines the amplification importance of the most relevant pair between the food image region features and the attended sentence vector, and W is the combined feature of the recipe features and the cooking method features;
[0045] Step S245: Calculate the context attention loss according to the fine-grained relationship. The calculation formula is as follows:
[0046]
[0047] where and are negative samples, and α1 is the margin.
[0048] Preferably, in step S26, the calculation formula of the triplet loss is as follows:
[0049]
[0050] Among them, and are negative samples, α is the margin parameter, d is the Euclidean distance, [X] + = max(X, 0); The formula for calculating the maximum mean error is as follows:
[0051]
[0052] Among them, φ is the feature mapping of the canonical form φ(x) = k(x, ·), which is a reproducing Hilbert space with a Gaussian kernel k.
[0053] Preferably, in step S27, the formula for calculating the translation consistency loss is as follows:
[0054]
[0055] Among them, the possibility of food types is calculated for the image and the recipe respectively, and the formula for calculating the possibility is as follows:
[0056]
[0057]
[0058] Among them, N represents the number of food types, and the classification loss is L cls = (p img , p rec , c gt ), where c gt is the true food type label;
[0059] Minimize the KL divergence between p img and p rec , and the formula is as follows:
[0060]
[0061]
[0062] Preferably, the total loss formula of the image recipe retrieval model is as follows:
[0063]
[0064] Among them, α, β, and γ are constant parameters.
[0065] Therefore, the present invention adopts the above-mentioned image recipe retrieval method based on multi-level features and attention mechanism, and has the following beneficial effects:
[0066] (1) An image encoder based on a multi - layer feature with an attention mechanism is introduced, which focuses on the important features of food images and suppresses the responses of unnecessary regions. Its performance in cross - modal retrieval tasks is significantly better than that of similar products based on pure CNNs.
[0067] (2) A triplet loss for cross - modal retrieval is introduced. By combining the maximum mean discrepancy with the triplet loss, corresponding image - text pairs are better pulled closer, and unmatched pairs are moved farther away.
[0068] (3) A context learning module is introduced to capture the fine - grained relationships between regions in the image and words in the recipe, improving the interpretability of the model.
[0069] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings
[0070] Figure 1 It is a flowchart of an image - recipe retrieval method based on multi - level features and an attention mechanism according to the present invention;
[0071] Figure 2 It is a flowchart of step S2 of the present invention;
[0072] Figure 3 It is a block diagram of an image - recipe retrieval model according to the present invention. Detailed Embodiments
[0073] Embodiment
[0074] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention usually described and illustrated herein can be arranged and designed in various different configurations.
[0075] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0076] The embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings.
[0077] Refer to Figure 1-2 , an image - recipe retrieval method based on multi - level features and an attention mechanism, the specific steps are as follows:
[0078] Step S1: Collect food image data and recipe data. The recipe data includes the formula and the cooking method data of the food.
[0079] Step S2: Construct an image-recipe retrieval model based on a multi-level feature and context-aware attention mechanism. The block diagram of the image-recipe retrieval model is as shown in Figure 3 Figure.
[0080] Step S21: Extract the initial image features of the food image, and at the same time extract the initial text features of the formula and the cooking method.
[0081] Step S211: Use the ResNet50 network pre-trained on ImageNet and the last avgpool and fc layers to extract features from the food image to obtain the initial image features. The initial image features are layer4, and the dimension of the Layer4 features is 1024x7x7, corresponding to 49 spatial grids of the image. Each grid is represented as a 1024-dimensional vector. Represent the features of layer4 as F region ={f1,f2,...,f 49}.
[0082] Step S212: For each formula, first use the word2vec model to obtain a feature representation, and then use Bi-LSTM to obtain the initial text features of the formula.
[0083] Step S213: Use a two-stage LSTM model to represent the features of the cooking method. That is, first, each sentence in the cooking method is represented as a skip-instructions vector, and then the sequence of skip-instructions vectors is trained by LSTM to extract the initial text features of the cooking method.
[0084] Step S22: Pass the extracted initial image features through a channel attention mechanism and a spatial attention mechanism to obtain regional picture features, thereby obtaining a set of regional picture features. Step S22 is specifically as follows:
[0085] Step S221: Use average pooling and max pooling to aggregate spatial information to generate two different spatial context description features ( and ). First, use W0 to compress the features, and then restore the features through W1. and share W0 and W1. Then use element-wise summation to merge the output feature vectors,
[0086] where σ represents the sigmoid function. To reduce the parameter overhead, W0 ∈ RC / r×C ,
[0087] Step S222: To generate a valid feature description, two spatial context description features are concatenated along the channel axis ( and ) to obtain a spatial attention map by convolving and fusing the feature maps with a standard convolutional layer,
[0088] where σ represents the sigmoid function, and f 7×7 represents that the size of the filter of the standard convolutional layer is 7x7.
[0089] Step S223: The spatial attention map passes through the channel attention module and the spatial attention module to obtain the optimized regional image feature H,
[0090] Step S23: The self-attention mechanism is used for the initial text features of the recipe and cooking method to obtain word feature representations with higher importance than the initial text features, thereby obtaining the recipe feature set and the cooking method feature set.
[0091] Step S24: The obtained regional image feature set, recipe feature set, and cooking method feature set are used as the input of the context learning module to obtain the fine-grained relationship between the regional image and the words, thereby obtaining the context attention loss. Step S24 is specifically as follows:
[0092] Use the regional image feature set H region ={h1, h2,..., h 49} and the recipe feature set R ingr ={e1, e2,..., e m} and the cooking method feature set R instr ={e m+1 , e m+2 ,..., e m+n} as the input of the context attention module. The main purpose of the context attention module is to measure the similarity of an image-recipe pair.
[0093] Step S241: Focus on the words related to each image region in each sentence, and calculate the cosine similarity matrix for each region-word pair in the image-recipe pair. The calculation formula of the cosine similarity matrix is as follows:
[0094]
[0095] where s ij represents the similarity between the i-th image region and the j-th word in the recipe, h i is an element in the regional image feature set, and ej is an element in the set of recipe feature sets and the set of cooking method feature sets, and when j ∈ [1, m], e j is an element in the set of recipe feature sets, and when j ∈ [m + 1, m + n], e j is an element in the set of cooking method feature sets;
[0096] Step S242: Calculate the weighted word representation through the cosine similarity matrix, and the calculation formula is as follows:
[0097]
[0098]
[0099] where α ij is the attention weight, is the sentence vector relative to the i-th food image region.
[0100] Step S243: Given the sentence context, determine the importance of each image region by calculating the cosine similarity, and the cosine similarity calculation formula is as follows:
[0101]
[0102] Step S244: Calculate the similarity between the regional image and the word through LogSumExp pooling to obtain the fine-grained relationship between the regional picture and the word, and the calculation formula is as follows:
[0103]
[0104] where k is the number of regions into which the picture is divided, λ is a parameter that determines the amplification importance of the most relevant pair between the food image region feature and the attended sentence vector, and W is the combined feature of the recipe feature and the cooking method feature;
[0105] Step S245: Calculate the context attention loss according to the fine-grained relationship, and the calculation formula is as follows:
[0106]
[0107] where, and are negative samples, and α1 is the margin.
[0108] Step S25: Overlap the regional picture feature and the initial image feature obtained in Step S21 and input them into the first FC layer to obtain the first image feature representation V, V = Tanh(W fc (AvgPool(F)+Mean(H))+b fc) F is a set of image features obtained by removing the last two layers of ResNet, containing 49 regions.
[0109] Concatenate the recipe and cooking method features in the recipe feature set and the cooking method feature set and input them into the first FC layer to obtain the first recipe representation R. R = Tanh(W fc ([R ingr ; R instr ) + b fc ) Ringr and Rinstr are the features obtained by passing the ingredients and recipes through the attention mechanism respectively. b fc is a bias of the fully connected layer.
[0110] Step S26: Calculate the triplet loss and the maximum mean discrepancy using the first image feature representation V and the first recipe representation R.
[0111] In step S26, the formula for calculating the triplet loss is as follows:
[0112]
[0113] Where, and are negative samples, α is the margin parameter, d is the Euclidean distance, [X] + = max(X, 0);
[0114] The formula for calculating the maximum mean discrepancy is as follows:
[0115]
[0116] Where, φ is the feature map of the canonical form φ(x) = k(x, ·), which is a reproducing Hilbert space with a Gaussian kernel k.
[0117] Step S27: The first image feature representation V and the first recipe representation R are respectively passed through their respective second FC layers to obtain the second image feature representation V' and the second recipe representation R'. Calculate the translation consistency loss using the second image feature representation V' and the second recipe representation R'. The formula for calculating the translation consistency loss is as follows:
[0118]
[0119] Where, calculate the possibility of food types for images and recipes respectively. The formula for calculating the possibility is as follows:
[0120]
[0121]
[0122] Where N represents the number of food types, and the classification loss is Lcls = (p img , p rec , c gt ), where c gt is the true food category label;
[0123] Minimize the KL divergence between p img and p rec , and the formula is as follows:
[0124]
[0125]
[0126] Step S28: The total loss of the image recipe retrieval model includes context attention loss, triplet loss, maximum mean error, and translation consistency loss. The formula for the total loss of the image recipe retrieval model is as follows: L = L Tri + αL MMD + βL CAM + γL TC
[0127] where α, β, and γ are constant parameters.
[0128] Step S3: Train the image recipe retrieval model in Step S2 with the food image data and recipe data in Step S1.
[0129] Step S4: Perform cross-modal retrieval on food images and recipes using the trained image recipe retrieval model.
[0130] Use median retrieval rank (MedR) and recall at top K (R@K) as evaluation metrics. Among them, MedR measures the median rank position of the returned positive sample rankings. Therefore, when the performance is better, the value of MedR is lower. R@K calculates the number of times the correct recipe is found among the top-K retrieved candidate samples. Therefore, the higher the R@K score, the higher the performance. As the largest available structured dataset containing food image and recipe pairs so far, Recipe1M is widely used. This dataset is extracted from multiple popular cooking websites and contains 1,029,720 recipes and 887,536 pictures, where the training set accounts for 70%, and the remaining data is approximately divided into a test set and a validation set in a 1:1 ratio. This dataset has a total of 1048 categories, and each category corresponds to multiple recipes. The average number of ingredients and instructions for each recipe is 9.3 and 10.5 respectively. Therefore, the recipes are long texts and are relatively complex to process.
[0131] To test the performance of the algorithm, image-recipe pairs are randomly selected from the test set. In a subset, the corresponding recipe is retrieved through the image, and at the same time, the corresponding image is retrieved through the recipe. Also, to demonstrate the scalability of the algorithm, the subsets are set to 1K and 10K pairs. Each experiment is repeated 10 times, and the average results are presented. As shown in the following table,
[0132]
[0133] It can be seen that the present invention not only outperforms other methods on the 1k test set, but is also equally effective on the 10k test set. For the 1k set, for retrieving images from recipes, the present invention significantly improves the value of R@K by approximately 3-4%. Similarly, for the value of MedR, the present invention reduces it to 1. Compared with other methods, the present invention uses the image regions and words in the sentence as context to better align the latent space and obtain better results. When we use the 10k test set, the method of the present invention shows good robustness.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. An image recipe retrieval method based on multi-level features and attention mechanism, characterized in that, The specific steps are as follows: Step S1: Collect food image data and recipe data; the recipe data includes the formula and cooking method data of the food; Step S2: Construct an image-recipe retrieval model based on a multi-level feature and context-aware attention mechanism; Step S2 is specifically as follows: Step S21: Extract the initial image features of the food image, and at the same time extract the initial text features of the formula and cooking method; Step S22: Pass the extracted initial image features through a channel attention mechanism and a spatial attention mechanism to obtain regional picture features, thereby obtaining a set of regional picture features; Step S23: Use the self-attention mechanism for the initial text features of the formula and cooking method to obtain a word feature representation with higher importance than the initial text features, thereby obtaining a set of formula features and a set of cooking method features; Step S24: Use the obtained set of regional picture features and the combined feature set of formula features and cooking method features as the input of the context learning module to obtain the fine-grained relationship between the regional picture and the words, thereby obtaining the context attention loss; Step S25: Overlap the regional picture features and the initial image features obtained in Step S21 and input them into the first FC layer to obtain the first image feature representation V, splice the formula and cooking method features in the set of formula features and the set of cooking method features together and input them into the first FC layer to obtain the first recipe representation R; Step S26: Calculate the triplet loss and the maximum mean error through the first image feature representation V and the first recipe representation R; Step S27: The first image feature representation V and the first recipe representation R respectively pass through their respective second FC layers to obtain the second image feature representation and the second recipe representation , and calculate the calculation translation consistency loss through the second image feature representation and the second recipe representation ; Step S28: The total loss of the image-recipe retrieval model includes the context attention loss, the triplet loss, the maximum mean error, and the translation consistency loss; Step S3: Train the image-recipe retrieval model in Step S2 with the food image data and recipe data in Step S1; Step S4: Perform cross-modal retrieval on the food image and the recipe through the trained image-recipe retrieval model.
2. The image recipe retrieval method based on multi-level features and attention mechanism according to claim 1, wherein, Step S21 is specifically as follows: Step S211: Use the ResNet50 network pre-trained on ImageNet and the last avgpool and fc layers to extract features from the food image to obtain the initial image features; Step S212: For each formula, first use the word2vec model to obtain a feature representation, and then use Bi-LSTM to obtain the initial text features of the formula; Step S213: Use a two-stage LSTM model to represent the cooking method, that is, first, each sentence in the cooking method is represented as a skip-instructions vector, and then the sequence of skip-instructions vectors is trained by LSTM to extract the initial text features of the cooking method.
3. The image recipe retrieval method based on multi-level features and attention mechanism according to claim 2, characterized in that, Step S22 is specifically as follows: Step S221: Use average pooling and max pooling to aggregate spatial information to generate two different spatial context description features, first compress the spatial context description features, then restore the spatial context description features, and use element-wise summation to merge the output feature vectors; Step S222: Connect two spatially contextual description features along the channel axis to obtain a spatial attention map by convolving and fusing the feature maps with a standard convolutional layer; Step S223: The spatial attention map passes through the channel attention module and the spatial attention module to obtain the optimized regional image features .
4. A method for retrieving image recipes based on multi-level features and attention mechanism according to claim 3, characterized in that, Step S24 is as follows: Step S241: Focus on the words in each sentence related to each image region, and calculate the cosine similarity matrix for each region-word pair in the image-recipe pair. The formula for the cosine similarity matrix is as follows: ; wherein is an element in the set of regional image features, is an element in the set of formula features and the set of cooking method features, and when then is an element in the set of formula features, and when then is an element in the set of cooking method features; Step S242: Calculate the weighted word representation through the cosine similarity matrix. The formula is as follows: ; ; Among them, is the attention weight, is the sentence vector relative to the th food image region; Step S243: Given the sentence context, determine the importance of each image region by calculating the cosine similarity. The formula for calculating the cosine similarity is as follows: ; Step S244: The similarity between the regional image and the word is calculated by LogSumExp pooling to obtain the fine-grained relationship between the regional image and the word. The formula is as follows: ; Where k is the number of regions into which the image is divided, and λ is a parameter that determines the importance of magnifying the most relevant pair of the regional features of the food image and the attended sentence vector. is the combined feature of the recipe feature and the cooking method feature; Step S245: Calculate the context attention loss according to the fine-grained relationship. The formula is as follows: ; Among them, and are negative samples, is the margin.
5. The image recipe retrieval method based on multi-level features and attention mechanism according to claim 4, characterized in that: In step S26, the formula for the triplet loss is as follows: ; Among them, and are negative samples, is the margin parameter, d is the Euclidean distance, ; The formula for the maximum mean error is as follows: ; Among them, is a typical form of feature mapping, which is a reproducing Hilbert space with a Gaussian kernel k.
6. The image recipe retrieval method based on multi-level features and attention mechanism according to claim 5, wherein: In step S27, the formula for the translation consistency loss is as follows: ; Among them, calculate the possibility of food types for the image and the recipe respectively. The formula for the possibility is as follows: ; ; where N represents the number of food categories, and the classification loss is , where is the true food category label; Minimize and the KL divergence between them, where the formula is as follows: ; 。 7. A method for retrieving image recipes based on multi-level features and attention mechanism according to claim 6, characterized in that: The total loss formula of the image-recipe retrieval model is as follows: ; Among them, , and are constant parameters.
Citation Information
Patent Citations
Cross-modal image text retrieval method based on credibility self-adaptive matching network
CN111026894A
Cross-modal image-text retrieval method based on multi-granularity feature fusion
CN115033670A