Image-text retrieval method based on information enhancement and multimodal global-local feature alignment
By extracting and fusing multimodal features of pictures and text, and using attention mechanisms and semantic consistency constraints, the problem of difficult semantic consistency and coarse-grained alignment in the prior art is solved, and a more accurate graphic and text retrieval effect is achieved.
Patent Information
- Application Number
- CN202510186603.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-20
AI Technical Summary
The existing multimodal feature alignment technology is difficult to effectively implement semantic consistency constraints, and the coarse-grained alignment of multimodal is not fully considered.
By extracting the global features of the picture, local features and sentence features and word features of the text, and using the attention mechanism to achieve intramodal feature enhancement and intermodal feature fusion, combined with semantic consistency constraints, the information fusion and alignment of the picture and text are achieved.
It realizes more full integration and alignment of picture and text information, improves the accuracy and effect of graphic and text retrieval, and overcomes the limitation of multimodal alignment in Euclidean space.
Smart Images

Figure CN119646272B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image and text retrieval, and in particular to an image and text retrieval method based on information enhancement and multi-modal global and local feature alignment. Background Art
[0002] In the multimodal field, multimodal feature alignment is a key step in multimodal learning, which involves how to discover and establish correspondences between different data modalities. In multimodal data, different modalities may contain complementary information, and alignment is to associate this information so that it can be transferred from one modality to another.
[0003] However, how to align different modalities is currently a difficult point; in the process of multimodal feature alignment, the selection of data sets, preprocessing, methods for extracting data features, and feature alignment models have a very important impact on the quality of the alignment results.
[0004] Wang et al. proposed a Chinese cross-modal entity alignment method based on multimodal knowledge graph (Wang Huan, Song Lijuan, Du Fang. Chinese cross-modal entity alignment method based on multimodal knowledge graph [J]. Computer Engineering, 2023, 49(12): 88-95. DOI: 10.19678 / j.issn.1000-3428.0066938). Image information is introduced into the entity alignment task, and a single- and double-stream interactive pre-trained language model (CCMEA) is designed for domain fine-grained images and Chinese text. Based on the self-supervised learning method, visual and text encoders are used to extract visual and text features, and fine modeling is performed through cross encoders. Finally, the contrastive learning method is used to calculate the matching degree of image and text entities.
[0005] Ma et al. used the RoBERTa pre-trained model to extract text features after the original text data was supplemented with image captions for text modality (Rang Yuchen, Ma Jing. Multimodal alignment sentiment analysis model based on image captions [J / OL]. Data Analysis and Knowledge Discovery, 1-15 [2025-02-17]). For the image modality, ClipVisionModel was used to extract image features. The extracted text and image features were enhanced by the multimodal alignment layer based on the multimodal Transformer, and finally the multimodal fusion features were input into the multilayer perceptron for sentiment recognition and classification.
[0006] Chen et al. started by constructing two different embedding spaces (Hui Chen, Guiguang Ding, ZijiaLin, Sicheng Zhao, and Jungong Han. 2019. Cross-Modal Image-Text Retrievalwith Semantic Consistency. In Proceedings of the 27th ACM InternationalConference on Multimedia (MM '19). Association for Computing Machinery, NewYork, NY, USA, 1749–1757), namely the embedding space of images and the embedding space of text. However, instead of learning the two embedding spaces separately, they added a semantic consistency constraint to the common ranking objective function so that the two embedding spaces can be learned simultaneously and benefit from each other to achieve performance improvement.
[0007] However, the above studies either have no good constraints on semantic consistency or only study local information without considering coarse-grained alignment of multimodal features. Summary of the invention
[0008] The purpose of the present invention is to provide an image-text retrieval method based on information enhancement and multimodal global-local feature alignment, which extracts the global and local features of images and the sentence and word features of texts, and uses the attention mechanism to achieve intra-modal feature enhancement and inter-modal feature fusion, so that the information fusion of images and texts is more complete, and semantic consistency constraints are added to the model to make images and texts more fully aligned.
[0009] To achieve the above object, the present invention provides the following technical solutions:
[0010] On the one hand, the present invention provides a method for image-text retrieval based on information enhancement and multimodal global-local feature alignment, comprising the following steps:
[0011] Step 1: Get image information and corresponding text description data;
[0012] Step 2: Extract features from the image data and text data respectively to obtain global features, local features, word features, and sentence features of the text;
[0013] Step 3: Fuse the local features of the image with the global features of the image to obtain local features with global information and achieve information enhancement;
[0014] Step 4: Put the image features and text features into the cross-attention mechanism model to perform coarse-grained and fine-grained feature fusion, obtain the global fusion features of the image and text and the local fusion features of the image and text, and realize the coarse-grained and fine-grained alignment of the image and text;
[0015] Step 5: Calculate the similarity between the local fusion features of the image and text and the local features of the image and text to achieve image and text retrieval.
[0016] In some embodiments, step 1 includes the following steps:
[0017] Resize each image, adjusting the length and width of the image to 224 pixels;
[0018] Organize and summarize the text data, classify all sentences describing the same picture, and put all the classified sentences into a dictionary. The key in the dictionary is each picture name, and the value corresponding to each key is a text list describing the picture.
[0019] Perform data segmentation on images and texts, split the image and text dictionaries into training sets and test sets respectively, and store them in corresponding folders;
[0020] use The pre-trained model extracts local area images from the image, and the extraction method is:
[0021] ;
[0022] In the formula, For the pictures; For the Area pictures; For use Model.
[0023] In some embodiments, in step 2, the image data is feature extracted using the ResNet50 pre-trained model to obtain global features and local features of the image; and the text data is feature extracted using the BeRT pre-trained model to obtain word features and sentence features of the text.
[0024] In some embodiments, the computational expression for feature extraction of the image data is:
[0025] ;
[0026] In the formula, and Respectively The global image features and Features of regional images; For the model using ResNet50;
[0027] The calculation expression for feature extraction of the text data is:
[0028] ;
[0029] In the formula, For the Sentence features; For the The word features of a sentence; For the Sentences; For models using BeRT.
[0030] In some embodiments, step 3 includes the following steps:
[0031] Obtain local features with global information;
[0032] After obtaining the local features containing global information, an attention mechanism model is executed to calculate the global features with enhanced information;
[0033] After obtaining the global features of information enhancement, an attention mechanism model is executed to obtain the local information of information enhancement.
[0034] In some embodiments, the calculation expression for obtaining the local features having global information is:
[0035] ;
[0036] ;
[0037] The calculation expression of the local information obtained by information enhancement is:
[0038] ;
[0039] ;
[0040] ;
[0041] ;
[0042] The calculation expression of the local information obtained by information enhancement is:
[0043] ;
[0044] ;
[0045] In the formula, is the global image feature and regional image features The importance weight of is a local feature that contains global information; Global features for information enhancement; It is a combination of mapping function, Tanh activation function and batch normalization; For splicing operation; for and The importance weight of is the average characteristic of regional characteristics; for The weight matrix of for and Importance score; is the number of regional images; Local features for information enhancement; for and The importance weight of .
[0046] In some embodiments, in step 4, first, using an autoencoder and Convert the image features and text features to the same feature dimension, so that their feature dimensions are both converted to 1024. The specific calculation expression is:
[0047] ;
[0048] ;
[0049] Secondly,
[0050] Add weight features to global features and regional features:
[0051] ;
[0052] Add weight features to sentence features and word features:
[0053] ;
[0054] In the formula, Respectively Global features and regional features after feature conversion; For use Model; Respectively The sentence feature dimension is 1024 sentence features and word features; For use Model; After adding weights, they are The global and regional features of the map; are the weight matrices of global features and regional features respectively; After adding weights, they are The sentence features and word features of a sentence; They are the weight matrices of sentence features and word features respectively; Respectively The sentence feature dimension is 768 sentence features and word features.
[0055] In some embodiments, the coarse-grained and fine-grained feature fusion comprises the following steps:
[0056] Put global features and sentence features into the cross-attention mechanism model to achieve coarse-grained feature fusion and achieve coarse-grained and fine-grained alignment of images and texts:
[0057] ;
[0058] ;
[0059] Put regional features and word features into the cross-attention mechanism model to achieve fine-grained feature fusion:
[0060] ;
[0061] ;
[0062] In the formula, For the first The global feature is Q (query), with the first The sentence features of a sentence are the image-sentence fusion features of K (key) and V (value); For the first The sentence feature of the sentence is Q, with the first The global features are K and V sentence-image fusion features; is the activation function, in order to positively activate the relevance score and improve stability; is a scaling parameter that controls the distribution range of the Softmax input value; For the first The features of all regions in a picture are Q (query), with the first The features of all words in a sentence are region-word fusion features of K (key) and V (value); For the first All the word features of the sentence are Q, with the first All regional features of the image are word-region fusion features of K and V; is the vector product operation.
[0063] In some embodiments, step 5 includes the following steps:
[0064] The similarity between the image-text regional fusion features and the regional image features and word features is calculated respectively to obtain the regional image-text similarity matrix and the regional text-image similarity matrix, where each element in each matrix represents the similarity between the regional image and the word:
[0065] ;
[0066] ;
[0067] The global fusion features of images and texts are respectively calculated for similarity with the global image features and sentence features to obtain the global image-text similarity matrix and the global text-image similarity matrix, where each element in each matrix represents the similarity between the image and the sentence:
[0068] ;
[0069] ;
[0070] In the formula, are the similarity scores for images and texts respectively; For the The average similarity of the regional image of the image to all words in all sentences, that is, the similarity to the sentences; For the The average similarity of the words in the sentence to all regional images of all images, that is, the similarity to the images; It is a modular operation; For the first The features of all regions in a picture are Q (query), with the first The features of all words in a sentence are region-word fusion features of K (key) and V (value); For the first All the word features of the sentence are Q, with the first All regional features of the image are word-region fusion features of K and V; For the first The global feature is Q (query), with the first The sentence features of a sentence are the image-sentence fusion features of K (key) and V (value); For the first The sentence feature of the sentence is Q, with the first The global features are K and V sentence-image fusion features; Respectively Global features and regional features after feature conversion; For use Model; Respectively The sentence features and word features of a sentence; For the The similarity of a picture to all sentences; For the Similarity of sentence to all images; The table is the image-text similarity matrix; is the text-image similarity matrix.
[0071] In some embodiments, in step 5, using a mean square error (MSE) loss function to implement semantic consistency constraints on the similarity matrix from the region image to the word and the similarity matrix from the word to the region image includes the following steps:
[0072] Use mean squared error for the region image-text similarity matrix and the region text-image similarity matrix:
[0073] ;
[0074] Use mean squared error for the global image-text similarity matrix and the global text-image similarity matrix:
[0075] ;
[0076] Use the triplet loss function for the region image-text similarity matrix and the region text-image similarity matrix:
[0077] ;
[0078] Use the triplet loss function for the global image-text similarity matrix and the global text-image similarity matrix:
[0079] ;
[0080] In the formula, is the mean square error loss of the region similarity matrix; is the region similarity matrix triplet loss; , They are The Line The elements of the column and The Line Elements of a column; for Anchor in; is the similarity between the anchor and the positive sample; is the similarity between the anchor and the hard negative sample; for Anchor in; is the similarity between the anchor and the positive sample; is the similarity between the anchor and the hard negative sample; is the mean square error loss of the region similarity matrix; is the triplet loss of the global similarity matrix; For the The similarity of a picture to all sentences; For the Similarity of sentence to all images; is the image-text similarity matrix; is the text-image similarity matrix.
[0081] On the other hand, the present invention provides a picture-text retrieval system based on information enhancement and multimodal global local feature alignment, using a picture-text retrieval method based on information enhancement and multimodal global local feature alignment, comprising:
[0082] Acquisition module: used to obtain image information and corresponding text description data;
[0083] Feature extraction module: used to extract features from image data and text data respectively, and obtain global features of the image, local features of the image, and features of text words and sentences;
[0084] Feature fusion module: used to fuse local features of the image with global features of the image to obtain local features with global information and achieve information enhancement;
[0085] Coarse-grained and fine-grained alignment module: used to put image features and text features into the cross-attention mechanism model for coarse-grained and fine-grained feature fusion, obtain global fusion features of images and texts and local fusion features of images and texts, and realize coarse-grained and fine-grained alignment of images and texts;
[0086] Similarity calculation module: used to calculate the similarity between the local fusion features of the image and text and the local features of the image and text to achieve image and text retrieval.
[0087] Compared with the prior art, the present invention has the following beneficial effects:
[0088] The present invention performs information enhancement and feature fusion on global information and local information respectively, ensuring coarse-grained and fine-grained alignment; the present invention uses mean square error to achieve semantic consistency between images and texts, improving the alignment of images and texts. The present invention also provides a new idea for multimodal feature alignment, combining global information and local information to achieve multimodal feature alignment, which can overcome the limitation of inaccurate alignment caused by multimodal alignment in Euclidean space. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1 This is a schematic diagram of the process of Example 1 of the present invention;
[0090] Figure 2 This is a schematic diagram of the process of the image and text retrieval method according to Embodiment 1 of the present invention;
[0091] Figure 3 This is a schematic diagram of the cross-attention mechanism model structure of Example 1 of the present invention;
[0092] Figure 4 This is a schematic diagram of the information enhancement structure of Example 1 of the present invention;
[0093] Figure 5 This is a schematic diagram of the structure of Example 2 of the present invention. DETAILED DESCRIPTION
[0094] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0095] Embodiment 1:
[0096] See also Figure 1-Figure 4 , a method for image-text retrieval based on information enhancement and multimodal global-local feature alignment, comprising the following steps:
[0097] Step 1: Get image information and corresponding text description data.
[0098] This embodiment uses the Flickr8K dataset to illustrate the image-text retrieval method proposed in the present invention, and the method is also applicable to other image-text data. The data in the Flickr8K dataset is preprocessed. The Flickr8K dataset is multimodally aligned image data and text data. The dataset selects a total of 8,000 images of certain behaviors of people or animals from the image social networking site Flickr. Each image corresponds to 5 manually labeled sentence descriptions.
[0099] There is a token file in the dataset, which contains 40,000 lines of data. Each line of data consists of a picture name and a text description of the corresponding picture. There is an images file, which contains 8,000 pictures of different styles. There are three text files. Each line in the text file consists of a different picture name. The text files are divided into training set text, validation set text, and test set text. The training set text has 6,000 lines of data, the validation set text has 1,000 lines of data, and the test set text has 1,000 lines of data.
[0100] The data preprocessing comprises the following steps:
[0101] The Flickr8K dataset is divided into a training set and a test set. The training set includes the training set text and the corresponding pictures and picture descriptions, with a total of 6,000 pictures and 30,000 sentences; the test set includes the test set text and the corresponding pictures and picture descriptions, with a total of 1,000 pictures and 5,000 sentences.
[0102] Based on the above data set, preprocess the image data and text data:
[0103] First, resize the image data and modify the length and width of each image to 224 pixels. The shape of the image is (3, 224, 224), where 3 represents the number of channels and 224 represents the length and width.
[0104] The text data is sorted and summarized, all sentences describing the same picture are classified, and all the classified sentences are put into a dictionary. Specifically, the key in the dictionary is each picture name, and the value corresponding to each key in the dictionary is a list of all the classified sentences describing the picture.
[0105] Secondly, perform data segmentation on images and texts. Specifically, images are segmented into training sets and test sets, and stored in corresponding folders; text dictionaries are segmented into training sets and test sets, and stored in corresponding folders;
[0106] Finally, extract the region of the image; use Pre-train the model and set the maximum number of images for target detection to 4, that is, generate at most 4 region images containing image information. Take each image as input, and output 4 images containing part of the image as region images. If the number of detected region images is less than 4, copy the detected region images until there are 4 region images.
[0107] The calculation expression for regional image extraction is:
[0108] ;
[0109] In the formula, For the pictures; For the Area pictures; For use Model.
[0110] Step 2: Extract features from the image data and text data respectively to obtain global features of the image, local features of the image, and word features and sentence features of the text.
[0111] First, use the ResNet50 pre-trained model to extract features from the global image and the four regional images. Each image is input into the ResNet50 pre-trained model, and each image obtains an image feature with a shape of (1, 1000), where 1 represents this image and 1000 represents that each image has 1000 numbers as its feature. The regional images are stacked to obtain an image feature of (4, 1000), where 4 represents that 4 regional images are included.
[0112] The calculation expression for the image feature extraction is:
[0113] ;
[0114] In the formula, and Respectively The global image features and Features of regional images; This is a model using ResNet50.
[0115] Finally, the BeRT pre-trained model is used to extract features from the text, where the length of each sentence is fixed to 32 words. Each sentence is input into the BeRT pre-trained model, and each sentence obtains a word feature of the shape of (1, 32, 768) and a sentence feature of (1, 768), where 1 represents a sentence, 32 represents 32 words, and 768 represents 768 numbers as the features of each word / sentence.
[0116] The calculation expression for text feature extraction is:
[0117] ;
[0118] In the formula, For the Sentence features; For the The word features of a sentence; For the Sentences; For models using BeRT.
[0119] The image data and the text data are batch processed, and each batch of the image data and the text data is set to 40.
[0120] Step 3: Fuse the local features of the image with the global features of the image to obtain local features with global information and achieve information enhancement.
[0121] Obtain local features with global information:
[0122] ;
[0123] ;
[0124] After obtaining the local features containing global information, an attention mechanism model is executed to calculate the global features with enhanced information:
[0125] ;
[0126] ;
[0127] ;
[0128] ;
[0129] After obtaining the global features of information enhancement, an attention mechanism model is executed to obtain the local information of information enhancement:
[0130] ;
[0131] ;
[0132] In the formula, is the global image feature and regional image features The importance weight of is a local feature that contains global information; Global features for information enhancement; It is a combination of mapping function, Tanh activation function and batch normalization; For splicing operation; for and The importance weight of is the average characteristic of regional characteristics; for The weight matrix of for and Importance score; is the number of regional images; Local features for information enhancement; for and The importance weight of .
[0133] Using Autoencoders and Convert image features and text features so that their feature dimensions are converted to 1024:
[0134] ;
[0135] at this time The shape is (1, 1024), The shape is (4, 1024).
[0136] ;
[0137] at this time The shape is (1, 1024), The shape is (1, 32, 1024).
[0138] In the formula, Respectively Global features and regional features after feature conversion; For use Model; Respectively The sentence features and word features of a sentence; For use Model; Respectively sentence features and The word features of a sentence.
[0139] Step 4: Put the image features and text features into the cross-attention mechanism model to perform coarse-grained and fine-grained feature fusion, obtain the global fusion features of the image and text and the local fusion features of the image and text, and achieve coarse-grained and fine-grained alignment of the image and text.
[0140] First, add weight features to global features and regional features:
[0141] ;
[0142] Add weight features to sentence features and word features:
[0143] ;
[0144] In the formula, After adding weights, they are The global and regional features of the map; are the weight matrices of global features and regional features respectively; After adding weights, they are The sentence features and word features of a sentence; They are the weight matrices of sentence features and word features respectively.
[0145] Secondly, the global features and sentence features are put into the cross-attention mechanism for coarse-grained feature fusion to achieve coarse-grained and fine-grained alignment of images and texts:
[0146] ;
[0147] ;
[0148] Put regional features and word features into the cross-attention mechanism model to achieve fine-grained feature fusion:
[0149] ;
[0150] ;
[0151] In the formula, For the first The global feature is Q (query), with the first The sentence features of a sentence are the image-sentence fusion features of K (key) and V (value); For the first The sentence feature of the sentence is Q, with the first The global features are K and V sentence-image fusion features; is the activation function, in order to positively activate the relevance score and improve stability; is a scaling parameter that controls the distribution range of the Softmax input value; For the first The features of all regions in a picture are Q (query), with the first The features of all words in a sentence are region-word fusion features of K (key) and V (value); For the first All the word features of the sentence are Q, with the first All regional features of the image are word-region fusion features of K and V; is the vector product operation.
[0152] Step 5: Calculate the similarity between the local fusion features of the image and text and the local features of the image and text to achieve image and text retrieval.
[0153] The similarity between the image-text region fusion features and the regional image features and word features is calculated respectively to obtain the regional image-text similarity matrix and the regional text-image similarity matrix, where each element in each matrix represents the similarity between the regional image and the word.
[0154] The similarity calculation expression is:
[0155] ;
[0156] ;
[0157] The mean square error (MSE) is used to implement the semantic consistency constraints of the similarity matrix from the regional image to the word and the similarity matrix from the word to the regional image. The specific operations include:
[0158] The global fusion features of images and texts are respectively used to calculate similarity with the global image features and sentence features to obtain the global image-text similarity matrix and the global text-image similarity matrix, where each element in each matrix represents the similarity between an image and a sentence.
[0159] The similarity calculation expression is:
[0160] ;
[0161] ;
[0162] In the formula, are the similarity scores for images and texts respectively; For the The average similarity of the regional image of the image to all words in all sentences, that is, the similarity to the sentences; For the The average similarity of the words in the sentence to all regional images of all images, that is, the similarity to the images; It is a modular operation; For the first The features of all regions in a picture are Q (query), with the first The features of all words in a sentence are region-word fusion features of K (key) and V (value); For the first All the word features of the sentence are Q, with the first All regional features of the image are word-region fusion features of K and V; For the first The global feature is Q (query), with the first The sentence features of a sentence are the image-sentence fusion features of K (key) and V (value); For the first The first sentence is Q, The global features are K and V sentence-image fusion features; Respectively Global features and regional features after feature conversion; For use Model; Respectively The sentence features and word features of a sentence; For the The similarity of the image to all sentences; For the Similarity of sentence to all images; The table is the image-text similarity matrix; is the text-image similarity matrix.
[0163] Use mean squared error for the region image-text similarity matrix and the region text-image similarity matrix:
[0164] ;
[0165] Use mean squared error for the global image-text similarity matrix and the global text-image similarity matrix:
[0166] ;
[0167] Use the triplet loss function for the region image-text similarity matrix and the region text-image similarity matrix:
[0168] ;
[0169] Use the triplet loss function for the global image-text similarity matrix and the global text-image similarity matrix:
[0170] ;
[0171] In the formula, is the mean square error loss of the region similarity matrix; is the region similarity matrix triplet loss; , They are The Line The elements of the column and The Line Elements of a column; for Anchor in; is the similarity between the anchor and the positive sample; is the similarity between the anchor and the hard negative sample; for Anchor in; is the similarity between the anchor and the positive sample; is the similarity between the anchor and the hard negative sample; is the mean square error loss of the region similarity matrix; is the triplet loss of the global similarity matrix; For the The similarity of a picture to all sentences; For the Similarity of sentence to all images; is the image-text similarity matrix; is the text-image similarity matrix.
[0172] Example 2
[0173] like Figure 5 As shown, a picture-text retrieval system based on information enhancement and multimodal global-local feature alignment includes:
[0174] Acquisition module: used to obtain image information and corresponding text description data;
[0175] Feature extraction module: used to extract features from image data and text data respectively, and obtain global features of the image, local features of the image, and features of text words and sentences;
[0176] Feature fusion module: used to fuse local features of the image with global features of the image to obtain local features with global information and achieve information enhancement;
[0177] Coarse-grained and fine-grained alignment module: used to put image features and text features into the cross-attention mechanism model for coarse-grained and fine-grained feature fusion, obtain global fusion features of images and texts and local fusion features of images and texts, and realize coarse-grained and fine-grained alignment of images and texts;
[0178] Similarity calculation module: used to calculate the similarity between the local fusion features of the image and text and the local features of the image and text to achieve image and text retrieval.
[0179] The image-text retrieval system based on information enhancement and multimodal global-local feature alignment of the present invention can be installed in a computer device. The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as an image-text retrieval program based on information enhancement and multimodal global-local feature alignment. The memory includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. The processor is the control core of the electronic device, and uses various interfaces and lines to connect the various components of the entire computer device, and executes various functions of the computer device and processes data by running or executing programs or modules stored in the memory, and calling data stored in the memory.
[0180] The module described in the present invention refers to a series of computer program segments that can be executed by a processor of a computer device and can complete fixed functions, and is stored in a memory of the computer device.
[0181] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for image-text retrieval based on information enhancement and multimodal global-local feature alignment, characterized in that: The following steps are involved: Step 1: Get image information and corresponding text description data; Step 2: Extract features from the image data and text data respectively to obtain global features, local features, word features, and sentence features of the text; Step 3: Fusing the local features of the image with the global features of the image to obtain local features with global information and achieve information enhancement; including the following steps: Obtain local features with global information: ;in, ; After obtaining the local features containing global information, an attention mechanism model is executed to calculate the global features with enhanced information: ; in, ; ; ; After obtaining the global features of information enhancement, an attention mechanism model is executed to obtain the local information of information enhancement: ;in, ; In the formula, is the global image feature and regional image features Importance weight of is a local feature that contains global information; Global features for information enhancement; It is a combination of mapping function, Tanh activation function and batch normalization; For splicing operation; for and The importance weight of is the average characteristic of regional characteristics; for The weight matrix of for and Importance score; is the number of regional images; Local features for information enhancement; for and The importance weight of Step 4: Put the image features and text features into the cross-attention mechanism model to perform coarse-grained and fine-grained feature fusion, obtain the global fusion features of the image and text and the local fusion features of the image and text, and realize the coarse-grained and fine-grained alignment of the image and text; Step 5: Calculate the similarity between the local fusion features of the image and text and the local features of the image and text to achieve image and text retrieval.
2. According to claim 1, a method for image-text retrieval based on information enhancement and multimodal global-local feature alignment is characterized in that: Step 1 includes the following steps: Adjust the length and width of the image to 224 pixels; Classify all sentences describing the same picture and put all the classified sentences into a dictionary. The key in the dictionary is each picture name, and the value corresponding to each key is a text list describing the picture. Divide the image and text dictionaries into training sets and test sets respectively, and store them in corresponding folders; The Faster-rcnn pre-trained model is used to extract local area images from the image. The extraction method is as follows: ; In the formula, For the pictures; For the Area pictures; For use Model.
3. The image-text retrieval method based on information enhancement and multimodal global-local feature alignment according to claim 1, characterized in that: In step 2, the ResNet50 pre-trained model is used to extract features from the image data to obtain global features and local features of the image; the BeRT pre-trained model is used to extract features from the text data to obtain word features and sentence features of the text.
4. The image-text retrieval method based on information enhancement and multimodal global-local feature alignment according to claim 3, characterized in that: The calculation expression for feature extraction of the image data is: ; In the formula, and Respectively The global image features and Features of regional images; For the model using ResNet50; The calculation expression for feature extraction of the text data is: ; In the formula, For the Sentence features; For the The word features of a sentence; For the Sentences; For models using BeRT; For the pictures; For the A picture of the area.
5. The image-text retrieval method based on information enhancement and multimodal global-local feature alignment according to claim 1, characterized in that: In step 4, first, use the autoencoder and Convert image features and text features to the same feature dimension: ; ; Secondly, Add weight features to global features and regional features: ; Add weight features to sentence features and word features: ; In the formula, Respectively Global features and regional features after feature conversion; For use Model; Respectively The sentence feature dimension is 1024 sentence features and word features; For use Model; After adding weights, they are The global and regional features of the map; are the weight matrices of global features and regional features respectively; After adding weights, they are The sentence features and word features of a sentence; They are the weight matrices of sentence features and word features respectively; Respectively The sentence feature dimension is 768 sentence features and word features.
6. The image-text retrieval method based on information enhancement and multimodal global-local feature alignment according to claim 5, characterized in that: The coarse-grained and fine-grained feature fusion comprises the following steps: Put global features and sentence features into the cross-attention mechanism model to achieve coarse-grained feature fusion: ; ; Put regional features and word features into the cross-attention mechanism model to achieve fine-grained feature fusion: ; ; In the formula, For the first The global feature is Q, The sentence features of a sentence are the image-sentence fusion features of K and V; For the first The first sentence is Q, The global features are K and V sentence-image fusion features; is the activation function; is the scaling parameter; For the first The features of all regions in the image are Q, with the first The features of all words in a sentence are the region-word fusion features of K and V; For the first All the word features of the sentence are Q, with the first All regional features of the image are word-region fusion features of K and V; is the vector product operation.
7. The image-text retrieval method based on information enhancement and multimodal global-local feature alignment according to claim 1, characterized in that: Step 5 includes the following steps: The similarity between the image-text regional fusion features and the regional image features and word features is calculated to obtain the regional image-text similarity matrix and the regional text-image similarity matrix: ; ; The global fusion features of images and texts are respectively calculated for similarity with the global image features and sentence features to obtain the global image-text similarity matrix and the global text-image similarity matrix: ; ; In the formula, are the similarity scores for images and texts respectively; For the The average similarity of the regional image of the image to all words in all sentences, that is, the similarity to the sentences; For the The average similarity of the words in the sentence to all regional images of all images, that is, the similarity to the images; is a modular operation; For the first The features of all regions in the image are Q, with the first The features of all words in a sentence are the region-word fusion features of K and V; For the first All the word features of the sentence are Q, with the first All regional features of the image are word-region fusion features of K and V; For the first The global feature is Q, The sentence features of a sentence are the image-sentence fusion features of K and V; For the first The sentence feature of the sentence is Q, with the first The global features are K and V sentence-image fusion features; Respectively Global features and regional features after feature conversion; For use Model; Respectively The sentence features and word features of a sentence; For the The similarity of the image to all sentences; For the Similarity of sentence to all images; The table is the image-text similarity matrix; is the text-image similarity matrix.
8. The image-text retrieval method based on information enhancement and multimodal global-local feature alignment according to claim 7, characterized in that: In step 5, the mean square error loss function is used to implement the semantic consistency constraint of the similarity matrix from the region image to the word and the similarity matrix from the word to the region image, including the following steps: Use mean squared error for the region image-text similarity matrix and the region text-image similarity matrix: ; Use mean squared error for the global image-text similarity matrix and the global text-image similarity matrix: ; Use the triplet loss function for the region image-text similarity matrix and the region text-image similarity matrix: ; Use the triplet loss function for the global image-text similarity matrix and the global text-image similarity matrix: ; In the formula, is the mean square error loss of the region similarity matrix; is the region similarity matrix triplet loss; , They are The Line The elements of the column and The Line Elements of a column; for Anchor in; is the similarity between the anchor and the positive sample; is the similarity between the anchor and the hard negative sample; for Anchor in; is the similarity between the anchor and the positive sample; is the similarity between the anchor and the hard negative sample; is the mean square error loss of the region similarity matrix; is the triplet loss of the global similarity matrix; For the The similarity of the image to all sentences; For the Similarity of sentence to all images; is the image-text similarity matrix; is the text-image similarity matrix.
Citation Information
Patent Citations
Cross-modal image-text retrieval method based on multi-level semantic alignment
CN116821391A