Cross-modal image-text retrieval method based on multi-window attention mechanism

By introducing a multi-window attention mechanism and cross-mapping alignment network in cross-modal graphic and text retrieval, the problem of insufficient interaction between image and text in the prior art is solved, and higher semantic retrieval accuracy and stability are achieved.

CN120011609AInactive Publication Date: 2025-05-16OCEAN UNIV OF CHINA

Patent Information

Application Number
CN202510495885.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing cross-modal graphic and text retrieval methods have shortcomings in the local information alignment of images and text, global feature considerations, interaction between modals and fine-grained alignment, resulting in insufficient semantic correlation.

Method used

The cross-modal graphic and text retrieval method based on the multi-window attention mechanism is adopted to extract the local and global features of the image through the local window and row window attention mechanism, and fine-grained semantic alignment between modes and within modes is combined with the cross-mapping alignment network.

Benefits of technology

It improves the matching and retrieval effect of graphics and text, enhances the representation ability of visual features, improves the stability and accuracy of semantic alignment, and achieves higher semantic retrieval accuracy of image and text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011609A_ABST
    Figure CN120011609A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image-text retrieval, and discloses a cross-modal image-text retrieval method based on a multi-window attention mechanism, which comprises the following steps: S1, data preprocessing and embedding: preprocessing an image and a text, and converting the image and the text into a vector matrix; s2, image and text feature extraction: for the image vector matrix, extracting local and global features of image blocks by using a Transform module based on a multi-window attention mechanism to obtain a visual feature matrix; extracting text features from the text vector matrix; s3, performing cross-modal semantic alignment and semantic similarity calculation: performing cross-modal semantic alignment on the visual feature matrix and the text feature matrix, and obtaining a fine-grained semantic relationship matrix of the image blocks and the text words by adopting a cross-mapping alignment network; mapping representation and semantic alignment between the two modes are achieved, and the semantic similarity of the image and the text is calculated; and S4, outputting a retrieval result. Through the method, the image-text matching and retrieval effects are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image and text retrieval technology, and in particular relates to a cross-modal image and text retrieval method based on a multi-window attention mechanism. Background Art

[0002] Cross-modal retrieval is to retrieve samples of one modality with similar semantics through samples of another modality. Cross-modal image-text retrieval is to retrieve semantically related samples from text (or image) based on a given image (or text), which usually includes two subtasks: image-text retrieval (i2t) and text-image retrieval (t2i).

[0003] The calculation of cross-modal image-text similarity can be mainly divided into one-to-one matching method, many-to-many matching method and method based on multimodal pre-trained model. (1) One-to-one matching method. (2) Many-to-many matching method. For example, some researchers proposed a stacked cross attention network (SCAN), which models the similarity between multiple regions of an image and multiple words of a text. (3) Image-text retrieval based on multimodal pre-trained model. The pre-trained model is used for cross-modal retrieval, mainly by pre-training on ultra-large-scale datasets to learn effective cross-modal general representations, and then migrating these representations to downstream tasks through fine-tuning. Cross-modal models based on pre-training are divided into two categories: two-stream models and single-stream models.

[0004] By analyzing the existing cross-modal research methods, the following existing problems are summarized: (1) Although the one-to-one matching method can achieve global information alignment, it ignores the local information of images and texts, and ignores the fine-grained alignment between image regions and sentence words; (2) Although the many-to-many matching method considers the correspondence between fine-grained segments, it lacks consideration of global features and still has deficiencies in combining intra-modal interactions with inter-modal correspondence. (3) Image-text retrieval method based on multimodal pre-trained model. The single-stream model uses a deep fusion encoder with cross-modal attention to model the relationship between image and text. The two modalities can fully interact, but cannot work offline. The two-stream model uses an image encoder and a text encoder to represent the image and text respectively, which can achieve offline work. However, due to this, the two-stream model has the problem of insufficient interaction between modalities. Therefore, how to accurately understand the semantic content in the image, how to better associate the correlation between the two modalities, and how to combine fine-grained information to make the correlation between modalities more accurate are all challenges faced by the two-stream model cross-modal image-text retrieval task. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides a cross-modal image-text retrieval method based on a multi-window attention mechanism to improve image-text matching and retrieval effects. The present invention can realize the tasks of searching for text with images and searching for images with text in Internet resource retrieval. In the retrieval process, different processing branches are matched according to the type of search data (image or text), and the present invention is used to complete data processing, feature extraction, modality alignment, similarity matching and other tasks. When obtaining the retrieval results, a search engine is used to retrieve the image-text representation containing semantic information, and recall the Top-N semantically similar samples. The present invention can fully extract the interactive features and semantic information of image and text descriptions. Compared with previous image-text retrieval methods, the present invention retains the representation characteristics of the deep learning semantic model vector and improves the accuracy of image and text semantic retrieval.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: The cross-modal image-text retrieval method based on the multi-window attention mechanism includes the following steps: S1. Data preprocessing and embedding: Preprocess images and texts, and use pre-trained visual models and language models to perform image block embedding and word embedding to convert images and texts into vector matrices; S2, image and text feature extraction: For the image vector matrix obtained by S1, the Transformer module based on the multi-window attention mechanism is used to extract the local and global features of the image block to obtain the visual feature matrix ; For the text vector matrix obtained by S1, use the Bert model to extract text features and obtain the text feature matrix ; S3, cross-modal semantic alignment and semantic similarity calculation: The visual feature matrix obtained by S2 and the text feature matrix Perform cross-modal semantic alignment and use a cross-mapping alignment network to obtain a fine-grained semantic relationship matrix between image blocks and text words; achieve mapping representation and semantic alignment between the two modalities; then calculate the semantic similarity between the image and text through cosine similarity based on the obtained fine-grained semantic relationship matrix; S4. Output search results: The results of cross-modal cross-alignment and intra-modal mapping alignment are comprehensively considered according to the similarity.

[0007] Furthermore, in step S1, it is divided into two branches, namely, a visual branch and a text branch. In the visual branch, the image is first preprocessed, an image is evenly divided into a number of image blocks, and then the image blocks are embedded, and the image blocks are converted into vector matrices; in the text branch, the text is preprocessed, and a pre-trained language model is used for word embedding to convert the text into a vector matrix.

[0008] Furthermore, in step S2, in the visual branch, the image block vector matrix obtained in S1 is processed by a Transformer module based on a multi-window attention mechanism to extract local and global features of the image block and obtain a fused feature matrix as the visual feature matrix , and output; The multi-window attention mechanism includes local window attention and row and column window attention. The specific processing steps are as follows: S2-1. Extract local features of image blocks using local window attention: (1) Image block rearrangement: After the original image is segmented, a group of one-dimensionally arranged image blocks are obtained. First, according to the position information of the segmented image blocks in the original image, the segmented image blocks are rearranged so that their relative position relationship in the original image is maintained; (2) Local window selection: Use a sliding window of preset size to select a local image block on the rearranged image block; (3) Local attention calculation: When calculating attention, only the relationship between the anchor image block and other image blocks in the window is considered, and the relationship between image blocks outside the window is ignored; (4) Feature extraction: Based on the calculated attention, weighted summation is performed on the local image blocks to obtain local features; S2-2. Use row and column window attention to extract contextual features of image blocks as global features: The row and column window attention includes row window attention and column window attention. The division of row and column windows is to expand the window size horizontally and vertically with the current image block, i.e., the anchor image block, as the center. The calculation of row and column window attention is the same as that of local window attention. S2-3. Feature fusion: The global features extracted by row window attention and column window attention are fused with the local features extracted by local window attention. The fused features are used as the final feature representation of the image, namely the visual feature matrix .

[0009] Furthermore, the calculation process of row and column window attention in step S2-2 is as follows: (1) Row window attention extraction: With the anchor image block as the center, the window is expanded horizontally to form a row window; the attention weights between the anchor image block and other image blocks in the row window are calculated; the attention between non-anchor image blocks in the row window is masked; (2) Column window attention extraction: With the anchor image block as the center, the window is expanded in the vertical direction to form a column window; the attention weights between the anchor image block and other image blocks in the column window are calculated; and the attention between non-anchor image blocks in the column window is masked.

[0010] Furthermore, in the text branch of S2, the word vector matrix obtained in S1 is used to extract text features using Bert; in the text branch, the deep semantic representation of the text is learned based on the pre-trained BERT model, and the feature representation of the text is ,in, Indicates the number of words in the text. Indicates The characteristics of the word, , Represents the dimension of the word feature vector; after a fully connected layer, the dimension of the text feature vector is unified to the same dimension as the image feature, and the converted text feature is represented as , that is, the text feature matrix .

[0011] Furthermore, in S3, the specific steps are as follows: S3-1. Mapping representation between visual and textual modalities: The visual feature matrix Transformed into a text feature matrix through a linear embedding layer Vectors of the same dimension, converting the visual feature matrix and text feature matrix Perform a connection operation to obtain a new feature matrix Z as the input of the cross-mapping alignment network, and calculate the attention coefficient matrix S between the two modalities; use the attention coefficient matrix S to calculate the weighted feature matrix and the mapping feature matrix; S3-2, cross alignment and mapping alignment: Performing an alignment operation based on the weighted feature matrix and the mapped feature matrix obtained in step S3-1 includes: Inter-modal cross-alignment: Directly match the fine-grained relationship between image patches and text words through a cross-modal attention matrix; Intra-modal mapping alignment: Use the attention score of the other modality to reconstruct the feature representation of the own modality and enhance intra-modal consistency; S3-3. Final similarity calculation: The obtained attention coefficient matrix and mapping feature matrix constitute a fine-grained semantic relationship matrix of image blocks and text words, which is used to calculate similarity.

[0012] Furthermore, the attention coefficient matrix in step S3-1 is calculated as follows: Through the attention mechanism Get the complete attention coefficient matrix S between the two modalities: ; Mapping representation refers to representing one modality based on another modality. Specifically, it uses a collection of text segments as the representation of visual segments, and vice versa. Represents the mapping representation of vision to text, Each line of is regarded as a representation of a text segment by a certain visual segment. Represents the mapping of text to vision, Each column of is regarded as a representation of a certain text segment set to a visual segment; and Do matrix multiplication to get the visual attention matrix The mapping matrix : ; Depend on and The weighted visual feature matrix is ​​obtained respectively and the mapping feature matrix : ; ; Where V is the basis matrix used for linear transformation; Similarly, we get the text attention matrix The mapping matrix , weighted text feature matrix and the mapping feature matrix : ; ; .

[0013] Further, in step S3-2, according to the obtained weighted visual feature matrix and the text feature matrix , perform alignment operations as follows: (1) Cross-alignment between modes: Inter-modal alignment seeks fine-grained associations between images and text. For each segment in vision or text, cross-alignment uses cross-modal attention to find the most relevant segment from the opposite modality; including: Calculating image-text similarity: measuring the token level The image and Then, based on this, we calculate the similarity of the object level and get the The image and Similarity of texts ; Calculate text-image similarity: Similarly, the method of calculating image-text similarity is used to obtain the first The text sub- The similarity of images ; (2) Intra-modal mapping alignment: Intra-modal mapping alignment is to find the fine-grained semantic association between images and texts from another perspective. It is implemented based on the original feature representation and the mapped feature representation. The attention mechanism is used to map the data of different modalities, and the correlation between the two is measured through the consistency relationship. For the visual modality, is the original feature representation, It is a mapping feature representation obtained based on the attention coefficient matrix of the language modality. The two are descriptions of the image from the visual perspective and the language perspective respectively. The strength of their consistency relationship represents the strength of the correlation between the image and the text. images and The similarity of texts is expressed as: ; ; in, It is The original feature representation of the image Tokens, By The text obtained The mapping representation of an image is Indicates images and The intermediate result of the similarity measurement between texts, Indicates the final images and The similarity measure between the texts, n represents the number of tokens in the visual modality; Similarly, within the language mode, The text and The similarity of images is expressed as: ; ; in, It is The original feature representation of the text Tokens, By The image obtained The mapping representation of the text, Indicates The text and The intermediate result of the similarity measurement between images, Indicates the final The text and The similarity measure between images, m represents the number of tokens in the language modality; S3-3. Final similarity calculation: The final similarity between image and text is expressed as: ; Similarly, the final similarity between text and image is expressed as: .

[0014] The final image-text and text-image similarities are obtained by combining the results of cross-modal alignment and intra-modal mapping alignment.

[0015] Compared with the prior art, the present invention has the advantages of: (1) First, for image feature extraction, the present invention adopts multi-window attention that combines local window attention and row and column window attention to extract local details of the image (such as the association between adjacent image blocks) and global context (long-distance dependence in the horizontal / vertical direction), respectively, enhance the representation ability of visual features, and make the mined feature information more complete.

[0016] (2) Then, in the image-text alignment work, a cross-mapping alignment network is proposed to effectively explore the token-level fine-grained alignment of images and texts from two levels: inter-modality (cross-modal cross-alignment) and intra-modality (mapping alignment). The intra-modal cross-alignment enhances the interaction between modalities and the intra-modal mapping alignment complements each other, reducing the interference of meaningless alignment and erroneous alignment, thereby improving the stability and accuracy of semantic alignment.

[0017] (3) Two-stream model design: a Transformer-based visual branch and a BERT-based language branch, optimizing cross-modal interaction efficiency through unified dimensional mapping and feature fusion.

[0018] (4) The present invention is particularly suitable for pest and disease image and text retrieval. The method of the present invention can provide a pest and disease image and text retrieval method based on deep learning. Compared with the existing common methods, the method uses a multi-window attention mechanism to extract the interactive information of the images to be matched. At the same time, the cross-mapping alignment network performs fine-grained semantic alignment of image blocks and text words from both the intra-modal and inter-modal levels, fully exploring the potential connection between visual and language modalities. Therefore, the accuracy of the image-text semantic similarity calculation of this method is higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0020] Figure 1 This is a flowchart of the method of Embodiment 1 of the present invention; Figure 2 This is a cross-mapping network structure diagram. DETAILED DESCRIPTION

[0021] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0022] Example 1 Combination Figure 1 As shown, this embodiment provides a cross-modal image and text retrieval method based on a multi-window attention mechanism. When performing cross-modal image and text retrieval, the user inputs query data through a browser. After the server obtains the user input data, it first determines the data type, that is, whether it is image data or text data, and then performs data preprocessing and feature extraction on the obtained data respectively. Then, cross-alignment between modalities is performed through the cross-mapping alignment network of the present invention. Finally, the retrieval result is obtained through the result of the similarity measurement and displayed to the user.

[0023] The present invention can realize the tasks of searching for text with images and searching for images with text in Internet resource retrieval. In the retrieval process, different processing branches are matched according to the type of search data (image or text), and a deep learning image-text similarity matching model (i.e., the model of the present invention) is used to complete data processing, feature extraction, modality alignment, similarity matching, and the like. When obtaining the retrieval results, a search engine is used to retrieve the image-text representation containing semantic information, and recall the Top-N semantically similar samples. The present invention can fully extract the interactive features and semantic information of image and text descriptions. Compared with previous image-text retrieval methods, the present invention retains the representation characteristics of the deep learning semantic model vector and improves the accuracy of image and text semantic retrieval.

[0024] The following is a detailed introduction to the processing process of each step.

[0025] S1. Data preprocessing and embedding: The images and texts are preprocessed, and the pre-trained visual model and language model are used for image block embedding and word embedding to convert the images and texts into vector matrices.

[0026] In step S1, it is divided into two branches, namely the visual branch and the text branch. In the visual branch, the image is first preprocessed to evenly divide an image into several image blocks, and then the image blocks are embedded to convert the image blocks into vector matrices; in the text branch, the text is preprocessed and word embedding is performed using a pre-trained language model to convert the text into a vector matrix.

[0027] As an example, in this embodiment, in the visual branch, for each resolution The image is evenly cut into non-overlapping local blocks, each of which is considered a "token". In the implementation, the initial dimension of each block is After a linear embedding layer, we get Dimensional local region representation ,in .

[0028] S2, image and text feature extraction: For the image vector matrix obtained by S1, the Transformer module based on the multi-window attention mechanism is used to extract the local and global features of the image block to obtain the visual feature matrix .

[0029] As a preferred implementation, in step S2, in the visual branch, the image block vector matrix obtained in S1 is processed by a Transformer module based on a multi-window attention mechanism to extract local and global features of the image block, and a fused feature matrix is ​​obtained as the visual feature matrix , and output.

[0030] The multi-window attention mechanism includes local window attention and row and column window attention. The specific processing steps are as follows: S2-1. Use local window attention to extract local features of image blocks to ensure continuity between blocks.

[0031] (1) Image block rearrangement: After the original image is segmented, a group of one-dimensionally arranged image blocks are obtained. First, according to the position information of the segmented image blocks in the original image, the segmented image blocks are rearranged so that their relative position relationship in the original image is maintained; (2) Local window selection: Use a sliding window of preset size to select a local image block on the rearranged image block; (3) Local attention calculation: When calculating attention, only the relationship between the anchor image block and other image blocks in the window is considered, and the relationship between image blocks outside the window is ignored; (4) Feature extraction: Based on the calculated attention, weighted summation is performed on the local image blocks to obtain local features.

[0032] As an example, in this embodiment, the patch sequence is sorted according to the original position information. Perform a reshape operation to restore the relative arrangement to the same state as the original image. Then use The sliding window of size selects the local image block ,in , and calculate the local attention. When calculating the attention, only the anchor point is retained The relationship between the other patches is masked, which simplifies the attention calculation process. The overall attention calculation is as follows: : ; in, , , are the query matrix, key matrix and value matrix, is the number of blocks in the window. The relative position deviation of each head is included in the calculation of similarity. , the bias matrix ), The values ​​in are taken from .

[0033] S2-2. Use row and column window attention to extract contextual features of image blocks as global features to expand the field of view.

[0034] The row and column window attention includes row window attention and column window attention. The row window mainly focuses on the contextual relationship of the image block in the horizontal direction, and the column window mainly focuses on the contextual relationship of the image block in the vertical direction. The two together form a global attention to the image block in the two-dimensional space including the horizontal and vertical directions. The division of the row and column windows expands the window size in the horizontal and vertical directions with the current image block (i.e., the anchor image block) as the center. Since the row and column windows mainly focus on the importance of the anchor image block in the context, the relationship between the anchor image block and other image blocks in the window needs to be focused on, while the relationship between other image blocks is not important. Therefore, in order to reduce unnecessary calculations, the attention between them is masked.

[0035] Specifically, the calculation process of row and column window attention in step S2-2 is as follows: (1) Row window attention extraction: With the anchor image block as the center, the window is expanded horizontally to form a row window; the attention weights between the anchor image block and other image blocks in the row window are calculated; the attention between non-anchor image blocks in the row window is masked; (2) Column window attention extraction: With the anchor image block as the center, the window is expanded in the vertical direction to form a column window; the attention weights between the anchor image block and other image blocks in the column window are calculated; and the attention between non-anchor image blocks in the column window is masked.

[0036] The calculation of row and column window attention is the same as that of local window attention, so we will not repeat it here. We can get the global feature .

[0037] S2-3. Feature fusion: The global features extracted by row window attention and column window attention Local features extracted with local window attention The fused features are used as the final feature representation of the image, that is, the visual feature matrix : .

[0038] For the text vector matrix obtained by S1, the Bert model is used to extract text features to obtain the text feature matrix .

[0039] As a preferred implementation, in the text branch of S2, the word vector matrix obtained in S1 is used to extract text features using Bert; in the text branch, the deep semantic representation of the text is learned based on the pre-trained BERT model, and the feature representation of the text is: ,in, Indicates the number of words in the text. Indicates The characteristics of the word, , Represents the dimension of the word feature vector; after a fully connected layer, the dimension of the text feature vector is unified to the same dimension as the image feature, and the converted text feature is represented as , that is, the text feature matrix .

[0040] S3, cross-modal semantic alignment and semantic similarity calculation: The visual feature matrix obtained by S2 and the text feature matrix Cross-modal semantic alignment is performed, and the cross-mapping alignment network (CACM) is used to obtain the fine-grained semantic relationship matrix between image blocks and text words; the mapping representation and semantic alignment between the two modalities are realized. Then, based on the obtained fine-grained semantic relationship matrix, the semantic similarity between the image and the text is calculated by cosine similarity.

[0041] Combination Figure 2 As shown in the figure, a cross-mapping alignment network (CACM) is used to achieve mapping representation and semantic alignment between two modalities. Including mapping representation between modalities, cross-modal cross alignment and mapping alignment, through self-attention and cross-attention mechanisms, feature extraction and interaction between different modal data are achieved, and through the mapping and alignment process, data of different modalities can be processed and analyzed in a unified feature space.

[0042] As a preferred implementation, in S3, the specific steps are as follows: S3-1. Mapping representation between visual and textual modalities: The visual feature matrix Transformed into a text feature matrix through a linear embedding layer The d-dimensional vector of the same dimension converts the converted visual feature matrix and text feature matrix A connection operation is performed to obtain a new feature matrix Z as the input of the cross-mapping alignment network, and the attention coefficient matrix S between the two modalities is calculated; the attention coefficient matrix S is used to calculate the weighted feature matrix and the mapping feature matrix.

[0043] Specifically, the attention coefficient matrix in step S3-1 is calculated as follows: Through the attention mechanism Get the complete attention coefficient matrix S between the two modalities: ; Mapping representation refers to representing one modality based on another modality. Specifically, it uses a collection of text segments as the representation of visual segments, and vice versa, where Represents the mapping representation of vision to text, Each line of is regarded as a representation of a text segment by a certain visual segment. Represents the mapping of text to vision, Each column of is regarded as a representation of a certain text segment set to a visual segment; and Do matrix multiplication to get the visual attention matrix The mapping matrix : ; Depend on and The weighted visual feature matrices are obtained respectively and the mapping feature matrix : ; ; Where V is the basis matrix used for linear transformation; Similarly, we get the text attention matrix The mapping matrix , weighted text feature matrix and the mapping feature matrix : ; ; .

[0044] The attention coefficient matrix and mapping feature matrix obtained by the above method constitute the fine-grained semantic relationship matrix of image blocks and text words, which reflects the fine-grained semantic relationship between image blocks and text words.

[0045] S3-2. Cross alignment and mapping alignment, as well as semantic similarity of image patches and text.

[0046] The weighted feature matrix and the mapping feature matrix obtained in step S3-1 are aligned. Different from the previous work, the present invention not only adopts cross-modal alignment but also adopts intra-modal mapping alignment.

[0047] Inter-modal cross-alignment: Directly match the fine-grained relationship between image patches and text words through a cross-modal attention matrix; Intra-modal mapping alignment: Use the attention scores of the other modality to reconstruct the feature representation of the own modality and enhance intra-modal consistency (for example, correcting image features with text attention).

[0048] Specifically, in step S3-2, according to the obtained weighted visual feature matrix and the text feature matrix , perform alignment operations as follows: (1) Cross-alignment between modes: Inter-modal alignment seeks fine-grained associations between images and text. For each segment in vision or text, cross-alignment uses cross-modal attention to find the most relevant segment from the opposite modality; and Indicates images and The number of tokens in a text, and the corresponding encoding feature is and .

[0049] Cross-modal alignment is achieved by calculating token-level similarity and object-level similarity.

[0050] Calculating image-text similarity: measuring the token level The image and The similarity of the text. Visual token , calculate it with all text tokens similarity, and use the largest one as its Token-level similarity of sentences .

[0051] Then, based on this, we calculate the similarity at the object level and get The image and Similarity of texts : .

[0052] Calculate text-image similarity: Similarly, the method of calculating image-text similarity is used to obtain the first The text sub- The similarity of images .

[0053] (2) Intra-modal mapping alignment: Intra-modal mapping alignment is to find fine-grained semantic associations between images and texts from another perspective. It is implemented based on the original feature representation and the mapped feature representation. It uses the attention mechanism to map data from different modalities and measures the correlation between the two through the consistency relationship.

[0054] For intra-visual modality alignment, the consistency of the original visual features and the mapped features is calculated. is the original feature representation, It is a mapping feature representation obtained based on the attention coefficient matrix of the language modality. The two are descriptions of the image from the visual perspective and the language perspective respectively. The strength of their consistency relationship represents the strength of the correlation between the image and the text. images and The similarity of texts is expressed as: ; ; in, It is The original feature representation of the image Tokens, By The text obtained The mapping representation of an image is Indicates images and The intermediate result of the similarity measurement between texts, Indicates the final images and is the similarity measure between texts, and n represents the number of tokens in the visual modality.

[0055] Similarly, for text intra-modal alignment, the consistency between the original visual features and the mapped features is calculated. The text and The similarity of images is expressed as: ; ; in, It is The original feature representation of the text Tokens, By The image obtained The mapping representation of the text, Indicates The text and The intermediate result of the similarity measurement between images, Indicates the final The text and is the similarity measure between images, and m represents the number of tokens in the language modality.

[0056] S3-1. Final similarity calculation: In summary, we obtain the inter-modal alignment and intra-modal alignment of image and text, and the final similarity between image and text is expressed as: ; Similarly, the final similarity between text and image is expressed as: .

[0057] S4. Output search results: The results of cross-modal cross-alignment and intra-modal mapping alignment are comprehensively considered according to the similarity.

[0058] Example 2 This embodiment is based on the cross-modal image-text retrieval method provided in Example 1, and experiments are conducted based on Flickr30K and MSCOCO image-text datasets. Flickr30K and MSCOCO are commonly used benchmark datasets in cross-modal image-text retrieval tasks. The Flickr30K dataset is a image-text dataset consisting of image-text pairs, including 31,000 images, each image corresponds to 5 annotated titles, and the content is mainly about humans and animals, with a wide variety of types and a wide range. The MSCOCO dataset is also a large-scale image-text dataset widely used in image-text matching tasks.

[0059] In the cross-modal image-text matching task, recall at K is usually used as a model evaluation indicator. In order to refine the comparison results of the model, this paper uses R@K (K is 1, 5, and 10) and Rsum (the sum of the first three) as model evaluation indicators. The higher the R@K value, the better the model performance.

[0060] In order to verify the effectiveness of this method, two sets of comparative experiments were set up. Experiment 1: Comparative experiment with the classic model. Experiment 2: Comparative experiment in which the multi-window attention mechanism and the cross-mapping network are successively integrated into the experimental process.

[0061] In Experiment 1, the model of the present invention is compared with classic models such as SCAN and MMCA. The experimental results on the Flickr30K and MSCOCO image and text datasets are shown in Tables 1 and 2, respectively.

[0062] Table 1 Experimental results on Flickr30K

[0063] Table 2 Experimental results on MSCOCO

[0064] Table 1 shows the results of each model on the Flickr30K test set. In the R@1 and R@5 indicators in the image-text retrieval and text-image retrieval tasks, the present invention can obtain the best retrieval results, and the Rsum indicator is improved by 1.6 percentage points compared with the most advanced FILIP, which verifies the superiority of this method. Unlike the CNN+RNN two-stream model, which has model differences, the Transformer+Transformer two-stream model eliminates the model differences between the visual model and the language model, and the attention mechanism combined with the FCN structure enables the visual model to integrate excellent context modeling capabilities and hierarchical multi-resolution feature extraction capabilities, so it can achieve better results in the R@K indicator, especially in the R@1 indicator.

[0065] Table 2 shows the results of each model on the MSCOCO test set. In the Rsum index of the image-text retrieval and text-image retrieval tasks, the performance of the present invention reached the highest 436.2. Moreover, it surpassed the FILIP model based on the Transformer visual model in almost all indicators, reflecting the importance of the present invention in exploring the relationship between text and image based on the Transformer visual model. The use of a multi-window attention mechanism to fuse the global and local features of the image can better capture the multi-granularity features of the image than the double-layer local attention. The fusion of cross-alignment and mapping alignment can make the cross-modal alignment of images more complete than the use of cross-alignment alone, so that the model has better matching performance. Compared with other models that only use the method of cross-alignment between modalities, the inter-modal and intra-modal alignment methods used by the model in this chapter are more conducive to making full use of information. At the same time, the cross-alignment between modalities and the mapping alignment within modalities are balanced with each other, reducing the impact of possible misalignment in a single alignment process on the image-text matching effect, and increasing the robustness of the model.

[0066] In Experiment 2, a comparative experiment of multi-window attention mechanism and cross-mapping network was successively integrated into the experimental process. The experimental results are shown in Table 3.

[0067] Table 3 Ablation experiment results

[0068] As can be seen from Table 3, on the pest and disease dataset, the model performs better by using the multi-window attention mechanism and the cross-modal interactive guidance network method. The comparison of the results of MAMR-WRC and MAMR-NONE shows that the addition of multi-window attention can effectively improve the performance of the model. It can help the visual model obtain richer semantic information of agricultural pest and disease images and improve the feature quality of the visual modality in cross-modal alignment, which is more conducive to the semantic alignment of images and texts. The comparison of the results of MAMR-CAMA and MAMR-NONE shows that compared with only using inter-modal cross alignment, intra-modal and inter-modal alignment using cross alignment and mapping alignment is more beneficial to improving the performance of the model. The results of MAMR-BOTH show that the fusion of multi-window attention mechanism and cross-mapping alignment network can more significantly improve the performance of the model in image-text retrieval.

[0069] Experimental summary: On the Flickr30K dataset, Rsum reaches 555.1, an increase of 1.6% over FILIP; on the MSCOCO dataset, Rsum reaches 436.2, a significant improvement over the baseline model (Table 1, Table 2). The ablation experiment (Table 3) shows that the fusion of multi-window attention and CACM makes the performance improvement most obvious (Rsum increases from 535.0 to 541.1).

[0070] Comparison conclusion: The multi-window attention mechanism is more adaptable to the spatial characteristics of images than single-window or global attention, especially in fine-grained alignment tasks. The dual alignment mechanism (cross-modal + intra-modal) makes up for the shortcomings of single alignment and verifies the importance of modal interaction and feature consistency.

[0071] In summary, the present invention (1) proposes a multi-window attention mechanism. In view of the limitations of the prior art: the global self-attention of the traditional Transformer has high computational complexity and it is difficult to take into account both local details and global structures; although the existing cross-modal methods (such as SCAN and MMCA) introduce local attention, they lack modeling of the two-dimensional spatial relationship of the image (such as long-distance dependence in the row and column directions). The present invention proposes: Local window attention: Limit the attention range by sliding the window, retain the spatial continuity of adjacent image blocks, and reduce the amount of calculation. Row and column window attention: Expand the window in the horizontal and vertical directions to model the two-dimensional contextual relationship of the image and make up for the field of view limitations of the local window. It integrates local and global information in visual feature extraction, which is better than the traditional single window or global attention mechanism (such as ViT and Swin Transformer).

[0072] (2) A cross-mapping alignment network (CACM) was designed. In view of the limitations of existing technologies: existing methods (such as SCAN) only focus on inter-modal alignment and ignore intra-modal feature consistency; two-stream models (such as CLIP) lack modal interaction, and single-stream models (such as FILIP) cannot be processed offline. The present invention proposes: inter-modal cross-alignment: directly match the fine-grained relationship between image blocks and text words through a cross-modal attention matrix. Intra-modal mapping alignment: use the attention score of the other modality to reconstruct the feature representation of the own modality and enhance intra-modal consistency (for example, use text attention to correct image features). Through the dual alignment mechanism, noise alignment (such as mismatching of irrelevant background and text) is reduced, and the robustness of semantic alignment is improved.

[0073] (3) Optimization through dual-stream model. In view of the limitations of existing technologies: the traditional dual-stream model has insufficient interaction due to the independent design of the modal encoder; the single-stream model has high computational cost and cannot be applied offline. The present invention proposes: the visual branch adopts multi-window Transformer, the language branch adopts fine-tuned BERT, and the feature dimension is unified through the fully connected layer to balance the computational efficiency and interaction depth. Combining offline feature extraction (dual-stream advantage) with online cross-modal alignment (single-stream advantage), it takes into account both real-time performance and accuracy.

[0074] (4) This invention combines local window and row-column window attention with cross-modal alignment for the first time. Through the multi-window attention mechanism and cross-mapping alignment network, the attention of the other modality is used to reconstruct the own features, enhance cross-modal semantic consistency, and solve the shortcomings of existing methods in local-global feature fusion. In addition, this invention balances the offline processing capability of the dual-stream model with the interaction depth of the single-stream model, and is suitable for real-time retrieval scenarios (such as public security and medical image retrieval).

[0075] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Any changes, modifications, additions or substitutions made by ordinary technicians in this technical field within the essential scope of the present invention should fall within the protection scope of the present invention.

Claims

1. A cross-modal image-text retrieval method based on a multi-window attention mechanism, characterized in that: The following steps are involved: S1. Data preprocessing and embedding: Preprocess images and texts, and use pre-trained visual models and language models to perform image block embedding and word embedding to convert images and texts into vector matrices; S2, image and text feature extraction: For the image vector matrix obtained by S1, the Transformer module based on the multi-window attention mechanism is used to extract the local and global features of the image block to obtain the visual feature matrix ; For the text vector matrix obtained by S1, use the Bert model to extract text features and obtain the text feature matrix ; S3, cross-modal semantic alignment and semantic similarity calculation: The visual feature matrix obtained by S2 and the text feature matrix Perform cross-modal semantic alignment and use a cross-mapping alignment network to obtain a fine-grained semantic relationship matrix between image blocks and text words; achieve mapping representation and semantic alignment between the two modalities; then calculate the semantic similarity between the image and text through cosine similarity based on the obtained fine-grained semantic relationship matrix; S4. Output search results: The results of cross-modal cross-alignment and intra-modal mapping alignment are comprehensively considered according to the similarity.

2. The cross-modal image-text retrieval method based on multi-window attention mechanism according to claim 1 is characterized in that: In step S1, it is divided into two branches, namely the visual branch and the text branch. In the visual branch, the image is first preprocessed to evenly divide an image into several image blocks, and then the image blocks are embedded to convert the image blocks into vector matrices; in the text branch, the text is preprocessed and word embedding is performed using a pre-trained language model to convert the text into a vector matrix.

3. The cross-modal image-text retrieval method based on multi-window attention mechanism according to claim 2 is characterized in that: In step S2, in the visual branch, the image block vector matrix obtained in S1 is processed by a Transformer module based on a multi-window attention mechanism to extract local and global features of the image block and obtain a fused feature matrix as the visual feature matrix , and output; The multi-window attention mechanism includes local window attention and row and column window attention. The specific processing steps are as follows: S2-1. Extract local features of image blocks using local window attention: (1) Image block rearrangement: After the original image is segmented, a group of one-dimensionally arranged image blocks are obtained. First, according to the position information of the segmented image blocks in the original image, the segmented image blocks are rearranged so that their relative position relationship in the original image is maintained; (2) Local window selection: Use a sliding window of preset size to select a local image block on the rearranged image block; (3) Local attention calculation: When calculating attention, only the relationship between the anchor image block and other image blocks in the window is considered, and the relationship between image blocks outside the window is ignored; (4) Feature extraction: Based on the calculated attention, weighted summation is performed on the local image blocks to obtain local features; S2-2. Use row and column window attention to extract contextual features of image blocks as global features: The row and column window attention includes row window attention and column window attention. The division of row and column windows is to expand the window size horizontally and vertically with the current image block, i.e., the anchor image block, as the center. The calculation of row and column window attention is the same as that of local window attention. S2-3. Feature fusion: The global features extracted by row window attention and column window attention are fused with the local features extracted by local window attention. The fused features are used as the final feature representation of the image, namely the visual feature matrix .

4. The cross-modal image-text retrieval method based on multi-window attention mechanism according to claim 3 is characterized in that: The calculation process of row and column window attention in step S2-2 is as follows: (1) Row window attention extraction: With the anchor image block as the center, the window is expanded horizontally to form a row window; the attention weights between the anchor image block and other image blocks in the row window are calculated; the attention between non-anchor image blocks in the row window is masked; (2) Column window attention extraction: With the anchor image block as the center, the window is expanded in the vertical direction to form a column window; the attention weights between the anchor image block and other image blocks in the column window are calculated; and the attention between non-anchor image blocks in the column window is masked.

5. The cross-modal image-text retrieval method based on multi-window attention mechanism according to claim 2, characterized in that: In the text branch of S2, the word vector matrix obtained in S1 is used to extract text features using Bert. In the text branch, the deep semantic representation of the text is learned based on the pre-trained BERT model, and the feature representation of the text is: ,in, Indicates the number of words in the text. Indicates The characteristics of the word, , Represents the dimension of the word feature vector; after a fully connected layer, the dimension of the text feature vector is unified to the same dimension as the image feature, and the converted text feature is represented as , that is, the text feature matrix .

6. The cross-modal image-text retrieval method based on multi-window attention mechanism according to claim 1, characterized in that: In the S3, the specific steps are as follows: S3-1. Mapping representation between visual and textual modalities: The visual feature matrix Transformed into a text feature matrix through a linear embedding layer Vectors of the same dimension, converting the visual feature matrix and text feature matrix Perform a connection operation to obtain a new feature matrix Z as the input of the cross-mapping alignment network, and calculate the attention coefficient matrix S between the two modalities; Use the attention coefficient matrix S to calculate the weighted feature matrix and the mapping feature matrix; S3-2, cross alignment and mapping alignment: Performing an alignment operation based on the weighted feature matrix and the mapped feature matrix obtained in step S3-1 includes: Inter-modal cross-alignment: Directly match the fine-grained relationship between image patches and text words through a cross-modal attention matrix; Intra-modal mapping alignment: Use the attention score of the other modality to reconstruct the feature representation of the own modality and enhance intra-modal consistency; S3-3. Final similarity calculation: The obtained attention coefficient matrix and mapping feature matrix constitute a fine-grained semantic relationship matrix of image blocks and text words, which is used to calculate similarity.

7. The cross-modal image-text retrieval method based on multi-window attention mechanism according to claim 6, characterized in that: The calculation of the attention coefficient matrix in step S3-1 is as follows: Through the attention mechanism Get the complete attention coefficient matrix S between the two modalities: ; Mapping representation refers to representing one modality based on another modality. Specifically, it uses a collection of text segments as the representation of visual segments, and vice versa. Represents the mapping representation of vision to text, Each line of is regarded as a representation of a text segment by a certain visual segment. Represents the mapping of text to vision, Each column of is regarded as a representation of a certain text segment set to a visual segment; and Do matrix multiplication to get the visual attention matrix The mapping matrix : ; Depend on and The weighted visual feature matrix is ​​obtained respectively and the mapping feature matrix : ; ; Where V is the basis matrix used for linear transformation; Similarly, we get the text attention matrix The mapping matrix , weighted text feature matrix and the mapping feature matrix : ; ; 。 8. The cross-modal image-text retrieval method based on multi-window attention mechanism according to claim 7, characterized in that: In step S3-2, according to the obtained weighted visual feature matrix and the text feature matrix , perform alignment operations as follows: (1) Cross-alignment between modes: Inter-modal alignment seeks fine-grained associations between images and text. For each segment in vision or text, cross-alignment uses cross-modal attention to find the most relevant segment from the opposite modality; including: Calculating image-text similarity: measuring the token level The image and Then, based on this, we calculate the similarity of the object level and get the The image and Similarity of texts ; Calculate text-image similarity: Similarly, the method of calculating image-text similarity is used to obtain the first The text sub- The similarity of images ; (2) Intra-modal mapping alignment: Intra-modal mapping alignment is to find the fine-grained semantic association between images and texts from another perspective. It is implemented based on the original feature representation and the mapped feature representation. The attention mechanism is used to map the data of different modalities, and the correlation between the two is measured through the consistency relationship. For the visual modality, is the original feature representation, It is a mapping feature representation obtained based on the attention coefficient matrix of the language modality. The two are descriptions of the image from the visual perspective and the language perspective respectively. The strength of their consistency relationship represents the strength of the correlation between the image and the text. images and The similarity of texts is expressed as: ; ; in, It is The original feature representation of the image Tokens, By The text obtained The mapping representation of an image is Indicates images and The intermediate result of the similarity measurement between texts, Indicates the final images and The similarity measure between the texts, n represents the number of tokens in the visual modality; Similarly, within the language mode, The text and The similarity of images is expressed as: ; ; in, It is The first Tokens, By The image obtained The mapping representation of the text, Indicates The text and The intermediate result of the similarity measurement between images, Indicates the final The text and The similarity measure between images, m represents the number of tokens in the language modality; S3-3. Final similarity calculation: The final similarity between image and text is expressed as: ; Similarly, the final similarity between text and image is expressed as: 。

Citation Information

Patent Citations

  • Cross-modal image-text retrieval method based on multi-level semantic alignment

    CN116821391A

  • Gradual semantic aggregation and structured cognitive enhancement-based image-text matching method

    CN119397048A

Cited By

  • Multi-modal large model alignment method and system based on language perception and feature fusion

    CN120932247A

  • Multimodal large model alignment method and system based on language perception and feature fusion

    CN120932247B

  • Multi-modal retrieval method combining image features and semantic understanding

    CN121858757A

  • A Multimodal Data Fusion Method for Human Body Following in a Four-Wheel Independent Steering Vehicle

    CN122569523A