Image-text retrieval method based on cross-modal semantic alignment

Updated area-level image features through semantic enhancement module and adaptive attention factor, solving the problem of inaccurate regional relationship modeling in image-text retrieval, achieving more efficient image-text matching accuracy.

CN120296186AActive Publication Date: 2025-07-11STATE GRID ANHUI ULTRA HIGH VOLTAGE CO +1
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510363596.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-11
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

In the prior art, in image-text retrieval, there are problems in modeling object area relationships and ignoring network feedback adjustment capabilities, resulting in impairment of the accuracy of image-text matching.

Method used

Semantic enhancement module is used to process regional image features, and the correlation between regional image features and word features in text sentences is updated by constructing adaptive attention factor, and a similarity score is calculated using a self-attention algorithm.

Benefits of technology

It significantly improves the accuracy of regional image feature coding and word-region image matching effect, and improves the accuracy of image-text retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296186A_ABST
    Figure CN120296186A_ABST
Patent Text Reader

Abstract

The invention discloses an image-text retrieval method based on cross-modal semantic alignment, which comprises the following steps: firstly, converting a whole image into a group of region-level image features, and then processing the region-level image features by using a semantic enhancement module to obtain enhanced region-level image features; then updating the region-level image feature related to a certain word feature in the text sentence in the region-level image features through the two adaptive attention factors; calculating the similarity between each word feature in the text sentence and each region-level image feature related to the word feature, and calculating to obtain a similarity score between the whole text sentence and the whole image; and finally, according to the steps, when the text sentences or images are queried, L pictures or L text sentences with the highest similarity score with the text sentences or images in the database are retrieved as retrieval results. According to the method, the region-level image features can be coded more accurately, and the word-region-level image matching process is remarkably promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graphic and text data processing, and specifically to an image-text retrieval method based on cross-modal semantic alignment. Background Art

[0002] Image-text retrieval is a fundamental and challenging task that bridges the fields of language and vision. Previous work has mainly addressed cross-modal image-text retrieval problems in the following directions: 1) Coarse-grained retrieval methods directly calculate the global similarity between the input image and the entire text by mapping two heterogeneous modalities to a common embedding space; 2) Fine-grained matching methods automatically align image regions and text segments by exploring the fine-grained cross-modal correspondence between image regions and text segments.

[0003] The general process of fine-grained image-text matching is as follows: 1) Image and text feature representation. 2) Design of a cross-modal attention module. 3) Design of a loss function. The most classical method is SCAN (Stacked Cross Attention Network), which infers image-text similarity by discovering all potential alignment relationships between image regions and words in a sentence. This method can capture the fine-grained interaction between vision and language, making image-text matching more interpretable. Subsequently, image-text retrieval models have basically continued the main idea of SCAN.

[0004] However, most of these methods have the following problems: (1) Modeling the relationships between object regions and directly performing local-level or global-level matching, where each region is treated equally, which will damage the accuracy of the representation; (2) Adopting a one-time forward association or aggregation strategy with a complex architecture or additional information, while ignoring the regulatory ability of network feedback. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide an image-text retrieval method based on cross-modal semantic alignment, which can more accurately encode region-level image features and significantly promote the word-region-level image matching process.

[0006] The technical solution of the present invention is as follows:

[0007] An image-text retrieval method based on cross-modal semantic alignment specifically includes the following steps:

[0008] (1) First, convert the entire image into a set of region-level image features, and then use a semantic enhancement module to process the region-level image features to obtain enhanced region-level image features;

[0009] (2) Construct two adaptive attention factors, namely an adaptive channel weight vector for correcting the correlation between region-level image features and word features in a text sentence, and a weight vector for adjusting the attention distribution. Then, update the region-level image features related to a certain word feature in the text sentence in the region-level image features through the two adaptive attention factors;

[0010] (3) Calculate the similarity between each word feature in the text sentence and each region-level image feature related to it, and then calculate the similarity score between the entire text sentence and the entire image;

[0011] (4) According to the above steps (1)-(3), when querying a text sentence, retrieve the top L images in the database with the highest similarity score to this text sentence as the retrieval result; or when querying an image, retrieve the top L text sentences in the database with the highest similarity score to this query image as the retrieval result; L is an integer not less than 1.

[0012] The specific steps of first converting the entire image into a set of region-level image features and then using a semantic enhancement module to process the region-level image features to obtain enhanced region-level image features are as follows:

[0013] S11. First, convert the entire image I into a set of region-level image features where, m i represents a region image feature of the entire image I, and K represents the number of region image features in the entire image I;

[0014] S12. Use the semantic enhancement module to process each region-level image feature. First, input a set of region-level image features M, and perform initial global embedding using a region-level feature mapping function local(·) and a global embedding mapping function global(·), as shown in the following formula (1):

[0015]

[0016] In formula (1), W l and W g are learnable parameter matrices;

[0017] S13. Aggregate the initial global embeddings of all region-level image features through element-wise multiplication of matrices and the softmax function, as shown in the following formula (2), and then calculate the global embedding m g ;

[0018]

[0019] In formulas (2) and (3), represents the combination of the Tanh activation function and the batch normalization function, softmax(·) represents the softmax activation function, ⊙ represents element-wise matrix multiplication, and W a is a learnable parameter matrix, and w i is the weight coefficient, representing the correlation between the region-level image feature m i and the initial global embedding, ‖‖2 represents the L2 norm, i.e., the Euclidean distance;

[0020] S14. Obtain the enhanced region-level image features of the entire image by performing the self-attention algorithm Specifically, see the following formula (4):

[0021]

[0022] In formula (4), x i represents the region-level image feature m i The region-level image feature after semantic enhancement.

[0023] The construction of the two adaptive attention factors specifically includes the following steps:

[0024] S21. Input the enhanced region-level image features of the entire image A set of word features of the text sentence E where t j represents the j-th word feature, and N represents the number of word features in the text sentence;

[0025] S22. For the word feature t j and the region-level image feature related to the word feature t j whose initial value is the region-level image feature after semantic enhancement, first construct the alignment vector a to encode the difference relationship between t j and j and Specifically, see the following formula (5):

[0026]

[0027] In formula (5), Norm is the vector normalization operation, and linear is the linear transformation;

[0028] S23. The alignment vector a j is used to learn the two adaptive attention factors. The two adaptive attention factors are respectively the adaptive channel weight vector p i for correcting the correlation between the region-level image feature x j and the word feature t in the text sentence j and the weight vector q j, the calculation formulas are shown in the following formulas (6) and (7):

[0029] p j = tanh(W p′ (tanh(W p a j ))) (6);

[0030] q j = [W q′ (tanh(W q a j ))] + (7);

[0031] In formulas (6) and (7), W p , W p′ , W q and W q′ are learnable parameter matrices, tanh(·) represents the tanh activation function, and [·] + represents taking the larger value of the variable and 0.

[0032] Updating the region-level image feature related to a certain word feature in the text sentence by two adaptive attention factors, specifically as shown in the following formula (8):

[0033]

[0034] In formula (8), represents the region-level image feature related to the word feature t j ; ‖‖ represents the L1 norm, that is, the Manhattan distance; ⊙ represents the exclusive NOR operation.

[0035] The specific steps for calculating the similarity between each word feature in the text sentence and each related region-level image feature, and then calculating the similarity score between the entire text sentence and the entire image are as follows:

[0036] S31. First, calculate the cosine similarity between all word features t j in the text sentence and each region-level image feature j in the updated region-level image features related to t , and perform softmax normalization processing. The processing process is shown in the following formula (9):

[0037]

[0038] In formula (9), measures the similarity between the i-th region-level image feature and the j-th word feature, and [·] + represents taking the larger value of the variable and 0;

[0039] S32. Calculate the similarity score S between the entire image I and the entire text sentence E through average pooling AVG (I, E), and the calculation process is shown in the following formula (10):

[0040]

[0041] Advantages of the present invention:

[0042] (1). The present invention uses a semantic enhancement module to process regional image features, obtains enhanced regional image features, and enhances regional image features through global representation. Compared with directly performing regional or global matching, it can encode image visual features more accurately.

[0043] (2). The present invention updates the regional image features related to the feature of a certain word in the text sentence in the regional image features through two adaptive attention factors, explores the regulatory ability of the network itself, effectively updates the regional image features related to the word features, and thus significantly promotes the word-regional image matching process. Brief Description of the Drawings

[0044] Figure 1 is a flowchart of the present invention. Detailed Embodiments

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0046] An image-text retrieval method based on cross-modal semantic alignment specifically includes the following steps:

[0047] (1). First, convert the entire image into a set of regional image features, and then use a semantic enhancement module to process the regional image features to obtain enhanced regional image features. The specific steps are as follows:

[0048] S11. First, convert the entire image I into a set of regional image features where m i represents a regional image feature of the entire image I, and K represents the number of regional image features in the entire image I;

[0049] S12. Process each regional image feature using the semantic enhancement module. First, input a set of regional image features M, and perform initial global embedding using the regional feature mapping function local(·) and the global embedding mapping function global(·), as shown in the following formula (1):

[0050]

[0051] In formula (1), W l and W g are learnable parameter matrices;

[0052] S13. Aggregate the initial global embeddings of all regional image features through element-wise multiplication of matrices and the softmax function, as shown in the following formula (2), and then calculate the global embedding m through the following formula (3): g ;

[0053]

[0054] In formulas (2) and (3), represents the combination of the Tanh activation function and the batch normalization function, softmax(·) represents the softmax activation function, ⊙ represents element-wise multiplication of matrices, W a is a learnable parameter matrix, w i is the weight coefficient, representing the correlation between the regional image feature m i and the initial global embedding, ‖‖2 represents the L2 norm, i.e., the Euclidean distance;

[0055] S14. Obtain the enhanced regional image features of the entire image by performing the self-attention algorithm as shown in the following formula (4):

[0056]

[0057] In formula (4), x i represents the regional image feature m i after semantic enhancement;

[0058] (2). Construct two adaptive attention factors, and then update the regional image features related to the word features in a text sentence through the two adaptive attention factors, specifically including the following steps:

[0059] S21. Input the enhanced regional image features of the entire image a set of word features of the text sentence E where t j represents the j-th word feature, and N represents the number of word features in the text sentence;

[0060] S22. For word feature t j and the region-level image feature j related to word feature t The initial value is the region-level image feature after semantic enhancement. First, construct the alignment vector a j used to encode the difference relationship between t j and as shown in the following formula (5):

[0061]

[0062] In formula (5), Norm is the vector normalization operation, and linear is the linear transformation;

[0063] S23. The alignment vector a j is used to learn two adaptive attention factors, which are the adaptive channel weight vector p i for correcting the correlation between the region-level image feature x j and the word feature t in the text sentence j , and the weight vector q j for adjusting the attention distribution, and the calculation formulas are shown in the following formulas (6) and (7) respectively:

[0064] p j = tanh(W p′ (tanh(W p a j ))) (6);

[0065] q j = [W q′ (tanh(W q a j ))] + (7);

[0066] In formulas (6) and (7), W p , W p′ , W q and W q′ are learnable parameter matrices, tanh(·) represents the tanh activation function, and [·] + represents taking the larger value of the variable and 0;

[0067] S24. Update the region-level image feature related to a certain word feature in the text sentence in the region-level image feature through the two adaptive attention factors, as shown in the following formula (8):

[0068]

[0069] In formula (8), Represents the region-level image feature related to the word feature t j ; ‖‖ represents the L1 norm, i.e., the Manhattan distance; ⊙ represents the exclusive NOR operation;

[0070] (3) Calculate the similarity between each word feature in the text sentence and each related region-level image feature, and then calculate the similarity score between the entire text sentence and the entire image. The specific steps are as follows:

[0071] S31. First, calculate all the word features t in the text sentence j and the updated region-level image features related to t j for each region-level image feature in to calculate the cosine similarity and perform softmax normalization. The process is shown in Equation (9) below:

[0072]

[0073] In Equation (9), measures the similarity between the i-th region-level image feature and the j-th word feature, and [·] + represents taking the larger value of the variable and 0;

[0074] S32. Calculate the similarity score S AVG (I, E) between the entire image I and the entire text sentence E through average pooling. The calculation process is shown in Equation (10) below:

[0075]

[0076] (4) According to the above steps (1)-(3), when querying a text sentence, retrieve the L images with the highest similarity scores to this text sentence in the database as the retrieval results; or when querying an image, retrieve the L text sentences with the highest similarity scores to this query image in the database as the retrieval results; L is an integer not less than 1.

[0077] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An image-text retrieval method based on cross-modal semantic alignment, characterized in that: Specifically, it includes the following steps: (1) First, convert the entire image into a set of region-level image features, and then use a semantic enhancement module to process the region-level image features to obtain enhanced region-level image features; (2) Construct two adaptive attention factors, namely an adaptive channel weight vector for correcting the correlation between region-level image features and word features in the text sentence, and a weight vector for adjusting the attention distribution. Then, update the region-level image features related to a certain word feature in the text sentence through the two adaptive attention factors; (3) Calculate the similarity between each word feature in the text sentence and each region-level image feature related to it, and then calculate the similarity score between the entire text sentence and the entire image; (4) According to the above steps (1)-(3), when querying a text sentence, retrieve the top L images in the database with the highest similarity score to this text sentence as the retrieval result; or when querying an image, retrieve the top L text sentences in the database with the highest similarity score to this query image as the retrieval result; L is an integer not less than 1.

2. The image-text retrieval method based on cross-modal semantic alignment according to claim 1, characterized in that: The specific steps of first converting the entire image into a set of region-level image features and then using a semantic enhancement module to process the region-level image features to obtain enhanced region-level image features are as follows: S11. First, convert the entire image I into a set of region-level image features where m i represents a region image feature of the entire image I, and K represents the number of region image features in the entire image I; S12. Use the semantic enhancement module to process each region-level image feature. First, input a set of region-level image features M, and perform initial global embedding using the region-level feature mapping function local(·) and the global embedding mapping function global(·), as shown in the following formula (1): In formula (1), W l and W f are learnable parameter matrices; S13. Aggregate the initial global embeddings of all regional image features through element-wise multiplication of matrices and the softmax function, as shown in the following formula (2), and then obtain the global embedding m through the calculation of the following formula (3). g ; In formulas (2) and (3), represents the combination of the Tanh activation function and the batch normalization function, softmax(·) represents the softmax activation function, ⊙ represents element-wise matrix multiplication, and W a is a learnable parameter matrix, w i is the weight coefficient, representing the correlation between the region-level image feature m i and the initial global embedding, and ‖‖2 represents the L2 norm, i.e., the Euclidean distance; S14. Obtain the enhanced regional image features of the entire image by performing the self-attention algorithm Specifically, see the following formula (4): In formula (4), x i represents the region-level image feature m i which is the region-level image feature after semantic enhancement.

3. The image-text retrieval method based on cross-modal semantic alignment according to claim 2, wherein: The specific steps of constructing the two adaptive attention factors specifically include the following: S21. Input the region-level image features after enhancing the entire image A set of word features of text sentence E where t j represents the j-th word feature, and N represents the number of word features in the text sentence; S22. For word feature t j and the region-level image feature j related to word feature t whose initial value is the region-level image feature after semantic enhancement, first construct an alignment vector a j used to encode the difference relationship between t j and as shown in the following formula (5): In formula (5), Norm is the vector normalization operation, and linear is the linear transformation; S23, alignment vector a j For learning two adaptive attention factors, the two adaptive attention factors are respectively used to correct the region-level image feature x i and the word feature t in the text sentence j The adaptive channel weight vector p for the correlation between them j , and the weight vector q for adjusting the attention distribution j , and the calculation formulas are shown in the following formulas (6) and (7) respectively: p j = tanh(W p′ (tanh(W p a j ))) (6); q j = [W q′ (tanh(W q a j ))] + (7); In equations (6) and (7), W p , W p′ , W q and W q′ are learnable parameter matrices, tanh(·) represents the tanh activation function, and [·] + represents taking the larger value of the variable and 0.

4. The image-text retrieval method based on cross-modal semantic alignment according to claim 3, wherein: The specific steps of updating the region-level image features related to a certain word feature in the text sentence through the two adaptive attention factors are as shown in the following formula (8): In formula (8), represents the region-level image feature related to the word feature t j ; ‖‖ represents the L1 norm, that is, the Manhattan distance; ⊙ represents the exclusive NOR operation.

5. The image-text retrieval method based on cross-modal semantic alignment according to claim 4, wherein: The specific steps of calculating the similarity between each word feature in the text sentence and each region-level image feature related to it, and then calculating the similarity score between the entire text sentence and the entire image are as follows: S31. First, calculate all word features t in the text sentence j and the updated region-level image features related to t j for each region-level image feature in Calculate the cosine similarity between them and perform softmax normalization. The processing process is shown in the following formula (9):​ In Equation (9), measures the similarity between the i-th regional image feature and the j-th word feature, [·] + denotes taking the larger value of the variable and 0; S32. Calculate the similarity score S between the entire image I and the entire text sentence E through average pooling AVG (I, E), and the calculation process is shown in the following formula (10):

Citation Information

Patent Citations

  • Image text matching method based on double-vision-field semantic reasoning network

    CN111242197A

  • Image-text matching method based on region-enhanced network with topic constraints

    CN112084358A

  • Cross-modal retrieval method based on multilayer semantic alignment

    CN112966127A

  • Semantic-based image-text cross-modal retrieval method

    CN113902764A

  • News event searching method and system based on multistage image-text semantic alignment model

    CN114297473A