A cross-modal retrieval system, method and computer device based on counterfactual reasoning
Through the multi-level counterfactual comparison learning method, a cross-modal retrieval model is constructed, which solves the problem of difficult to extract semantic alignment relationships between data of different modalities, and achieves a more efficient cross-modal retrieval effect.
Patent Information
- Application Number
- CN202210716568.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-06-23
AI Technical Summary
It is difficult for the prior art to effectively extract the structural information and deep semantic correspondence of data of different modalities, and then conduct cross-modal retrieval.
Using a multi-level counterfactual comparison learning method, images and text features are extracted through Faster-RCNN and Bert models, and high-level semantic alignment is used to construct instance-level, image-level and semantic-level counterfactual comparison learning samples, which are integrated into a unified framework for overall training of cross-modal retrieval models.
The model's understanding and reasoning ability of diverse visual content, high-level text semantics and complex cross-modal relationships is improved, more discriminant features are extracted, and the model's semantic alignment ability and retrieval accuracy are improved.
Smart Images

Figure CN115146100B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multimedia computing, and particularly relates to a cross-modal retrieval model, method and computer device based on counterfactual reasoning. Background Art
[0002] With the wide application of artificial intelligence technology in various fields, the forms of data presentation are becoming more and more diverse. Multimodal data such as text, images, and videos has increased sharply. The information of single-modal data is limited, and interactive multimodal data can convey richer information. The same thing will have multiple descriptions in different modal data. These data are "heterogeneous and homologous" in form and semantically related. The diversification of data content forms can help people perceive and understand the surrounding world, because it is easy for people to align and complement different forms of information, so as to learn knowledge more comprehensively and accurately. In the field of artificial intelligence cross-modal, it has brought an urgent need for cross-modal retrieval. Cross-modal retrieval is one of the important applications of cross-modal learning, also known as cross-media retrieval. Its characteristic is that all modal data exists during the training process, but only one modal is available during the testing process. Cross-modal retrieval aims to achieve information interaction between two different modalities, and its fundamental purpose is to explore the relationship between different modal samples, that is, to retrieve another modal sample with approximate semantics through one modal sample. Cross-modal retrieval requires the retrieval set and the query set to be of different modalities, such as using text to search for pictures, pictures to search for videos, etc. How to extract the structural information and its deep semantic correspondence relationship of different modal data, and then model it is the difficulty in improving multimodal retrieval.
[0003] In recent years, with the application of the theory of causality in the field of deep learning, counterfactual methods based on causal reasoning have begun to be used for the deep semantic alignment relationship between different modal data in the multimodal field. Currently, very good results have been achieved in many cross-modal subtasks such as VQA (Visual Question Answering). Causality has strong interpretability, and counterfactual learning based on causality is to assume that there already exists a causal relationship structure, and by controlling variables, the shadow projection results can be obtained each time, so that the distribution of situations that did not exist in history can be modeled, and thus an unbiased estimate can be obtained. Swaminathan et al. first defined a machine learning framework for counterfactual learning in historical records, and promoted and further normalized its model structure in deep learning. In addition, counterfactual learning has also been extended to the fields of representation learning and log learning. Summary of the Invention
[0004] The objective of the present invention is to utilize multi-level counterfactual contrastive learning to facilitate the joint modeling of the model for diverse image content, high-level text semantics, and complex cross-modal relationships, thereby learning more discriminative feature representations. The technical solution for implementing the present invention is as follows:
[0005] A cross-modal retrieval model based on counterfactual reasoning, which is obtained through the following steps:
[0006] In step S1, the Faster-RCNN model and the pre-trained Bert model are respectively used to extract the original image features and text features. After independently mapping the obtained image features and text features to the same dimension, the feature vectors are obtained in the image feature encoder and text feature encoder composed of four layers of transformers respectively. The obtained image feature vectors and text feature vectors are aligned at a high level using a two-layer Transformer, so that the image feature vectors and text feature vectors are mapped to the same common space, and the model is optimized by calculating the loss.
[0007] Then, the image features and the original text are processed using the counterfactual reasoning method. The recognized image region labels are compared with the nouns extracted from the original text to provide a coefficient matrix for constructing positive and negative samples for the following three types of contrastive learning (instance level, image level, semantic level).
[0008] In step S2, counterfactual reasoning is used to construct positive and negative samples at the instance level. After independently mapping the instance-level image features and text features obtained in step S1 to the same dimension, the feature vectors are obtained in the image feature encoder and text feature encoder composed of four layers of transformers respectively. The obtained image feature vectors and text feature vectors are aligned using a two-layer Transformer, so that the image feature vectors and text feature vectors are mapped to a common space to construct contrastive learning at the instance level, enabling the model to perceive the visual objects in the image and learn fine-grained local feature alignment between the image and text, so that the model can perceive the detailed information of the image.
[0009] In step S3, counterfactual reasoning is used to generate positive and negative samples at the image level. After independently mapping the image-level image features and text features obtained in step S1 to the same dimension, the feature vectors are calculated in the image feature encoder and text feature encoder composed of four layers of transformers. The obtained image feature vectors and text feature vectors are processed using a two-layer Transformer, and the image feature vectors and text feature vectors are mapped to a common space to construct contrastive learning at the image level, guiding the model to learn fine-grained global feature alignment between the image and text, so that the model can perceive the global scene information of the image.
[0010] Step S4: Use counterfactual reasoning to generate counterfactual samples at the semantic level of the text. After independently mapping the semantic-level image features and text features obtained in Step S1 to the same dimension, obtain feature vectors in the image feature encoder and text feature encoder composed of four-layer Transformers respectively. Process the obtained image feature vectors and text feature vectors using a two-layer Transformer, map the image feature vectors and text feature vectors to a common space to construct contrastive learning at the semantic level, obtain the final feature vectors, and construct contrastive learning at the image level based on these features to guide the model to learn fine-grained global feature alignment of image-text, enabling the model to have cross-modal semantic relationships.
[0011] Step S5: Integrate all the above parts into a unified framework for overall training of the cross-modal retrieval model.
[0012] Advantages of the present invention:
[0013] (1) The present invention proposes a multi-level counterfactual contrastive learning cross-modal retrieval framework, which improves the model's understanding and reasoning ability for diverse visual content, high-level text semantics, and complex cross-modal relationships.
[0014] (2) Apply the multi-level contrastive learning method based on counterfactual reasoning to cross-modal retrieval, and construct counterfactual contrast samples at the instance level, image level, and semantic level respectively. This enables the model to extract more discriminative features, thereby comprehensively understanding image and language expressions and enhancing the model's semantic alignment ability.
[0015] (3) Through the counterfactual contrastive learning method, the problem of spurious correlation caused by uneven data distribution in the dataset can be alleviated. Brief Description of the Drawings
[0016] Figure 1 is the framework diagram of the cross-modal retrieval model based on counterfactual reasoning of the present invention;
[0017] Figure 2 is the basic cross-modal retrieval model constructed by the present invention.
[0018] Figure 3 is the flowchart of the cross-modal retrieval method based on counterfactual reasoning of the present invention. Detailed Embodiments
[0019] The present invention proposes a cross-modal retrieval model, method and computer device based on counterfactual reasoning. Counterfactual reasoning is used to construct multi-level contrastive learning. Counterfactual positive and negative samples are generated according to the importance of each object node. In the instance-level contrastive learning module, we generate counterfactual samples (negative samples) by masking important object regions on the original image, and the original image is used as the factual sample (positive sample) for contrastive learning. In the image-level counterfactual contrastive learning module, other images in the mini-batch are used as counterfactual samples, and objects with low importance are masked as factual samples, enabling the model to focus on learning the region features highly relevant to images and texts. At the semantic level, counterfactual samples are generated by randomly replacing nouns in the text, and the original text is used as the factual sample. During the process of using counterfactual samples for contrastive learning, it is possible to alleviate the potential spurious correlations and selection biases in the dataset. This enables the model to extract more discriminative features, thereby comprehensively understanding image and language expressions, enhancing the semantic alignment ability of the model, and improving the accuracy of the model. The present invention will be further described below with reference to the accompanying drawings.
[0020] Figure 1 FIG. is a framework diagram of the counterfactual reasoning-based cross-modal retrieval model proposed by the present invention. Counterfactual positive and negative samples for contrastive learning are generated, and multi-level contrastive learning is used to jointly model rich visual information and complex cross-modal relationships by the model. Specifically, the model is obtained through the following steps:
[0021] Step S1, as Figure 1 shown in the text feature extraction module and the image feature extraction module, extract the features of the original picture and the text, thereby providing data support for generating positive and negative samples in the subsequent steps;
[0022] The step S1 further includes the following steps:
[0023] Step S1.1: For each picture G and its corresponding text E representation in the training data, extract its text feature image feature location feature and the object region label of the picture Here D q , D v represents the dimension of the text feature and the image feature, D s represents the number of labeled regions in the picture, that is, the contour regions of all the recognized objects in this picture (not necessarily rectangles). D n represents the number of local regions extracted from the image, D p represents the dimension of the location feature of the local region, D l represents the length of the sentence.
[0024] The image features are extracted by Faster-RCNN to obtain 36 (i.e., D n = 36) regional visual features, which are then concatenated together as the image feature V. The dimension of each regional visual feature is 2048 (i.e., D v = 2048). The position feature P includes the upper-left and lower-right coordinates of each feature region and the area of the region. D p is 5.
[0025]
[0026] where x 1 , y 1 , x 2 , y 2 are the upper-left and lower-right coordinates of the region respectively, and W and H represent the width and height of the picture.
[0027] The text features can be extracted by BERT to obtain 768-dimensional text features (i.e., D q = 768).
[0028]
[0029] T F = FC t (Bert(T))#(3)
[0030] FC v and FC t represent two independent fully connected layers, Bert represents the Bert model, represents concatenating the previous and the following values in the same dimension. T F ∈R De , D e = 1024.
[0031] Construct a basic cross-modal retrieval model that includes an image feature encoder and a text feature encoder each composed of four layers of Transformer, and a two-layer parameter-sharing Transformer structure for high-level semantic feature alignment, as Figure 2 shown. On the basis of this basic cross-modal retrieval model, three types of contrastive learning are added to implement a complete cross-modal retrieval model.
[0032] After processing the image features and text features through equations (2) and (3), the obtained V F and T FInput it into the created basic cross-modal retrieval model, and use the feature vector corresponding to the [CLS] flag bit in the obtained features as the final feature vector. Establish a triplet loss function as the loss function for image and text alignment:
[0033]
[0034] where t and x represent positive samples, and t - , x - represent negative samples, that is, other text features and image features in the same batch. α is a hyperparameter, and [a] + = max(a, 0),
[0035] Step S1.2: Align the image region label I corresponding to the picture with the nouns parsed from the original text to determine which regions of the image the noun objects in the text appear in. Then, through comparison, it can be known which regional features in the picture are important. For example, if there is a description of a dog in the text, then the region where the dog is located in the corresponding picture is important. At this time, if the local features that intersect with this region are masked, then the similarity between the picture features and the text features will definitely be lower than that before masking. This is the principle of counterfactual reasoning. Divide the length and width of the rectangle formed by all local feature regions of the image into 14 equal parts each, and take a total of 196 intersection points. Then, calculate the importance coefficient of the image region feature for the text by counting the number of points falling inside the contour of the important region, and connect all the values to form a coefficient matrix F Dn×1 As shown in Equation (5), where F i The smaller the value, the more important the regional feature V i .
[0036]
[0037] P i represents the position feature of region i, represents the object region label of the picture, E represents the original text, represents the number of points in which region i falls inside the important region, and F i represents the value of the i-th row of.
[0038] Step S2, as shown in Figure 1 the instance-level contrastive learning module in, use the importance coefficient matrix generated in Step S1.2 to construct instance-level contrastive learning, so that the model focuses on the fine-grained information in the visual picture;
[0039] The said Step S2 further includes the following steps:
[0040] Step S2.1: Construct instance-level contrastive learning: Utilize F obtained in S1.2 Dn×1 to judge the importance degree of the local image features V i i∈D n for the current text features. According to the calculation process in S1.2, if the local image feature V i corresponds to F i and the value is 1, it can be considered that this area is not within the important range and is an unimportant area feature. Connect the features of the important areas (the corresponding values in F Dn×1 are not 1) to construct the positive sample O ins+ , for example Figure 1 “[sheep]” and “[meadow]” in . Connect the features of the unimportant areas (the corresponding values in F are equal to 1) (assuming there are k of them) as counterfactual samples (negative samples) Figure 1 . Here, i = 1 means masking one of the least important area features, and so on, i = k represents masking k. For example
[0041] “[dog]” and “[man]” in ins+ and T. Help the model perceive the visual objects in the image and learn the fine-grained local feature alignment between the picture and the text. Design the InfoNCE contrastive loss function, which is specifically expressed as follows:
[0042]
[0043] Here, exp(n)=e n , T is the original text feature, and τ = 0.15 is the temperature parameter. The similarity of the positive sample is closer to the similarity of the original sample compared to the similarity of the negative sample. Therefore, the similarity of the positive sample should be higher than the result of the negative sample.
[0044] Based on the above inferences, based on the contrastive loss corresponding to the positive and negative samples, improve the model's recognition ability for the main objects in the picture. By optimizing the above loss, the features (T, V ins+ ) of the positive sample pairs containing important objects are guided to be close in the feature space, while the features of the negative sample pairs are guided to be far away. Therefore, enable the model to focus on learning the region features highly relevant to the image and text, and alleviate the potential false correlations and selection biases in the dataset.
[0045] Step S3, asFigure 1 As shown in the image-level contrastive learning module, the important matrix generated in step S1.2 is used to construct image-level contrastive learning, enabling the model to focus on global scene information;
[0046] The step S3 further includes the following steps:
[0047] Step S3.1: Construct image-level contrastive learning: Use the one obtained in S1.2 to judge the importance degree of the image local feature V i i∈D n for the current text feature, with the same principle as in S2.1. Randomly mask 20% of the unimportant local features (the corresponding value in is 1) to construct the positive sample B img+ , and randomly select m images from the remaining part as negative samples
[0048] Step S3.2: Input the positive and negative samples obtained in step S3.1 and the corresponding text features into the basic cross-modal retrieval model created in S1.1, and use the feature vector of the [CLS] flag bit in the obtained features as the final feature vector V img+ and as well as T. Guide the cross-modal retrieval model to learn fine-grained picture-text global feature alignment. Design the InfoNCE contrastive loss function, which is specifically expressed as follows:
[0049]
[0050] Here, exp(n)=e n , T is the original text feature, τ = 0.15 is the temperature parameter. The similarity of the positive sample is closer to the similarity of the original sample compared to the similarity of the negative sample. Therefore, the similarity of the positive sample should be higher than the result of the negative sample. Based on the above inferences, based on the contrastive loss corresponding to the positive and negative samples, improve the model's recognition ability of the global scene of the picture. By optimizing the above loss, the features (T, V img+ ) of the positive sample pairs with similar global scenes are guided to be close in the feature space, while the features of the negative sample pairs with different scenes are guided to be far away. Therefore, the model can learn fine-grained picture-text feature alignment from a global perspective.
[0051] Step S4, as Figure 1 shown in the semantic-level contrastive learning module, considers the counterfactual samples of the correct option text at the semantic level, constructs semantic-level contrastive learning, and enables the model to model cross-modal semantic relationships;
[0052] The step S4 further includes the following steps:
[0053] Step S4.1: Construct semantic-level contrastive learning: Randomly replace the nouns in the original text using the counterfactual idea to generate k (assuming there are k nouns in the original text) negative samples The original text serves as the positive sample G + .
[0054] Step S4.2: Input the positive and negative text samples and image features obtained in S4.1 into the basic cross-modal retrieval model created in S1.1, and use the feature vector of the [CLS] flag bit in the obtained features as the final feature vector T sem+ and and T to train the model to pay more attention to the understanding of visual words, thereby establishing the relationship between different modalities and promoting the understanding of language expressions. The specific contrast loss function is as follows:
[0055]
[0056] where sexp(n) = e n , V is the image feature, and τ = 0.1 is the temperature parameter. The similarity of the positive sample is closer to the similarity of the original sample compared to the similarity of the negative sample. Therefore, the similarity of the positive sample should be higher than the result of the negative sample. Based on the above inferences, based on the contrast loss corresponding to the positive and negative samples, improve the model's recognition ability of visual words in the text. By optimizing the above loss, the image feature V and the original text feature T with similar semantics sem+ are guided to be close, while the features with dissimilar semantics are guided to be far from each other. Therefore, the semantic-level contrast loss promotes the model's understanding of language expressions and captures complex cross-modal relationships.
[0057] Step S5, by integrating all the above losses into a unified framework, design a loss function for the overall training of the visual cross-modal retrieval model:
[0058]
[0059] where λ 1 = 0.2, λ 2 = 0.2, λ 3 = 0.2, λ 4 = 0.4, which are balancing parameters.
[0060] The cross-modal retrieval method of the cross-modal retrieval model based on counterfactual reasoning is as follows:
[0061] After the model is fully trained on the training set, for any test image, it is input into the model to obtain the feature vector of the [CLS] flag bit of the final output. The similarity between this image and all texts in the test library is calculated through formula (6), and the text with the largest similarity is retrieved as the retrieval result; given a test text, it is input into the model to obtain the feature vector of the [CLS] flag bit of the final output. The similarity between this text and all images in the test library is calculated through formula (6), and the image with the largest similarity is retrieved as the retrieval result.
[0062] A computer device according to the present invention, wherein the computer device internally stores the execution program code of the cross-modal retrieval model based on counterfactual reasoning or the stored program code, or the execution program code of the cross-modal retrieval method described above or the stored program code.
[0063] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation manners of the present invention, and they are not intended to limit the protection scope of the present invention. Any equivalent manners or changes that do not depart from the technology created by the present invention should be included in the protection scope of the present invention.
Claims
1. A cross-modal retrieval system based on counterfactual reasoning, characterized in that, this system is obtained through the following steps: S1. Respectively extract the original image features and text features. After independently mapping the obtained image features and text features to the same dimension, use the image feature encoder and text feature encoder composed of four layers of transformers to obtain feature vectors respectively. Use a two-layer Transformer for high-level semantic alignment of the obtained image feature vectors and text feature vectors, so that the image feature vectors and text feature vectors are mapped to the same common space, and optimize by calculating the loss; Then process the image features and the original text using the counterfactual reasoning method, compare the recognized image region labels with the nouns extracted from the original text, and provide a coefficient matrix for constructing positive and negative samples for the subsequent instance-level, image-level, and semantic-level contrast learning; S2. Use counterfactual reasoning to construct positive and negative samples at the instance level, enabling the system to focus on the detailed information of the objects in the visual image; S3. Use counterfactual generation to generate positive and negative samples at the image level, enabling the system to focus on the global scene information of the image; S4. Use counterfactual generation to generate counterfactual samples of the text at the semantic level, construct semantic-level contrast learning, and enable the system to cross-modal semantic relationships; S5. Integrate the above processes to obtain a cross-modal retrieval system based on counterfactual reasoning.
2. A cross-modal retrieval system based on counterfactual reasoning according to claim 1, characterized in that, the specific implementation of S1 includes: S1.1: For each image I and its corresponding text description E in the training data, extract its text features Image features Location features and the object region labels of the image where D q , D v represents the dimensions of the text features and the image features, D s represents the number of labeled regions in the image, i.e., the contour regions of all the recognized objects in this image, D n represents the number of local regions extracted from the image, D p denotes the location feature dimension of the local regions, D l represents the length of the sentence; The image features are extracted by Faster-RCNN to obtain 36 regional visual features, which are then concatenated together as the image feature V. The dimension of each regional visual feature is 2048. The position feature P includes the upper-left coordinate and the lower-right coordinate of each feature region as well as the area of the region, and D p is 5; where x 1 , y 1 , x 2 , y 2 are the upper-left coordinates and the lower-right coordinates of the region respectively, and W and H represent the width and height of the picture; The text feature T is extracted by BERT to obtain a 768-dimensional text feature, T F = FC t (Bert(T))#(3) FC v and FC t represent two independent fully connected layers, Bert represents the Bert model, represents concatenating the two values before and after in the same dimension, T F ∈R De , D e = 1024; Construct a basic cross-modal retrieval system including an image feature encoder and a text feature encoder each composed of four layers of transformers and a two-layer parameter-sharing Transformer structure for high-level semantic feature alignment; The V obtained after processing the image features and text features through formulas (2) and (3) F and T F are input into the created basic cross-modal retrieval system. The feature vector corresponding to the [CLS] flag bit in the obtained features is used as the final feature vector, and a triplet loss function is established as the loss function for image and text alignment: where \(t, x\) represent positive sample pairs, \(t\) - , \(x\) - represents a negative sample, that is, other text features and image features of the same batch, \(\alpha\) is a hyperparameter, \([a]\) + =\(\max(a, 0)\), S1.2: Extract the image region label S, align it with the nouns extracted from the original text, and then compare it with the outline of the object to be masked. Divide the length and width of the rectangle formed by all local feature regions of the image into 14 equal parts each, take a total of 196 intersection points, and then calculate the importance coefficient of the image region feature to the text by dividing the number of points falling inside the object outline by 196. Then connect all the values to form a coefficient matrix As shown in Equation (5), where F i The smaller the value, the more important the regional feature V i is; P i Represents the position feature of region i Represents the object region label of the picture, E represents the original text, mask(P i , I i , E) represents the number of points where region i falls inside the important region, F i Represents The value of the i-th row of 3. A cross-modal retrieval system based on counterfactual reasoning according to claim 2, characterized in that, the specific implementation of S2 includes: S2.1: Construct instance-level contrastive learning: Use the matrix obtained in S1.2 to determine the importance of the local features V of the image i i ∈ D n for the current text features. Connect the features of the important regions where the corresponding values in are not 1 to construct the positive sample O ins+ and connect the features of the unimportant regions where the corresponding values in are 1 as the counterfactual sample, i.e., the negative sample where k represents the number of instance-level negative samples; S2.2: Input the positive and negative samples obtained in step S2.1 and their corresponding text features into the basic cross-modal retrieval system created in S1.1, and then use the feature vector corresponding to the [CLS] flag bit in the obtained features as the final feature vector V ins+ and and T, enabling the system to perceive visual objects in the image and learn fine-grained local feature alignment between pictures and texts, and designing an InfoNCE contrastive loss function, which is specifically expressed as follows: Here, exp(n) = e n , T is the original text feature, τ = 0.15, which is the temperature parameter.
4. A cross-modal retrieval system based on counterfactual reasoning according to claim 2, characterized in that, the specific implementation of S3 includes: Step S3.1: Construct image-level contrastive learning: Use the to judge the importance of the local image feature V i i∈D n to the current text feature, randomly mask 20% of the unimportant local features to construct the positive sample B img+ , and randomly select m images from the rest as negative samples m represents the number of image-level negative samples; Step S3.2: Input the positive and negative samples obtained in Step S3.1 and their corresponding text features into the cross-modal retrieval system created in S1.1, and then use the feature vector corresponding to the [CLS] flag bit in the obtained features as the final feature vector V img+ and and T, so that the cross-modal retrieval system learns fine-grained global feature alignment of image-text, and designs the InfoNCE contrastive loss function, which is specifically expressed as follows: Here, exp(n) = e n , T is the original text feature, and τ = 0.15 is the temperature parameter.
5. A cross-modal retrieval system based on counterfactual reasoning according to claim 1, characterized in that, the specific implementation of S4 includes: S4.1: Construct semantic-level contrastive learning: Assume that the original text has k nouns. Use the counterfactual idea to randomly replace the nouns in the original text to generate k negative samples The original text serves as the positive sample G + ; S4.2: Input the positive and negative text samples and image features obtained in S4.1 into the basic cross-modal retrieval system created in S1.1, and use the feature vector of the [CLS] flag bit in the obtained features as the final feature vector T sem+ and and T to train the system to pay more attention to the understanding of visual words, thereby establishing the relationship between different modalities and promoting the understanding of language expressions. The specific contrast loss function is as follows: Here, exp(n) = e n , V is the original image feature, and τ = 0.1 is the temperature parameter.
6. A cross-modal retrieval system based on counterfactual reasoning according to claim 1, characterized in that, the fusion method of S5 is as follows: Design loss function Conduct the overall training of the cross-modal retrieval system: Here, λ 1 = 0.2, λ 2 = 0.2, λ 3 = 0.2, λ 4 = 0.4 are the equalization parameters.
7. A cross-modal retrieval method of a cross-modal retrieval system based on counterfactual reasoning according to any one of claims 1-6, characterized in that, For any test image, input it into the system according to any one of claims 1-6, calculate the overall similarity between the image and all texts in the system test library, and retrieve the text with the largest similarity as the retrieval result; for any test text segment, calculate the similarity between the text and all images in the test library, and retrieve the image with the largest similarity as the retrieval result.
8. A computer device, characterized in that, The computer device internally stores the executable program code or the stored program code of the cross-modal retrieval system based on counterfactual reasoning according to any one of claims 1-6, or the executable program code or the stored program code of the cross-modal retrieval method according to claim 7.
Citation Information
Patent Citations
Construction method and application of cross-modal retrieval model based on multilayer attention mechanism
CN113779361A
News event searching method and system based on multistage image-text semantic alignment model
CN114297473A