Instance-level Cross-modal Retrieval Method Based on Relational Reasoning and Cross-modal Independent Matching Network
Through the instance-level cross-modal retrieval method of relational reasoning and cross-modal independent matching network, the problem of fine-grained instance-level cross-modal retrieval between image text is solved, and efficient semantic aggregation and precise matching of multimodal data is achieved.
Patent Information
- Application Number
- CN202310867080.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-07-14
AI Technical Summary
The prior art is difficult to effectively solve the problem of fine-grained instance-level cross-modal retrieval between image texts, especially when there are large differences in multimodal heterogeneity.
An instance-level cross-modal retrieval method for independent matching networks of relational inference and cross-modal matching networks is proposed. Local relationship inference and global semantic aggregation of multimodal data is realized through technical means such as modal feature extraction, feature relationship inference, graph pooling and gravitational loss function.
It realizes efficient and accurate completion of fine-grained instance-level cross-modal retrieval tasks in multimodal scenarios, reduces computational costs, and improves the accuracy of cross-modal matching.
Smart Images

Figure CN116881416B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of pattern recognition, computer vision, and natural language processing, and in particular to an instance-level cross-modal retrieval method for relational reasoning and cross-modal independent matching network. Background Art
[0002] With the rapid development of multimedia technology, more and more modal forms have become carriers for carrying information. Among them, vision and language, as two of the most important ways for humans to obtain and express information, have huge differences in data distribution and feature representation. The human brain can integrate various forms of multi-modal information through a complex nervous system and react to external changes in combination with existing experience, while machines do not have such a complex system. How to overcome the differences in multi-modal heterogeneity has become an urgent problem to be solved. In recent years, there have been many related studies, such as cross-modal retrieval, image captioning, visual question answering, and visual grounding. In recent years, great progress has been made in cross-modal retrieval between images and texts. However, due to the huge differences between the two and the similarity of internal individuals, fine-grained instance-level cross-modal retrieval between images and texts remains a challenging problem. Summary of the Invention
[0003] The present invention proposes an instance-level cross-modal retrieval method for relational reasoning and cross-modal independent matching network, which can effectively perform local relational reasoning and global semantic aggregation on multi-modal data, and efficiently and accurately complete the fine-grained instance-level cross-modal retrieval task in a multi-modal scenario.
[0004] The present invention adopts the following technical solutions.
[0005] The instance-level cross-modal retrieval method for relational reasoning and cross-modal independent matching network includes the following steps;
[0006] Step S1: For the image modality, use Faster R-CNN pre-trained on the Visual Genome dataset with ResNet-101 as the backbone to extract visual features of multiple significant regions of each image, and for the text modality, use a word embedding method to extract modality features;
[0007] Step S2: Use a feature relationship reasoning module composed of neighborhood relationship reasoning and potential relationship reasoning to perform image modality feature relationship reasoning on the image modality features obtained in Step S1; use pre-trained BERT to perform text modality feature relationship reasoning on the text modality features obtained in Step S1;
[0008] Step S3: Use the graph pooling method to perform modal global semantic aggregation on the image modal features and text modal features with fused context obtained in Step S2 respectively to obtain modal representations;
[0009] Step S4: Calculate the multi-modal feature similarity for the modal representations obtained in Step S3, and return the cross-modal retrieval results according to the similarity; Based on the retrieval results, use the gravitational loss function to guide and correct the learning process of the matching relationship within and between modalities during the neural network training process.
[0010] Step S11: For the image modality, the model uses Faster R-CNN pre-trained on the Visual Genome dataset to extract the visual features of 36 significant regions in each image. Specifically, the model first converts the input image into a feature map through the ResNet-101 backbone network, and then uses the Region Proposal Network (RPN) on the feature map to generate a number of candidate regions; Then use Region of Interest Pooling to map each candidate region to a feature vector of a fixed size, and this feature vector is the visual feature of the region;
[0011] Step S12: For the text modality, first use the WordPiece tokenization algorithm based on the greedy algorithm to decompose words into the smallest sub-word units. Then, input the tokenized text into the Transformer encoder, and the encoder encodes the sub-word sequence of each word to obtain the word embedding vector of the word. Finally, add the position embedding and segment embedding to the word embedding vector.
[0012] The specific implementation of Step S2 is as follows:
[0013] Step S21: For the regional features of the image modality, use the feature relationship reasoning module composed of neighborhood relationship reasoning and latent relationship reasoning to perform image modality feature relationship reasoning;
[0014] The specific method of neighborhood relationship reasoning is: The regional features of the image modality dispatcher are an unordered set of feature representations containing m = 36 significant regions. First, establish a two-dimensional grid matrix, and allocate these 36 significant region features to the two-dimensional grid matrix according to their block positions in the original image; Denote the relative positions of the centers of all features in the original image as the point set X = {(x 1 , y 1 ),..., (x m , y m )}, and denote the set of all points in the two-dimensional grid matrix as the point set Y = {(x′ 1 , y′ 1 ),..., (x′m , y' m )}; where x and y are the horizontal and vertical coordinates respectively; each point in the point set X must and can only match one point in the point set Y, and the distance D between two points is the matching cost, denoted as the distance between their relative positions, and expressed by the formula:
[0015]
[0016] where, A x , B x , A y , B y are the relative horizontal coordinate and the relative vertical coordinate respectively, which is the ratio of the image horizontal or vertical coordinate to the image width or height, and the value range is [0, 1];
[0017] To solve the element matching result between the point set X and the point set Y under the minimization of the cost C, the sum of the Euclidean distances of the relative position coordinates of all point pairs in the set X and the set Y is used as the matching cost, and the above problem is abstracted into the minimum cost matching problem of a bipartite graph, and the Hungarian algorithm is used to solve the matching result under the minimum matching cost, and expressed by the formula
[0018]
[0019] Thus, the two-dimensional structured modeling of the unordered regional features is completed; in the modeled two-dimensional grid matrix, the features of its neighborhood can be easily obtained from each feature;
[0020] The two-dimensional grid matrix is divided into small grids with side length a in a non-overlapping manner, and global self-attention calculation is performed in each small grid, and the formula is:
[0021] v' = MSA(LN(v)) + v Formula (3-3)
[0022] v'' = MLP(LN(v')) + v' Formula (3-4)
[0023] where, v represents the feature input of a small grid, v' is an intermediate variable, and v'' is the feature output that fuses local context. MSA(·) represents multi-head self-attention, and MLP(·) represents multi-layer perceptron;
[0024] The specific method of potential relationship reasoning is: for each regional feature v i (i = 1, 2,..., m), connect it with k regional features that do not belong to the same small grid as it to construct a path, and the path weight is the cosine similarity between the two regional features;
[0025] After constructing all the paths, find a connected circuit with the minimum weight sum, that is, find a feature string with the maximum distance; among this feature string, the sum of the differences between adjacent features is the largest, and the sum of similarities is the smallest; adopt a greedy strategy to approximately construct the circuit with the minimum weight sum, that is, start from a random node in the two-dimensional grid matrix, and take the farthest node among the k random nodes connected to it at each step;
[0026] After the above steps, a one-dimensional feature ring is obtained; the feature ring is segmented according to the length l, and the attention calculation is carried out within each segment by using formula (3-3) and formula (3-4);
[0027] The inference of the feature relationship of the image modality is based on the feature relationship inference module. For the image modality, the feature relationship inference module FRR is composed of the above neighborhood relationship inference NRR and potential relationship inference PRR in parallel; expressed by the formula as:
[0028] FRR(v l ) = NRR(v l ) + PRR(v l ), l = 1...N Formula (3-5)
[0029] where v l is the visual feature of the l-th layer, and N is the number of layers of the module. To improve the network expression ability and training stability, for each feature relationship inference module, an additional residual branch is introduced:
[0030] v l+1 = FRR(v l ) + v l+1 , l = 1...N Formula (3-6)
[0031] Infer the image region features through the feature relationship inference module with N = 2 layers, so as to obtain the region features that fuse the context;
[0032] Step S22, use the pre-trained BERT to perform text modality feature relationship inference on the word embedding vectors of the text modality; among them, the BERT model is composed of 12 layers of standard Transformer encoders, and the dimension of the word embedding is 768; retain the output corresponding to each element in the entire input sequence.
[0033] The specific implementation of step S3 is as follows:
[0034] Step S31, for the local features V = {v 1 , v 2 ,... v m} and T = {t 1 , t 2 ,... tn} Sort in descending order, rearrange all feature elements in each mini-batch according to the decreasing order of their values, and obtain the rearranged features and Among the sorted features, the i-th feature is the result of the i-th largest pooling. For example, the first feature is the result of the maximum pooling, the second feature is the result of the second largest pooling, and the last feature is the result of the minimum pooling;
[0035] Step S32: Take each feature as a node of the graph convolutional network GCN, construct a fully-connected graph of modal features, use graph convolution to calculate the weights of all rearranged local features, and aggregate the information of all non-maximum feature elements to the maximum feature element;
[0036] In GCN, each weight transfer calculation takes a certain node feature as input, and query Q, key K, and value V are calculated respectively through 3 independent one-dimensional convolutions; Q and K are dot-multiplied and normalized to obtain the weight of each element of V and multiply it with V; finally, a Conv1d and normalization are performed to obtain the transferred feature; the updated feature of each node is obtained by summing the transferred features of all current nodes; in this process, in order to avoid unnecessary calculations, the calculation of aggregating to non-maximum feature elements is not performed.
[0037] The specific implementation of step S4 is as follows:
[0038] Step S41: Design the sample quality based on sample saliency.
[0039] For an image sample, its sample quality m V is defined as the area of the intersection coverage area of its m significant regions and the image area S V The ratio; expressed by the formula:
[0040]
[0041] For a text sample, its sample quality m T is defined as the ratio of the number of its important words to the total number of words n contained in the sample; that is:
[0042]
[0043] where w i represents the weight of the i-th word word i in the text, and the formula is:
[0044]
[0045] Part-of-speech tagging is performed on all text data using Stanford CoreNLP to obtain the part of speech of each word. Among them, nouns (n.), verbs (v.), adjectives (adj.), locative words (p.), and cardinal numbers (cd.) are considered important words because they contain important information about the types, attributes, quantities, locations, and interaction relationships of entities in the samples. The remaining words such as and, or, where, and punctuation marks are considered unimportant words;
[0046] Step S42: Design a loss function based on the sample quality;
[0047] When using image V as the query, sample all the texts in each mini-batch to form positive sample pairs and negative sample pairs; the weighted similarity of the positive sample pairs should be higher than that of the negative sample pairs by a threshold γ; similarly, when using text T as the query, the weighted similarity of the positive sample pairs and negative sample pairs should also satisfy the above rules; the difference in weighted similarity has a non-linear square relationship with the loss; when the difference is large, the loss of pushing away is small, and when the difference is small, the loss increases sharply; combined with the online hard example mining technique, strive to optimize the most difficult negative samples in the mini-batch; the design of the gravitational loss function is as follows:
[0048]
[0049]
[0050]
[0051] Among them, L VT is the inter-modal gravitational loss, L V is the intra-modal gravitational loss of the image modality, L T is the intra-modal gravitational loss of the text modality. T′ and V′ are the most difficult negative samples in the mini-batch; κ is the similarity kernel, the global similarity of the image-text pair calculated using the Euclidean distance; the function [·] + is equivalent to max(·, 0). m V 、m T 、m′ V 、m′ T are all the qualities of the samples; q is a coefficient used to adjust the importance of quality in the gravitational loss, set to 8. Each image-text sample has its unique quality;
[0052] The total loss is expressed as the weighted sum of the inter-modal gravitational loss L inter and the intra-modal gravitational loss L intra ; where β is an adjustable parameter,
[0053] L = L inter + βL intra Formula (5-7)
[0054] L = L inter + βL intra
[0055] = L VT + βL V + βL T 。
[0056] When simplifying calculations, take β = 0.5.
[0057] The present invention relates to an instance-level cross-modal retrieval method based on relational reasoning and cross-modal independent matching network. First, a modal feature extractor is used to convert the input original image into regional features and the input text into a word sequence. Then, modal feature relationship reasoning is respectively performed on the image and text modalities to explore the interaction relationships between local features. Next, a graph pooling method based on a graph network is used to perform modal global semantic aggregation on the rearranged features. Finally, the similarity between multi-modal features is calculated, and the cross-modal retrieval results are returned according to the similarity. During the neural network training process, a gravitational loss function is used to guide and correct the learning process of the matching relationships within and between modalities; compared with the prior art, the present invention has the following beneficial effects:
[0058] 1. Aiming at the problem of low efficiency of the existing global relational reasoning method, the present invention proposes neighborhood relationship and potential relationship reasoning to achieve feature relationship reasoning, and can replace the existing method with a lower computational cost to achieve global relational reasoning of modal features.
[0059] 2. Aiming at the problems of resource waste, semantic ambiguity, sequence interference, etc. existing in the existing pooling schemes, the present invention proposes a graph pooling method to achieve efficient and comprehensive modal global semantic aggregation. Compared with the mainstream serialized pooling schemes, it not only has higher efficiency but also has good compatibility.
[0060] 3. The present invention incorporates sample quality into the graphic-text matching learning process and designs a gravitational loss function to learn more diverse matching relationships within and between modalities, and can achieve more accurate cross-modal matching.
[0061] 4. The design of the independent matching framework of the present invention can independently generate embeddings of different modalities. When reasoning modal features, only single calculations (linear time complexity) need to be performed on the image and text data respectively, instead of calculating all images and texts pairwise like the interactive matching method (quadratic time complexity), which greatly reduces the calculation time of the neural network and has high computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The present invention will be further described in detail below with reference to the drawings and specific embodiments:
[0063] Att Figure 1It is a schematic diagram of the principle of the present invention. DETAILED DESCRIPTION
[0064] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0065] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0066] The present invention relates to an instance-level cross-modal retrieval method based on relational reasoning and cross-modal independent matching network. First, a modal feature extractor is used to convert the input original image into regional features and the input text into a word sequence. Then, modal feature relationship reasoning is performed on the image and text modalities respectively to explore the interaction relationship between local features. Then, a graph pooling method based on a graph network is used to perform modal global semantic aggregation on the rearranged features. Finally, the similarity between multimodal features is calculated, and the cross-modal retrieval results are returned according to the similarity. During the neural network training process, the gravitational loss function is used to guide and correct the learning process of the intra-modal and inter-modal matching relationship.
[0067] As shown in the figure, the instance-level cross-modal retrieval method of relational reasoning and cross-modal independent matching network includes the following steps:
[0068] Step S1: For the image modality, the Faster R-CNN pre-trained on the Visual Genome dataset VisualGenomes with ResNet-101 as the backbone is used to extract the visual features of multiple salient areas of each image. For the text modality, the word embedding method is used to extract modality features;
[0069] Step S2: using a feature relationship reasoning module composed of neighborhood relationship reasoning and potential relationship reasoning to perform image modal feature relationship reasoning on the image modal feature obtained in step S1; using a pre-trained BERT to perform text modal feature relationship reasoning on the text modal feature obtained in step S1;
[0070] Step S3: using a graph pooling method to perform modality global semantic aggregation on the image modality features and text modality features of the fusion context obtained in step S2, so as to obtain a modality representation;
[0071] Step S4: Calculate the multi-modal feature similarity for the modal representation obtained in Step S3, and return the cross-modal retrieval result according to the similarity; based on the retrieval result, use the gravitational loss function to guide and correct the learning process of the matching relationship within and between modalities during the neural network training process.
[0072] Step S11: For the image modality, the model uses Faster R-CNN pre-trained on the Visual Genome dataset to extract the visual features of 36 significant regions in each image. Specifically, the model first converts the input image into a feature map through the ResNet-101 backbone network, and then uses the Region Proposal Network (RPN) on the feature map to generate a number of candidate regions; then uses Region of Interest Pooling to map each candidate region to a feature vector of a fixed size, and this feature vector is the visual feature of the region.
[0073] Step S12: For the text modality, first use the WordPiece tokenization algorithm based on the greedy algorithm to decompose words into the smallest sub-word units. Then, input the tokenized text into the Transformer encoder, and the encoder encodes the sub-word sequence of each word to obtain the word embedding vector of the word. Finally, add the position embedding and segment embedding to the word embedding vector.
[0074] The specific implementation of Step S2 is as follows:
[0075] Step S21: For the regional features of the image modality, use a feature relationship reasoning module composed of neighborhood relationship reasoning and potential relationship reasoning to perform image modality feature relationship reasoning.
[0076] The specific method of neighborhood relationship reasoning is as follows: The regional features of the image modality dispatcher are an unordered set of feature representations containing m = 36 significant regions. First, establish a two-dimensional grid matrix, and assign these 36 significant region features to the two-dimensional grid matrix according to their block positions in the original image; record the relative positions of the centers of all features in the original image as the point set X = {(x 1 , y 1 ),..., (x m , y m )}, and record the set of all points in the two-dimensional grid matrix as the point set Y = {(x' 1 , y' 1 ),..., (x' m , y' m )}; where x and y are the horizontal and vertical coordinates respectively; each point in the point set X must and can only be matched with one point in the point set Y, and the distance D between the two points, that is, the matching cost, is recorded as the distance between their relative positions, and is expressed by the formula:
[0077]
[0078] Among them, A x , B x , A y , B y are the relative abscissa and relative ordinate respectively, which are the ratios of the image abscissa or ordinate to the image width or height, and the value range is [0, 1];
[0079] In order to solve the element matching result between point set X and point set Y under the minimization of cost C, the sum of the Euclidean distances of the relative position coordinates of all point pairs in set X and set Y is used as the matching cost, and the above problem is abstracted into the minimum cost matching problem of a bipartite graph, and the Hungarian algorithm is used to solve the matching result under the minimum matching cost, which is expressed by the formula as
[0080]
[0081] Thus, the two-dimensional structured modeling of unordered regional features is completed; in the modeled two-dimensional grid matrix, the features of its neighborhood can be easily obtained from each feature;
[0082] The two-dimensional grid matrix is divided into small grids with side length a in a non-overlapping manner, and global self-attention calculation is performed in each small grid. The formula is:
[0083] v′ = MSA(LN(v)) + v Formula (3-3)
[0084] v″ = MLP(LN(v′)) + v′ Formula (3-4)
[0085] Among them, v represents the feature input of a small grid, v′ is an intermediate variable, and v″ is the feature output that fuses local context. MSA(·) represents multi-head self-attention, and MLP(·) represents multi-layer perceptron;
[0086] The specific method of potential relationship reasoning is: for each regional feature v i (i = 1, 2,..., m), connect it with k regional features that do not belong to the same small grid as it to construct a path, and the path weight is the cosine similarity between the two regional features;
[0087] After all paths are constructed, find a connected circuit with the minimum weight sum, that is, find a feature string with the farthest distance; among this feature string, the sum of the differences between adjacent features is the largest, and the sum of similarities is the smallest; adopt a greedy strategy to approximately construct the circuit with the minimum weight sum, that is, start from a random node in the two-dimensional grid matrix, and take the farthest node among the k random nodes connected to it at each step;
[0088] A one-dimensional feature ring is obtained through the above steps; the feature ring is segmented according to the length l, and attention calculation is performed within each segment using formulas (3-3) and (3-4).
[0089] The inference of the feature relationship of the image modality is based on the feature relationship inference module. For the image modality, the feature relationship inference module FRR is composed of the above neighborhood relationship inference NRR and potential relationship inference PRR in a parallel manner; it is expressed by the formula as follows:
[0090] FRR(v l ) = NRR(v l ) + PRR(v l ), l = 1...N Formula (3-5)
[0091] where v l is the visual feature of the l-th layer, and N is the number of layers of the module. To improve the network expression ability and training stability, for each feature relationship inference module, an additional residual branch is introduced:
[0092] v l+1 = FRR(v l ) + v i+1 , l = 1...N Formula (3-6)
[0093] The image region features are inferred through the feature relationship inference module with N = 2 layers to obtain the region features that fuse the context.
[0094] Step S22, the pre-trained BERT is used for the inference of the feature relationship of the text modality for the word embedding vectors of the text modality; the BERT model is composed of 12 standard Transformer encoders, and the dimension of the word embedding is 768; the output corresponding to each element in the entire input sequence is retained.
[0095] The specific implementation of step S3 is as follows:
[0096] Step S31, the local features V = {v 1 , v 2 ,...v m} and T = {t 1 , t 2 ,...t n} of the context of the image and text modalities are sorted in descending order, and all feature elements in each mini-batch are rearranged according to the decreasing order of their values to obtain the rearranged features And Among the sorted features, the i-th feature is the result of the i-th largest pooling. For example, the first feature is the result of the maximum pooling, the second feature is the result of the second largest pooling, and the last feature is the result of the minimum pooling.
[0097] Step S32: Take each feature as a node of the graph convolutional network GCN to construct a fully-connected graph of modal features, calculate the weights of all rearranged local features using graph convolution, and aggregate the information of all non-maximum feature elements onto the maximum feature element;
[0098] In GCN, each calculation of weight transfer takes a certain node feature as input, and the query Q, key K, and value V are calculated respectively through 3 independent one-dimensional convolutions; the weights of each element of V are obtained by multiplying Q and K through dot product and normalization and then multiplying with V; finally, a Conv1d and normalization are performed again to obtain the transferred features; the updated feature of each node is obtained by summing the transferred features of all current nodes; in this process, in order to avoid unnecessary calculations, the calculation of aggregating to non-maximum feature elements is not performed.
[0099] The specific implementation of step S4 is as follows:
[0100] Step S41: Design the sample quality based on sample saliency.
[0101] For an image sample, its sample quality m V is defined as the area of the intersection coverage region of its m significant regions and the area S V of the image; expressed by the formula:
[0102]
[0103] For a text sample, its sample quality m T is defined as the ratio of the number of its important words to the total number of words n contained in the sample; that is:
[0104]
[0105] where w i represents the weight of the i-th word word i in the text, and the formula is:
[0106]
[0107] Use StanfordCoreNLP to perform part-of-speech tokenization on all text data to obtain the part-of-speech of each word. Among them, nouns n., verbs v., adjectives adj., locative words p., and cardinal numbers cd. are considered important words because they contain important information about the types, attributes, quantities, locations, and interaction relationships of entities in the sample. The rest, such as and, or, where, and punctuation marks, are considered non-important words;
[0108] Step S42: Design the loss function based on the sample quality;
[0109] When using image V as a query, all texts in each mini-batch are sampled to form positive sample pairs and negative sample pairs; the weighted similarity of positive sample pairs should be higher than that of negative sample pairs by a threshold γ; similarly, when using text T as a query, the weighted similarity of positive sample pairs and negative sample pairs should also satisfy the above rules; the difference in weighted similarity and the loss show a non-linear square relationship; when the difference is large, the pushing-away loss is small, and when the difference is small, the loss increases sharply; combined with online hard example mining technology, efforts are made to optimize the most difficult negative samples in the mini-batch; the design of the gravitational loss function is as follows:
[0110]
[0111]
[0112]
[0113] Among them, L VT is the cross-modal gravitational loss, L V is the intra-modal gravitational loss of the image modality, and L T is the intra-modal gravitational loss of the text modality. T′ and V′ are respectively the most difficult negative samples in the mini-batch; κ is the similarity kernel, the global similarity of the image-text pair calculated using the Euclidean distance; the function [·] + is equivalent to max(·, 0). m V , m T , m′ V , m′ T are all the qualities of the samples; q is a coefficient used to adjust the importance of quality in the gravitational loss, set to 8. Each image-text sample has its unique quality;
[0114] The total loss is expressed as the weighted sum of the cross-modal gravitational loss L inter and the intra-modal gravitational loss L intra ; among them, β is an adjustable parameter,
[0115] L = L inter + βL intra Equation (5-7)
[0116] L = L inter + βL intra
[0117] = L VT + βL v + βL T .
[0118] When simplified calculations are required, take β = 0.5.
[0119] In this example, in view of the problem of low efficiency of the existing global relationship reasoning method, the present invention proposes neighborhood relationship and potential relationship reasoning to achieve feature relationship reasoning, which can replace the existing method with a lower computational cost to achieve global relationship reasoning of modal features. In view of the problems of resource waste, semantic ambiguity, sequence interference, etc. existing in the existing pooling scheme, the present invention proposes a graph pooling method to achieve efficient and comprehensive modal global semantic aggregation. Compared with the mainstream serialization pooling scheme, it not only has higher efficiency but also has good compatibility. The present invention incorporates sample quality into the process of image-text matching learning and designs a gravitational loss function to learn more diverse intra-modal and inter-modal matching relationships, which can achieve more accurate cross-modal matching. The design of the independent matching framework of the present invention can independently generate embeddings of different modalities. When reasoning about modal features, only a single calculation (linear time complexity) needs to be performed on the image and text data respectively, instead of calculating by pairing all images and texts one by one like the interactive matching method (quadratic time complexity), which greatly reduces the calculation time of the neural network and has high computational efficiency.
[0120] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention, when the functions and effects produced do not exceed the scope of the technical solution of the present invention, shall fall within the protection scope of the present invention.
Claims
1. Instance-level Cross-modal Retrieval Method Based on Relational Reasoning and Cross-modal Independent Matching Network, Characterized in that: It includes the following steps; Step S1: For the image modality, use Faster R-CNN with ResNet-101 as the backbone and pre-trained on the Visual Genomes dataset to extract visual features of multiple significant regions of each image. For the text modality, use the word embedding method for modality feature extraction; Step S2: Use the feature relationship reasoning module composed of neighborhood relationship reasoning and potential relationship reasoning to perform image modality feature relationship reasoning on the image modality features obtained in Step S1; Use pre-trained BERT to perform text modality feature relationship reasoning on the text modality features obtained in Step S1; Step S3: Use the graph pooling method to perform modality global semantic aggregation on the image modality features and text modality features with fused context obtained in Step S2 respectively to obtain modality representations; Step S4: Calculate the multi-modal feature similarity for the modality representations obtained in Step S3, and return the cross-modal retrieval results according to the similarity; Based on the retrieval results, use the gravitational loss function to guide and correct the learning process of the matching relationship within and between modalities during the neural network training process; The specific implementation of Step S2 is as follows: Step S21: Use the feature relationship reasoning module composed of neighborhood relationship reasoning and potential relationship reasoning to perform image modality feature relationship reasoning on the regional features of the image modality; The specific method for neighborhood relationship reasoning is as follows: The regional features of the image module dispatcher are an unordered set of feature representations containing m = 36 significant regions. First, a two-dimensional grid matrix is established, and these 36 significant region features are assigned to the two-dimensional grid matrix according to their block positions in the original image; the relative positions of the centers of all features in the original image are recorded as the point set X = {(x 1 , y 1 ),...,(x m , y m )}, and the set of all points in the two-dimensional grid matrix is recorded as the point set Y = {(x' 1 , y' 1 ),...,(x' m , y' m )}; where x and y are the horizontal and vertical coordinates respectively; each point in the point set X must and can only be matched with one point in the point set Y, and the distance D between the two points, that is, the matching cost, is recorded as the distance between their relative positions, expressed by the formula: Among them, A x , B x , A y , B y are the relative abscissa and relative ordinate respectively, which are the ratios of the image abscissa or ordinate to the image width or height, and the value range is [0, 1]; To solve the element matching result between point set X and point set Y under the minimization of cost C, take the sum of the Euclidean distances of the relative position coordinates of all point pairs in set X and set Y as the cost of matching, abstract the above solution process into the minimum cost matching problem of a bipartite graph, and use the Hungarian algorithm to solve the matching result under the minimum matching cost, which is expressed by the formula Thus, the two-dimensional structured modeling of the unordered regional features is completed; in the modeled two-dimensional grid matrix, the features of its neighborhood can be easily obtained from each feature; Divide the two-dimensional grid matrix into small grids with side length a in a non-overlapping manner, and perform global self-attention calculation in each small grid. The formula is: v′ = MSA(LN(v)) + v Formula (3-3) v" = MLP(LN(v')) + v′ Formula (3-4) Among them, v represents the feature input of a small grid, v′ is an intermediate variable, and v″ is the feature output integrating local context; MSA(·) represents multi-head self-attention, and MLP(·) represents a multi-layer perceptron; The specific method for potential relationship reasoning is as follows: for each regional feature v i (i = 1, 2,..., m), connect it to k regional features that do not belong to the same small grid with it to construct a path, and the path weight is the cosine similarity between the two regional features; After completing the construction of all paths, find a connected loop with the smallest weight sum, that is, find a feature string with the farthest distance; among this feature string, the sum of the differences of adjacent features is the largest, and the sum of similarities is the smallest; adopt a greedy strategy to construct the loop with the smallest weight sum, that is, start from a random node in the two-dimensional grid matrix, and take the farthest node among the k random nodes connected to it at each step; After the above steps, a one-dimensional feature ring is obtained; segment the feature ring according to the length l, and perform attention calculation within each segment using Formula (3-3) and Formula (3-4); The inference of the relationship between image modality features is based on the feature relationship inference module. For the image modality, the feature relationship inference module FRR is composed of the above-mentioned neighborhood relationship inference NRR and potential relationship inference PRR in parallel. It can be expressed by the formula: FRR(v l ) = NRR(v l ) + PRR(v l ), l = 1...N Equation (3-5) where v l is the visual feature of the l-th layer, and N is the number of layers of the module. To improve the network expression ability and training stability, for each feature relation reasoning module, a residual branch is additionally introduced: v l+1 = FRR(v l ) + v l+1 , l = 1...N Equation (3-6) The regional features of the image are inferred through the feature relationship inference module with N = 2 layers to obtain the regional features that fuse the context. Step S22: Use the pre-trained BERT to perform the inference of the text modality feature relationship for the word embedding vectors of the text modality. The BERT model is composed of 12 standard Transformer encoders, and the dimension of the word embedding is 768. The output corresponding to each element in the entire input sequence is retained. The specific implementation of step S3 is as follows: Step S31: Descendingly sort the local features V = {v 1 , v 2 ,..., v m} of the context of the image and text modalities and T = {t 1 , t 2 ,..., t n}, and rearrange all the feature elements in each mini-batch according to the decreasing order of their values to obtain the rearranged features and Among the sorted features, the i-th feature is the result of the i-th largest pooling. If the first feature is the result of the largest pooling, the second feature is the result of the second largest pooling, and the last feature is the result of the smallest pooling; Step S32: Take each feature as a node of the graph convolutional network GCN to construct a fully-connected graph of modality features, and use graph convolution to calculate the weights of all rearranged local features, and aggregate the information of all non-maximum feature elements to the maximum feature element. In GCN, each calculation of weight transfer takes a certain node feature as input, and the query Q, key K, and value V are calculated respectively through 3 independent one-dimensional convolutions. The weights of each element of V are obtained by multiplying Q and K through dot product and normalization and then multiplying with V. Finally, a Conv1d and normalization are performed again to obtain the transferred feature. The updated feature of each node is obtained by the sum of the transferred features of all current nodes. In this process, in order to avoid unnecessary calculations, the calculation of aggregating to non-maximum feature elements is not performed. The specific implementation of step S4 is as follows: Step S41: Design the sample quality based on sample saliency. For an image sample, its sample quality m V is defined as the area of the intersection coverage region of its m significant regions and the image area S V The ratio is expressed by the formula: For a text sample, define its sample quality m T as the ratio of the number of its important words to the total number of words n contained in the sample; that is: where w i represents the weight of the i-th word word i in the text, and the formula is: Use StanfordCoreNLP to perform part-of-speech tokenization on all text data to obtain the part-of-speech of each word. Among them, nouns (n.), verbs (v.), adjectives (adj.), locative words (p.), and cardinal numbers (cd.) are considered important words because they contain important information about the types, attributes, quantities, locations, and interaction relationships of entities in the sample. The rest, such as and, or, where, and punctuation marks, are considered unimportant words. Step S42: Design the loss function based on the sample quality. When using the image V as a query, sample all texts in each mini-batch to form positive sample pairs and negative sample pairs. The weighted similarity of the positive sample pairs should be higher than that of the negative sample pairs by a threshold γ. Similarly, when using the text T as a query, the weighted similarity of the positive sample pairs and negative sample pairs should also meet the above rules. The difference in weighted similarity and the loss show a non-linear square relationship. When the difference is large, the loss is small, and when the difference is small, the loss increases sharply. Combining the online hard example mining technique, it is committed to optimizing the most difficult negative samples in the mini-batch. The design of the gravitational loss function is as follows: Among them, L VT is the inter-modal gravitational loss, L V is the intra-modal gravitational loss of the image modality, L T is the intra-modal gravitational loss of the text modality; T′ and V′ are the most difficult negative samples in the mini-batch; κ is the similarity kernel, the global similarity of the image-text pair calculated using the Euclidean distance; the function [·] + i.e., max(·, 0); m V 、m T 、m′ V 、m′ T are all the qualities of the samples; q is the coefficient used to adjust the importance of quality in the gravitational loss, set to 8; each image-text sample has its unique quality; The total loss is expressed as the weighted sum of the inter-modal gravitational loss L inter and the intra-modal gravitational loss L intra ; where β is an adjustable parameter L = L inter + βL intra Equation (5-7) L = L inter + βL intra = L VT + βL V + βL T 。 2. The instance-level cross-modal retrieval method of the relationship inference and cross-modal independent matching network according to claim 1, characterized in that: Step S11: For the image modality, the model uses Faster R-CNN pre-trained on the Visual Genome dataset to extract visual features of 36 significant regions for each image. Specifically, the model first converts the input image into a feature map through the ResNet-101 backbone network, and then uses the Region Proposal Network (RPN) on the feature map to generate a number of candidate regions; then uses Region of Interest Pooling to map each candidate region to a feature vector of a fixed size, and this feature vector is the visual feature of the region. Step S12: For the text modality, first use the WordPiece tokenization algorithm based on the greedy algorithm to decompose words into the smallest sub-word units; then, input the tokenized text into the Transformer encoder, and the encoder encodes the sub-word sequence of each word to obtain the word embedding vector of the word; finally, add the position embedding and the segment embedding to the word embedding vector.
3. The instance-level cross-modal retrieval method of the relational reasoning and cross-modal independent matching network according to claim 1, characterized in that: When simplified calculations are required, take β = 0.5.
Citation Information
Cited By
Image-text matching method based on spatial position feature enhancement and deep semantic interaction
CN122654679A