A Structured Knowledge-Enhanced Image-Text Matching Method
By constructing co-occurring knowledge graphs and image knowledge structure graphs, extracting structured image knowledge features, and combining the features of images and text to match, the problem of insufficient use of knowledge and information in the existing technology is solved, and the effect of picture and text matching is improved.
Patent Information
- Application Number
- CN202210895904.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-07-27
AI Technical Summary
Existing graphic matching methods are difficult to effectively utilize knowledge information, especially the structural relationship between knowledge, resulting in poor matching effect.
By constructing a co-occurring knowledge graph, extracting high-frequency words as knowledge concepts, counting the co-occurring times of knowledge concepts, building a knowledge graph and obtaining knowledge concept features, combining the features of images and text, building an image knowledge structure graph, extracting structured image knowledge features, and fusing the knowledge features of images and text to match.
The accuracy and effect of picture and text matching are improved. By utilizing knowledge structure relationships, matching errors caused by the lack of structural information are reduced, and regional information and structural relationships are extracted more completely.
Smart Images

Figure CN115374289B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to cross-modal image-text matching technology, and specifically relates to an image-text matching method enhanced by structured knowledge. Background Art
[0002] Image-text matching is an important part of cross-modal matching. Specifically, it is to judge the matching relationship between image content and text content. The most direct applications of image-text matching are text retrieval based on images and image retrieval based on text. In addition, image-text matching can also serve many downstream tasks, such as image caption generation, visual question answering, etc., and has broad application prospects.
[0003] Existing image-text matching methods are divided into two categories according to whether external knowledge is added: image-text matching methods based on original features and image-text matching methods combined with knowledge, which are specifically described as follows:
[0004] I. Image-text matching methods based on original features are mainly divided into two directions:
[0005] (1) Mapping to a unified feature space for similarity calculation:
[0006] SCAN (Stacked cross attention for image-text matching) is a representative model of this kind. This model extracts the regional features of images and texts respectively, and uses two regional cross-attention modules for regional alignment to update the image-text features: one is the alignment module from image to text, which calculates the similarity between an image regional feature and all text regional features, and uses the similarity as a weight to perform weighted summation on the text regional features to obtain new image features. The other is the alignment module from text to image, which calculates the similarity between a text feature and all image features, and then uses the same algorithm as the alignment module from image to text to obtain new text features. Finally, the matching result is obtained by calculating the similarity of the new image-text features.
[0007] (2) Fusing the overall features of images and texts and then using a binary classification network to judge classification:
[0008] MTFN (Matching images and text with multi-modal tensor fusion and re-ranking) is a representative method of this kind. This method fuses the features of images and texts in the dimension of tensors, and uses a binary classification network to explicitly learn the image-text similarity function to judge the matching relationship between images and texts.
[0009] The above two image-text matching methods based on original features only use the features of images and texts themselves, without incorporating external knowledge, making it difficult to overcome the modality differences between images and texts, resulting in low accuracy.
[0010] II. Image-Text Matching Method Incorporating Knowledge:
[0011] Taking CVSE (Consensus-aware visual-semantic embedding for image-text matching) as a representative method. This method constructs a concept association graph for knowledge extraction by calculating the statistical co-occurrence associations between semantic concepts in the image caption corpus. Knowledge features are extracted from the original image and text features based on the knowledge vector to assist in the final matching calculation.
[0012] However, this method only extracts a set of knowledge concepts from images and texts, without considering the structural relationships between knowledge, resulting in incomplete knowledge information. Summary of the Invention
[0013] The technical problem to be solved by the present invention is: to propose a structured knowledge-enhanced image-text matching method, which fully utilizes knowledge to assist in matching by mining the structural relationship information between knowledge, thereby improving the effect of image-text matching.
[0014] The technical solution adopted by the present invention to solve the above technical problem is:
[0015] A structured knowledge-enhanced image-text matching method, comprising the following steps:
[0016] A. Training the image-text matching model:
[0017] A1. Constructing an image-text training data set, the image-text training data set includes a plurality of training data groups, each training data group includes a positive sample and two negative samples, the positive sample is composed of a correct image-text pair, and one of the two negative samples is an incorrect image-text pair composed of the image in the positive sample and a randomly selected incorrect text, and the other negative sample is an incorrect image-text pair composed of the text in the positive sample and a randomly selected incorrect image;
[0018] A2. Extracting high-frequency words as knowledge concepts from the image caption library composed of the texts of all positive samples, and counting the co-occurrence times of the knowledge concepts. A co-occurrence knowledge graph is constructed based on the co-occurrence times of the knowledge concepts, and then knowledge concept features are obtained based on the co-occurrence knowledge graph;
[0019] A3. Extracting features from the positive samples and their corresponding negative samples in the training data group. For each sample, the following steps are included:
[0020] A31. Extract features from the input image-text pair to obtain the features of each region of the image and the features of each region of the corresponding text.
[0021] A32. Based on the features of each region of the image obtained in step A31 and the knowledge concept features obtained in step A2, extract the regional visual knowledge features of the image; fuse the regional visual knowledge features of each image region with their corresponding regional spatial features to obtain the regional knowledge features of each image region; construct an image knowledge structure diagram according to the relationship between the regional knowledge features of the image, and then extract the structured image knowledge features according to the image knowledge structure diagram and the regional visual knowledge features.
[0022] Based on the features of each region of the text obtained in step A31 and the knowledge concept features obtained in step A2, extract the text knowledge features.
[0023] A33. Calculate the cosine similarity between the image knowledge features and the text knowledge features to obtain the cosine similarity of the positive samples and the cosine similarity of the corresponding negative samples respectively.
[0024] A4. Calculate the loss function according to the obtained cosine similarity of the positive samples and the cosine similarity of the corresponding negative samples, and then perform iterative training of steps A2 - A3 by the gradient descent method to generate an image-text matching model.
[0025] B. Perform image-text matching based on the image-text matching model:
[0026] Input an image and the text to be retrieved, calculate the similarity between the input image and all the texts to be retrieved through the image-text matching model, and determine the matched text according to the similarity.
[0027] Or, input a text and the images to be retrieved, calculate the similarity between the input text and all the images to be retrieved through the image-text matching model, and determine the matched images according to the similarity.
[0028] Further, in step A2, extract high-frequency words as knowledge concepts from the image caption library composed of all the texts of the positive samples, and count the co-occurrence times of the knowledge concepts, specifically including:
[0029] Count the occurrence times of all the words in the image caption library, and sort the words in descending order according to the occurrence times of the words. Select q words as knowledge concepts from high to low according to the standard of covering a preset proportion (such as 80%) of the total number of words.
[0030] Combine any two words in the knowledge concepts into word pairs, and count the co-occurrence times of each word pair in the image caption library.
[0031] Further, in step A2, constructing a co-occurrence knowledge graph according to the co-occurrence times of knowledge concepts specifically includes:
[0032] Taking knowledge concepts as nodes and the co-occurrence frequencies between knowledge concepts as edges to construct a co-occurrence knowledge graph; the co-occurrence frequency is the ratio of the co-occurrence times to the number of texts in the image caption library.
[0033] Further, in step A2, obtaining knowledge concept features based on the co-occurrence knowledge graph specifically includes:
[0034] Using the Glove model to map q knowledge concepts into initial knowledge concept features C = {c 1 , c 2 , …, c q}, and then based on the knowledge concept features C and the adjacency matrix P of the co-occurrence knowledge graph, using the GRU network to calculate the final knowledge concept features K = {k 1 , k 2 , …, k q}.
[0035] Further, in step A31, extracting features from the input image-text pair to obtain the features of each region of the image and the features of each region of the corresponding text specifically includes:
[0036] Using the Faster R-CNN model to extract n region features of the image in the image-text pair to obtain the regional visual features of the image Extracting the spatial features of each image region to obtain the regional spatial features of the image
[0037] Using the Bi-GRU network to extract m region features of the text in the image-text pair to obtain the text region features
[0038] Further, in step A32, based on the features of each region of the image obtained in step A31 and the knowledge features obtained in step A2, extracting the regional visual knowledge features of the image; fusing the regional visual knowledge features of each image region with its corresponding regional spatial features to obtain the regional knowledge features of each image region specifically includes:
[0039] First, based on the regional visual features of the image Calculating the similarity between each image region and all knowledge concept features, and its calculation formula is as follows:
[0040]
[0041] Among them, Denote the similarity between the visual feature of the \(i\)-th region and the knowledge concept feature \(k\) of the \(m\)-th m , Denote the visual feature of the \(i\)-th region, \(k\) m Denote the knowledge concept feature of the \(m\)-th, \(T\) represents matrix transpose, \(W\) I is the parameter matrix, and \(\beta\) is the amplification factor;
[0042] Then, using the similarity as the weight of each knowledge concept feature of the visual feature of each region, perform weighted summation according to the following formula to obtain the regional visual knowledge features of \(n\) image regions
[0043]
[0044] where, Denote the regional visual knowledge feature of the \(i\)-th image region;
[0045] Finally, according to the following formula, fuse the regional visual knowledge features of each image region and their regional spatial features to obtain the regional knowledge features of each image region containing spatial position information
[0046]
[0047] where, \(f\) concat represents the splicing operation.
[0048] Furthermore, in step A32, construct an image knowledge structure diagram according to the relationship between the regional knowledge features of the image, specifically including:
[0049] First, calculate the correlation between the regional knowledge feature of each region of the image and the regional knowledge features of other regions respectively to obtain the correlation matrix:
[0050] \(R = WV\) f (V f ) T
[0051] where, the element of \(R\) \(r\) ij is the element of the \(i\)-th row and \(j\)-th column of the correlation matrix \(R\), \(W\) is the learnable parameter matrix, \(w\) ij is the element of the \(i\)-th row and \(j\)-th column of \(W\);
[0052] Then, according to the correlation matrix, use the regional knowledge features of each region of the image as nodes and the correlation between it and the regional knowledge features of other regions of the image as edges to construct an image knowledge structure diagram.
[0053] Furthermore, in step A32, extract the structured image knowledge features according to the image knowledge structure diagram and the regional visual knowledge features, specifically including:
[0054] First, based on the adjacency matrix of the image knowledge structure diagram and the regional visual knowledge features They are fused through the GCN network to obtain the structured knowledge features of the image
[0055] Then, the structured knowledge features of the n regions of the image Are fused into an image knowledge feature v by using the method of taking the mean c .
[0056] Furthermore, in step A32, based on the regional features of the text obtained in step A31 and the knowledge features obtained in step A2, text knowledge features are extracted, specifically including:
[0057] First, the m text regional features of the text Are fused into a feature t by using the method of taking the mean r ;
[0058] Then, based on the feature t r , calculate its similarity with all knowledge concept features, and its calculation formula is as follows:
[0059]
[0060] Among them, a j Represents the similarity between t r And the j-th knowledge concept feature, k j Represents the j-th knowledge concept feature, T represents matrix transpose, W I Is the parameter matrix, and α is the amplification coefficient;
[0061] Finally, using the similarity between t r And each knowledge concept feature as the weight value of the corresponding knowledge concept feature, perform weighted summation on all knowledge concept features K = {k 1 , k 2 , …, k q} to obtain the text knowledge feature t c , and its calculation formula is as follows:
[0062]
[0063] Among them, a j Represents the weight value of the j-th knowledge concept feature, and k j Represents the j-th knowledge concept feature.
[0064] Further, in step A4, the loss function is calculated based on the cosine similarity of the obtained positive samples and the cosine similarity of their corresponding negative samples. Among them, the loss function adopts the triplet loss composed of a positive sample and two corresponding negative samples, which is specifically as follows:
[0065] Loss(v, t) = Σ (v,t) {[0, δ - d(v, t) + d(v - , t)] + max[0, δ - d(v, t) + d(v, t - )]}
[0066] Among them, d(v, t) is the calculation result of the cosine similarity of the positive samples in the training data set, d(v - , t) is the calculation result of the cosine similarity of a negative sample in the training data set, d(v, t - ) is the calculation result of the cosine similarity of another negative sample in the training data set, and δ is the margin parameter.
[0067] Further, in step B, the matching text determined according to the similarity includes: sorting the similarities between the input image and all texts to be retrieved from high to low, and selecting the top y texts in the similarity ranking as the matching results according to the actual scenario requirements;
[0068] The matching image determined according to the similarity includes: sorting the similarities between the input text and all images to be retrieved from high to low, and selecting the top y images in the similarity ranking as the matching results according to the actual scenario requirements.
[0069] The beneficial effects of the present invention are:
[0070] (1) Aiming at the problem of incorrect matching results caused by the lack of structured information in the process of image-text matching, the present invention designs a method for image-text matching enhanced by structured knowledge, extracts knowledge containing structured information from the image, and reduces matching errors caused by the lack of structural information;
[0071] (2) Aiming at the problem of ignoring the spatial information of image regions, the present invention adds spatial information to the information of the image regions themselves, so as to more completely extract the regional information. Furthermore, the structural relationship between regions is extracted to complete the image-text matching enhanced by structured knowledge and improve the matching effect. Brief Description of the Drawings
[0072] Figure 1 It is a flowchart for training the image-text matching model in the embodiment. Detailed Embodiment
[0073] The present invention aims to propose a method for enhancing image-text matching with structured knowledge. By mining the structural relationship information between knowledge, it makes full use of knowledge to assist in matching, thereby improving the effect of image-text matching. This method adds spatial information to the knowledge features of image regions, captures the relationships between image regions through the fused features, and constructs a regional knowledge structure graph. Through a graph convolutional neural network, regional knowledge features containing structured information are trained, and these regional features are fused into the structured knowledge features of the image. Finally, the structured knowledge features of the image and text are used for similarity calculation to complete the image-text matching with enhanced structured knowledge.
[0074] Embodiment:
[0075] The method for enhancing image-text matching with structured knowledge in this embodiment mainly includes two major parts: training an image-text matching model and performing image-text matching based on the image-text matching model, which are specifically described as follows:
[0076] I. Training of the image-text matching model:
[0077] The training process is as Figure 1 shown, and it includes steps such as text and image input, image-text feature extraction, external knowledge training, image-text knowledge feature extraction, structured knowledge extraction, feature fusion, similarity calculation, and gradient update. The specific introduction is as follows:
[0078] Step 1: Text and image input:
[0079] Download common image-text datasets such as MSCOCO or Flickr30k, etc. Each image and text in the image-text dataset has a corresponding number. The images and texts are correspondingly constructed into image-text pairs. Among them, the images and the correctly corresponding texts are constructed into correct image-text pairs as positive samples; each positive sample corresponds to two negative samples. One negative sample is an incorrect image-text pair constructed from the image in the positive sample and a randomly selected incorrectly corresponding text, and the other negative sample is an incorrect image-text pair constructed from the text in the positive sample and a randomly selected incorrectly corresponding image.
[0080] One positive sample and its corresponding two negative samples form a training data group; by processing the image-text dataset, multiple training data groups can be obtained for model training.
[0081] Since it is generally not possible to calculate all image-text pairs at one time during actual training, the training data needs to be input in batches according to the performance of the computer. For example, the entire training data is divided into multiple batches, and a certain number of training data groups are input in each batch.
[0082] Step 2: Image-text feature extraction:
[0083] Step 2.1: Image feature extraction:
[0084] Input a sample image with length N and width M into the Faster R-CNN model; the backbone network of the Faster R-CNN model consists of VGG-16, which includes 5 convolutional pooling layers. The structures of the first two convolutional pooling layers are sequentially connected as: conv_layer (convolutional layer), relu (activation function), conv_layer (convolutional layer), relu (activation function), pooling_layer (pooling layer); the structures of the third and fourth convolutional pooling layers are both sequentially connected as: conv_layer (convolutional layer), relu (activation function), conv_layer (convolutional layer), relu (activation function), conv_layer (convolutional layer), relu (activation function), pooling_layer (pooling layer), conv_layer (convolutional layer), relu (activation function), conv_layer (convolutional layer), relu (activation function). The structure of the fifth convolutional pooling layer is sequentially connected as: conv_layer (convolutional layer), relu (activation function). Among them, all convolutional operations use a 3*3 sliding window, with a stride of 1 and padding of 1; all pooling operations use a 2*2 sliding window, with a stride of 2 and padding of 0; for an image of size N*M*1, because the padding of the convolutional operation is 1, the convolutional layer does not change the length and width of the feature map, while the pooling operation reduces the length and width of the feature map to half of the original. Therefore, after the sample image is processed by the VGG-16 backbone network, a feature map of size (N / 16)*(M / 16)*512 is output.
[0085] Next, using the feature map with dimensions (N / 16)*(M / 16)*512 as the input, set the number of candidate boxes to 36. Through a 1*1*256*64 convolutional kernel, 64-dimensional and 128-dimensional vectors are obtained. The 64-dimensional vector represents the probability that 36 anchors (bounding boxes) are the background, and the 128-dimensional vector represents the four coordinate value information of 36 anchors, serving as the regional spatial features of the image. Among them, the number n of the regional spatial features is determined by the number of candidate bounding boxes set by the convolutional kernel.
[0086] Then, through RoI Pooling, downsampling is performed on the spatial regions represented by the 64-dimensional and 128-dimensional vectors, and visual features of 7*7*512 dimensions for each spatial region are output.
[0087] The 7*7*512-dimensional visual features can be reduced to one dimension by row using the flatten function of numpy, resulting in a one-dimensional vector of length 25088. Then, multiply the obtained one-dimensional vector by a learnable parameter matrix of size 1024*25088 to get a 1024-dimensional vector. 36 1024-dimensional vectors are used as regional visual features
[0088] Step 2.2, Text feature extraction:
[0089] First, preprocess the text in the input text-image pair data. The preprocessing includes: using the replace method of the string class to remove punctuation marks such as '-' and '.' in the text. Using the split method of the string class to separate the text with ' / ' as the delimiter and store it in an array.
[0090] Next, use the torchtext.vocab.Glove class in the Pytorch deep learning framework to obtain the word vector library trained by the glove model (Global Vectors for Word Representation). For each word in a text, extract the corresponding vector from the word vector library to form the initial vector representation of the text, and count the number of words contained in each text as vocab_length during this process.
[0091] Next, a text is input with the corresponding set of word vectors. First, it passes through an nn.embedding model network with 1024 dimensions for the number of words as parameters. Then, it passes through an nn.Dropout model network with a dropout of 0.35. Since the subsequent recurrent neural network requires vectors of the same size, all vectors are padded to the size of the longest vector through the pad_packed_sequence operation, and a vector sequence of the same size is output.
[0092] Next, the vector sequence passes through an nn.GRU model network with an input dimension of 300 and an embedding dimension of 1024. A 1024-dimensional vector of the number of words in the output is used as the text region feature
[0093] Step 3, External knowledge training:
[0094] Step 3.1, Extract high-frequency concepts from the image caption library and count the concept co-occurrence relationship:
[0095] The image caption library refers to the set of texts corresponding to all images in the graphic-text training dataset. In this step, first, count the occurrence times of all words in the caption library, and select the q words with the most occurrence times as knowledge concepts according to the goal of covering a certain proportion (such as 80%) of all words. Then, count the co-occurrence times of these words, and calculate the co-occurrence frequency of the word pair according to the ratio of the co-occurrence times of a pair of words to the number of texts.
[0096] The value of q here is based on the following: If fewer words are selected, the co-occurrence relationships contained in the caption library cannot be accurately extracted; if more words are selected, the calculation amount will increase. Due to the long-tail distribution of words, after statistics, the top 300 words in terms of occurrence frequency can cover 80% of all words. Therefore, the preferred value of q is 300.
[0097] Step 3.2: Construct a co-occurrence knowledge graph based on the co-occurrence relationships:
[0098] Take the knowledge concepts as the nodes in the node graph, and their co-occurrence frequencies as the edges. Generally speaking, the edges with higher co-occurrence frequencies can be regarded as a meaningful co-occurrence relationship, while some edges with lower co-occurrence frequencies can be regarded as accidental co-occurrence relationships and have no guiding significance and can be removed. Therefore, we construct a co-occurrence knowledge graph by removing the edges with smaller co-occurrence frequencies. Specifically, a confidence level of 0.3 can be set as the threshold for screening and filtering. And this confidence level can be adjusted, and 0.3 is just the optimal value selected after multiple experiments. After constructing the co-occurrence knowledge graph, describe the knowledge graph in the form of an adjacency list and store it in the adj file.
[0099] Step 3.3: Train knowledge features based on the co-occurrence knowledge graph:
[0100] First, call torchtext.vocab.Glove to obtain the word vector library trained by the glove model. And select the word vectors corresponding to 300 knowledge concepts as the initial knowledge concept features C = {c 1 , c 2 , …, c q}.
[0101] Then, input the 300 initial knowledge concept features and the adjacency matrix P of the co-occurrence knowledge graph obtained in the previous step into a three-layer convolutional neural network. The number of input channels is 300. The first layer is a 300 * 512 graph convolutional neural network; the second layer is a 512 * 512-dimensional graph convolutional neural network; the third layer is a 512 * 1024-dimensional graph convolutional neural network. Finally, output 300 trained 1024-dimensional knowledge concept features K = {k 1 , k 2 , …, k q}.
[0102] Step 4: Extraction of Graphic and Text Structured Knowledge Features
[0103] Step 4.1: Extract knowledge from text based on text region features and knowledge concept features
[0104] For the input and K = {k 1 , k 2 , …, k q}, first calculate the mean of to fuse m text region features into a 1*1024-dimensional t r feature. The specific calculation process is as follows
[0105]
[0106] Next, calculate the similarity between t r and all k j , and use this similarity to perform weighted summation on k 1 , k 2 , …, k q to obtain text knowledge features. The calculation operation process is to construct a learnable parameter matrix W I . For a k j , multiply the result of multiplying the 1*1024 t r by the 1024*1024 W I with the transpose 1024*1 of k j , k j T . Finally, obtain a single number result, multiply this number by α for amplification, and use it as the exponent of the exponential function with the natural constant e as the base to find the similarity of k j corresponding to t r . After finding the similarities corresponding to all k j , perform normalization to obtain the final weight a j of k j .
[0107] The mathematical description of this process is as follows
[0108]
[0109] where a j represents the similarity between t r and the j-th knowledge concept feature, k j represents the j-th knowledge concept feature, T represents matrix transpose, W I is the parameter matrix, α is the amplification coefficient, used to control the smoothness of the function and improve the generalization ability of the model, and is set to 10 according to relevant research experience
[0110] Next, use these weights to perform a weighted sum on the set \(K = \{k 1 , k 2 , \cdots, k q \}\) to obtain the knowledge feature \(t c \) corresponding to the text. The mathematical description is as follows:
[0111]
[0112] where \(a j \) represents the weight of the \(j\)-th knowledge concept feature, and \(k j \) represents the \(j\)-th knowledge concept feature.
[0113] Step 4.2: Extract knowledge in the image based on the original features and knowledge concept features of the image regions:
[0114] For the regional visual features of the input image and \(K = \{k 1 , k 2 , \cdots, k q \}\), calculate the similarity between each and all \(k m \), and use this similarity to perform a weighted sum on \(k 1 , k 2 , \cdots, k q \) to obtain the regional visual knowledge feature of this image region. The calculation process is also to construct a \(1024\times1024\) parameter matrix. For a \(k m \), multiply the result of multiplying the \(1\times1024\) by the \(1024\times1024\) parameter matrix with the transpose \(1024\times1\) of \(k m \), which is \(k m T \). Finally, obtain a scalar result, multiply this number by \(β\) for amplification, and use it as the exponent of the exponential function with the natural constant \(e\) as the base to find the similarity corresponding to \(k m \). After finding the similarities corresponding to all \(k \), perform normalization to obtain the final m weights. The calculation formula is as follows. where
[0115]
[0116] where represents the similarity between the \(i\)-th regional visual feature and the \(m\)-th knowledge concept feature \(k m \), represents the \(i\)-th regional visual feature, \(k m represents the \(m\)-th knowledge concept feature, \(T\) represents matrix transpose, and \(W Iis the parameter matrix, and β is the amplification factor, which is used to control the smoothness of the function and improve the generalization ability of the model. It is set to 10 according to relevant research experience;
[0117] These weights are used to perform weighted summation on K = {k 1 , k 2 , …, k q} to obtain the regional visual knowledge features corresponding to the image regions The calculation formula is as follows:
[0118]
[0119] Among them, represents the regional visual knowledge feature of the i-th image region;
[0120] Through the above calculation, the knowledge features of all image regions can be obtained
[0121] Next, using and the regional spatial features of the image obtained in step 2.1 as inputs, the 1024-dimensional and the 1024-dimensional are concatenated to obtain a 2048-dimensional i ∈ 1, 2, …, n, and the process is as follows:
[0122]
[0123] For each and the same operation is performed to obtain the image region knowledge features containing spatial position information
[0124] Step 4.3: Calculate the relationships between the image region knowledge features and construct an image knowledge structure diagram:
[0125] First, calculate the similarity between the image region knowledge features. For each pair of and calculate their product, and then multiply by a parameter w ij as and 's similarity r ij . The mathematical description of the calculation process is as follows:
[0126] R = WV f (V f ) T
[0127] Among them, the elements of R rij is the element at the i-th row and j-th column of the correlation matrix R, W is a learnable parameter matrix, and w ij is the element at the i-th row and j-th column of W;
[0128] Next, take and as the nodes of the graph, and r ij as the edges between the nodes. Filter these edges. The edges with a higher co-occurrence frequency can be regarded as a meaningful co-occurrence relationship, while some edges with a lower co-occurrence frequency can be regarded as accidental co-occurrence relationships and have no guiding significance, so they can be removed. Therefore, remove the r ij with a weaker relationship (smaller value), and retain the r ij with a stronger relationship (larger value). Specifically, set a confidence level of 0.3 as the threshold for screening and filtering. This confidence level can be adjusted, and 0.3 is the optimal value selected after multiple experiments. Save the final knowledge structure graph as an adj file.
[0129] Step 4.4: Train the structured knowledge features corresponding to the text and image according to the text-image knowledge structure graph:
[0130] First, input the visual knowledge features of the graph region and the adjacency matrix of the image knowledge structure graph obtained in the previous step into a three-layer graph convolutional neural network. The number of input channels is n. The first layer is a graph convolutional neural network of n * 2048; the second layer is a graph convolutional neural network of 2048 * 1024 dimensions; the third layer is a graph convolutional neural network of 1024 * 1024 dimensions. Output the image structured knowledge features
[0131] Next, is fused by the method of calculating the mean to obtain the knowledge feature v c corresponding to the image. The calculation process is as follows.
[0132]
[0133] So far, we have obtained the structured image knowledge feature v c and the text knowledge feature t c .
[0134] Step 5: Calculate the similarity:
[0135] In this step, calculate the cosine similarity between the image structure knowledge feature v c and the text knowledge feature t c as the basis for matching.
[0136] The mathematical description is:
[0137] sim(v c ,tc ) = cos(v c , t c ).
[0138] Step 6, Gradient update:
[0139] Perform gradient descent update based on the similarity calculation result and the loss function. Among them, the triplet loss is adopted for the loss function. For each training data group, it is specifically as follows:
[0140] Loss(v, t) = ∑ (v,t) {[0, δ - d(v, t) + d(v - , t)]
[0141] + max[0, δ - d(v, t) + d(v, t - )]}
[0142] Among them, d(v, t) is the cosine similarity calculation result of the positive sample in the training data group, v is the image in the sample, t is the text in the sample, d(v - , t) is the cosine similarity calculation result of a negative sample in the training data group, d(v, t - ) is the cosine similarity calculation result of another negative sample in the training data group, and δ is the margin parameter; it can be set to 0.2 according to experience.
[0143] During the iterative training process, sum up the losses of all training data groups in one round of input as the total loss of this round of input, then perform gradient descent update, and enter the next round of training; and so on. After completing the iterative training of the preset number of rounds or reaching the convergence condition, output the image - text matching model.
[0144] II. Image - text matching based on the image - text matching model:
[0145] After obtaining the image - text matching model, the matching model can be used for image - text retrieval. For example: input an image and the text to be retrieved, calculate the similarity between the input image and all texts to be retrieved through the image - text matching model, and determine the matching text according to the similarity; or, input a text and the images to be retrieved, calculate the similarity between the input text and all images to be retrieved through the image - text matching model, and determine the matching image according to the similarity.
[0146] Specifically, for a given image, calculate its similarity with all texts to be matched according to the image-text matching model, and then perform similarity ranking. If y matching texts are required, take the texts with the top y similarity rankings as the matching results. Similarly, for a given text, calculate its cosine similarity with all images to be matched, and then perform similarity ranking. If y matching images are required, take the images with the top y similarity rankings as the matching results. Thus, the image-text matching with enhanced structured knowledge is completed.
[0147] Although the present invention has been described herein with reference to its embodiments, the above embodiments are only preferred embodiments of the present invention, and the embodiments of the present invention are not limited to the above embodiments. It should be understood that those skilled in the art can design many other modifications and embodiments, and these modifications and embodiments will fall within the scope of the principles and spirit disclosed in this application.
Claims
1. A method for enhancing image-text matching with structured knowledge, characterized in that, it includes the following steps: A. Training an image-text matching model: A1. Construct an image-text training data set, which includes multiple training data groups. Each training data group includes a positive sample and two negative samples. The positive sample consists of a correct image-text pair. One of the two negative samples is an incorrect image-text pair formed by the image in the positive sample and a randomly selected incorrect text, and the other negative sample is an incorrect image-text pair formed by the text in the positive sample and a randomly selected incorrect image; A2. Extract high-frequency words as knowledge concepts from the image caption library composed of the texts of all positive samples, and count the co-occurrence times of the knowledge concepts. Construct a co-occurrence knowledge graph based on the co-occurrence times of the knowledge concepts, and then obtain knowledge concept features based on the co-occurrence knowledge graph; A3. Extract features from the positive sample and its corresponding negative sample in the training data group. For each sample, it includes the following steps: A31. Extract features from the input image-text pair to obtain the features of each region of the image and the features of each region of the corresponding text; A32. Based on the features of each region of the image obtained in step A31 and the knowledge concept features obtained in step A2, extract the regional visual knowledge features of the image; fuse the regional visual knowledge features of each image region with their corresponding regional spatial features to obtain the regional knowledge features of each image region; construct an image knowledge structure graph according to the relationship between the regional knowledge features of the image, and then extract structured image knowledge features according to the image knowledge structure graph and the regional visual knowledge features; Extract text knowledge features based on the features of each region of the text obtained in step A31 and the knowledge concept features obtained in step A2; A33. Calculate the cosine similarity between the image knowledge features and the text knowledge features, and obtain the cosine similarity of the positive sample and the cosine similarity of the corresponding negative sample respectively; A4. Calculate the loss function according to the obtained cosine similarity of the positive sample and its corresponding negative sample, and then perform iterative training of steps A2 - A3 by the gradient descent method to generate an image-text matching model; B. Perform image-text matching based on the image-text matching model; In step A32, constructing an image knowledge structure graph according to the relationship between the regional knowledge features of the image specifically includes: First, calculate the correlation between the knowledge feature of each region of the image and the knowledge features of other regions respectively to obtain a correlation matrix; Among them, , , is a learnable parameter matrix, is the element in the i-th row and j-th column of Then, based on the correlation matrix, use the knowledge features of each region of the image as nodes and the correlation between it and the knowledge features of other regions of the image as edges to construct an image knowledge structure graph; In step A32, extracting structured image knowledge features according to the image knowledge structure graph and the regional visual knowledge features specifically includes: First, based on the adjacency matrix of the image knowledge structure diagram and the regional visual knowledge features , they are fused through the GCN network to obtain the structured knowledge features of the image ; Then, the structured knowledge features of n regions of the image are fused into an image knowledge feature by using the method of taking the mean .
2. A method for enhancing image-text matching with structured knowledge as described in claim 1, characterized in that, in step A2, extracting high-frequency words as knowledge concepts from the image caption library composed of the texts of all positive samples, and counting the co-occurrence times of the knowledge concepts specifically includes: Count the occurrences of all words in the image caption library, and sort the words in descending order according to the number of occurrences. Select q words as knowledge concepts from high to low based on the standard of covering a preset proportion of the total number of words. Combine any two words in the knowledge concepts into word pairs, and count the co-occurrence times of each word pair in the image caption library.
3. A method for enhancing image-text matching with structured knowledge as described in claim 2, characterized in that, In step A2, construct a co-occurrence knowledge graph according to the co-occurrence times of knowledge concepts, specifically including: Construct a co-occurrence knowledge graph with knowledge concepts as nodes and the co-occurrence frequencies between knowledge concepts as edges; the co-occurrence frequency is the ratio of the co-occurrence times to the number of texts in the image caption library.
4. A method for enhancing image-text matching with structured knowledge as described in claim 2, characterized in that, In step A2, obtain knowledge concept features based on the co-occurrence knowledge graph, specifically including: Using the Glove model, map knowledge concepts into initial knowledge concept features , and then based on the knowledge concept features and the adjacency matrix of the co-occurrence knowledge graph , use the GRU network to calculate the final knowledge concept features .
5. A method for enhancing image-text matching with structured knowledge as described in claim 1, characterized in that, In step A31, extract features from the input image-text pair to obtain the features of each region of the image and the features of each region of the corresponding text, specifically including: Use the Faster R-CNN model to extract the regional features of the image in the text-image pair, and obtain the regional visual features of the image , extract the spatial features of each image region, and obtain the regional spatial features of the image ; Extract the regional features of the text in the text-image pair using the Bi-GRU network to obtain the text region features .
6. A method for enhancing image-text matching with structured knowledge as described in claim 5, characterized in that, In step A32, based on the features of each region of the image obtained in step A31 and the knowledge features obtained in step A2, extract the regional visual knowledge features of the image; fuse the regional visual knowledge features of each image region with its corresponding regional spatial features to obtain the regional knowledge features of each image region, specifically including: First, based on the regional visual features of the image , calculate the similarity between each image region and all knowledge concept features. The calculation formula is as follows: Among them, represents the similarity between the visual feature of the i-th region and the knowledge concept feature of the m-th , represents the visual feature of the i-th region, represents the knowledge concept feature of the m-th, T represents matrix transpose, is the parameter matrix, is the amplification factor; Then, using the similarity as the weight of each knowledge concept feature of the visual features of each region, weighted summation is performed according to the following formula to obtain the regional visual knowledge features of n image regions ; Among them, represents the regional visual knowledge feature of the i-th image region; Finally, according to the following formula, the regional visual knowledge features and regional spatial features of each image region are fused to obtain the regional knowledge features of each image region that contain spatial position information : Among them, represents a splicing operation.
7. A method for enhancing image-text matching with structured knowledge as described in claim 1, characterized in that, In step A32, based on the features of each region of the text obtained in step A31 and the knowledge features obtained in step A2, extract text knowledge features, specifically including: First, m text region features of the text are fused into one feature by using the method of taking the mean value ; Then, based on the feature , calculate its similarity with all knowledge concept features, and its calculation formula is as follows: Among them, denotes the similarity with the j-th knowledge concept feature, denotes the j-th knowledge concept feature, T denotes matrix transpose, is the parameter matrix, is the amplification factor; Finally, using the similarity with each knowledge concept feature as the weight value of the corresponding knowledge concept feature, for all knowledge concept features perform a weighted sum to obtain the text knowledge feature , and its calculation formula is as follows: Among them, represents the weight of the j-th knowledge concept feature, represents the j-th knowledge concept feature.
8. A method for enhancing image-text matching with structured knowledge as described in claim 1, characterized in that, In step A4, calculate the loss function according to the cosine similarity of the obtained positive samples and the cosine similarity of their corresponding negative samples, including: For a training data group, the loss function adopts the triplet loss composed of a positive sample and two corresponding negative samples, specifically as follows: Among them, is the cosine similarity calculation result of the positive sample in the training data set, is the cosine similarity calculation result of a negative sample in the training data set, is the cosine similarity calculation result of another negative sample in the training data set, is the margin parameter; During the iterative training process, sum up the losses of all training data groups input in one round as the total loss of the input in this round, and then perform gradient descent update to enter the next round of training.
9. A method for enhancing image-text matching with structured knowledge as described in claim 1, characterized in that, In step B, the image-text matching based on the image-text matching model is as follows: Input an image and a text to be retrieved, calculate the similarity between the input image and all texts to be retrieved through the image-text matching model, and determine the retrieved text according to the similarity; Or, input a text and an image to be retrieved, calculate the similarity between the input text and all images to be retrieved through the image-text matching model, and determine the retrieved image according to the similarity.
10. A method for enhancing image-text matching with structured knowledge as described in any one of claims 1-9, It is characterized in that In step B, the text matched according to the similarity includes: sorting the similarities between the input image and all texts to be retrieved from high to low, and selecting the top y texts in the similarity ranking as the matching results according to the actual scenario requirements; The image matched according to the similarity includes: sorting the similarities between the input text and all images to be retrieved from high to low, and selecting the top y images in the similarity ranking as the matching results according to the actual scenario requirements.
Citation Information
Patent Citations
Image-text matching method based on cross-modal mutual attention mechanism
CN114492646A
Method for constructing image text matching model based on prior knowledge graph
CN114547235A