A Cross-modal Retrieval Method and System Based on Multi-granularity Feature Fusion
By using a multi-grained feature fusion method in cross-modal retrieval, multiple fine-grained features are extracted and embedded, and feature fusion is used by self-attention mechanism, the problems of heterogeneous divide and overfitting between modals in the existing technology are solved, and the accuracy of cross-modal retrieval is significantly improved.
Patent Information
- Application Number
- CN202210901615.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-07-28
AI Technical Summary
The existing cross-modal retrieval methods have a heterogeneous gap between modals, making it difficult to effectively utilize local regional information and global information, resulting in low retrieval accuracy and prone to overfitting problems.
Using a cross-modal search method based on multi-grained feature fusion, the image fine-grained features and position fine-grained features of the image data are extracted, and the word fine-grained features of text data are embedded in the image fine-grained features, and the regional fine-grained features are obtained, and input them into the cross-modal search model. Multi-grained features are fused through the self-attention mechanism, and the local fine-grained similarity and fine-grained sum similarity are calculated, and the final loss function is constructed to optimize the model.
It effectively overcomes the heterogeneous gap in the cross-modal search method, takes into account local regional information and global information, integrates multi-grained features, significantly improves the accuracy of cross-modal search, and alleviates the overfitting phenomenon.
Smart Images

Figure CN115391625B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cross-modal two-way retrieval of graphics and texts, and more specifically, to a cross-modal retrieval method and system based on multi-granularity feature fusion. Background Art
[0002] With the rise of deep learning technology and the explosive growth of multi-modal data from the Internet, the research on the combination of multi-modal data and deep learning has gradually become a research hotspot in recent years. However, there is no natural connection between image visual features and text features themselves. Visual features are often represented as raw pixel arrays, and the information of each pixel point is recorded by three-channel RGB values; while text features often have a higher level of meaning, and a single word is generally represented by one-hot encoding. These two forms of features may have the same meaning, but their feature representations are extremely different. Therefore, there is a heterogeneous gap between these two modalities, making it difficult to match and retrieve between modalities. The research on cross-modal retrieval provides a solution to the above problems. It bridges the heterogeneous gap between different modalities by learning the feature representations of the two modalities in a common subspace and reducing their distances in the common subspace, so as to achieve cross-modal retrieval. Early work realized inter-modal retrieval by learning a network to embed coarse-grained features such as the entire picture and the entire sentence into a common subspace. However, coarse-grained features cannot well express the high-level semantic associations of local region details. In recent years, fine-grained methods based on local alignment have gradually become a research hotspot. Most methods design a two-branch network to guide image regions and words to align fine-grained information in a semi-supervised manner through an attention mechanism in the case of only global labels and no local labels, and achieve good results in many method experiments. However, this fine-grained method treats each object extracted by the object detection network and each word separated from the sentence equally when feeding them into the network, ignoring the difference in the importance degree of local regions and also ignoring the global information with important semantic connections. Therefore, the problems existing in the existing cross-modal retrieval methods lie in the heterogeneous gap between modalities, that is, how to use more effective information to improve the retrieval performance of the model, and at the same time how to solve the overfitting problem accompanying the model training process.
[0003] The prior art discloses a cross-modal retrieval method, device, storage medium, and terminal based on semantic enhancement. The method includes constructing a cross-modal retrieval model and training the cross-modal retrieval model based on an image-text retrieval data training set to obtain a trained cross-modal retrieval model; determining target query data and a target modal dataset, and obtaining the overall semantic similarity between the target query data and each target modal data based on the trained cross-modal retrieval model; selecting a preset number of target modal data corresponding to the overall semantic similarity from largest to smallest in the target modal dataset, and determining the retrieval result. This invention ignores the differences in the importance levels of local regions, does not utilize more effective information to improve the retrieval performance of the model, and has a low retrieval accuracy. Summary of the Invention
[0004] To overcome the defect of poor retrieval effect caused by the heterogeneous gap in the existing cross-modal retrieval method, the present invention provides a cross-modal retrieval method and system based on multi-granularity feature fusion, which overcomes the heterogeneous gap in the cross-modal retrieval method, takes into account local region information and global information at the same time, fuses multi-granularity features, and greatly improves the accuracy of cross-modal retrieval.
[0005] To solve the above technical problems, the technical solution of the present invention is as follows:
[0006] The present invention provides a cross-modal retrieval method based on multi-granularity feature fusion, including:
[0007] S1: Obtain a cross-modal dataset, where the cross-modal dataset includes corresponding image data and text data;
[0008] S2: Extract the image fine-grained features and position fine-grained features of the image data, and extract the word fine-grained features of the text data;
[0009] S3: Embed the position fine-grained features into the image fine-grained features to obtain region fine-grained features;
[0010] S4: Input the region fine-grained features and word fine-grained features into a constructed cross-modal retrieval model to extract visual modal features and text modal features;
[0011] S5: Calculate the local fine-grained similarity and the fine-grained total similarity according to the visual modal features and text modal features;
[0012] S6: Construct a final loss function according to the local fine-grained similarity and the fine-grained total similarity, and optimize with the goal of minimizing the final loss function to obtain a trained cross-modal retrieval model;
[0013] S7: Input the image data or text data to be retrieved into the trained cross-modal retrieval model for cross-modal retrieval to obtain the retrieval results.
[0014] Preferably, in the step S2, the method for extracting the image fine-grained feature and the position fine-grained feature of the image data is as follows:
[0015] Input the image data into an existing object detection network to extract the position of the target area in the image data and the coordinate information of the target box in the image data; obtain the image fine-grained feature I according to the position of the target area i ;
[0016] Calculate the position fine-grained feature O according to the coordinate information of the target box in the image data i ; First, calculate the overlap degree between two target boxes in the image data. The calculation formula is:
[0017]
[0018] In the formula, represents the overlap degree between the i-th target box and the j-th target box in the image, i, j ∈ [1, r], and r represents the number of target boxes in the image data; col ij represents the vertical overlap length of the overlapping area between the i-th target box and the j-th target box, and row ij represents the horizontal overlap length of the overlapping area between the i-th target box and the j-th target box; and represent the coordinates of the upper left corner and the lower right corner of the i-th target box respectively; and represent the coordinates of the upper left corner and the lower right corner of the j-th target box respectively;
[0019] Record the set of overlap degrees between the i-th target box and all target boxes as the position fine-grained feature O of this target box i , then where O i represents the position fine-grained feature of the i-th target box in the image data, and sum represents the total overlap degree between the i-th target box in the image data and all the remaining target boxes.
[0020] Preferably, in the step S2, the method for extracting the word fine-grained feature of the text data is as follows:
[0021] Input the text data into an existing BERT network to extract the word fine-grained feature t containing context semantic associations m , m ∈ [1, M], where M represents the number of words in the text data.
[0022] Preferably, in the step S3, the specific method for obtaining the region fine-grained feature is as follows:
[0023] Concatenate the image fine-grained feature I i and the position fine-grained feature O i to obtain the region fine-grained feature:
[0024] v i = linear([I i , O i ; θ f )
[0025] In the formula, v i represents the region fine-grained feature of the i-th target region, I i represents the image fine-grained feature of the i-th target region, linear represents the linear mapping operation, and θ f represents the linear mapping parameter.
[0026] Preferably, in the step S4, the constructed cross-modal retrieval model includes a parallel visual modality feature extraction unit and a text modality feature extraction unit; the visual modality feature extraction unit includes a first transformer encoder, a first linear mapping layer, a second transformer encoder, a first addition point, and a first normalization layer connected in sequence; the text modality feature extraction unit includes a second linear mapping layer, a third transformer encoder, a second addition point, and a second normalization layer connected in sequence;
[0027] The first transformer encoder, the second transformer encoder, and the third transformer encoder are all composed of several transformer encoding layers, and each transformer encoding layer has the same structure.
[0028] Preferably, in the step S4, the specific method for extracting the visual modality feature and the text modality feature is as follows:
[0029] For the image modality, input the region fine-grained feature into the first transformer encoder, and the processed output is mapped to the first intermediate fine-grained feature through the first linear mapping layer; input the first intermediate fine-grained feature into the second transformer encoder, and the second intermediate fine-grained feature is obtained after processing; sum the first intermediate fine-grained feature and the second intermediate fine-grained feature at the first addition point, and then input it into the first normalization layer for normalization processing to obtain the feature representation of the image modality in the common subspace, that is, the visual modality feature. The formula is:
[0030]
[0031] In the formula, Denote the visual modality features of the $i$-th target region. $TE1$ and $TE2$ respectively denote the processing operations of the first Transformer encoder and the second Transformer encoder. linear 1 Denote the linear mapping operation of the first linear mapping layer. norm 1 Denote the normalization operation of the first normalization layer;
[0032] The position of the target region and the coordinate information of the target box in the image data are a kind of fine-grained features. The regional fine-grained features contain the coordinate information of the target box in the image data and have rich image semantic associations, making the semantic information more sufficient in the process of image-text matching. Using the self-attention mechanism inside the Transformer encoder, a network combining coarse-fine fluency is formed. All the fine-grained features are weighted and summed to obtain the coarse-grained features. The coarse-grained features contain the information of all the fine-grained features, and the fine-grained features are extracted from the single local target region of the image data. The self-attention mechanism learns the fine-grained features and the coarse-grained features to obtain the multi-granularity fused visual modality features, improving the accuracy of subsequent retrieval.
[0033] For the text modality, the word fine-grained features are mapped by the second linear mapping layer and then input into the third Transformer encoder. After processing, the third intermediate fine-grained features are output. After summing the input and output of the third Transformer encoder at the second summing point, it is input into the second normalization layer for normalization processing to obtain the feature representation of the text modality in the common subspace, that is, the text modality features. The formula is:
[0034]
[0035] In the formula, Denote the $m$-th text modality feature. $TE3$ denotes the processing operation of the third Transformer encoder. t m Denote the $m$-th word fine-grained feature extracted from the text data through the BERT network. linear 2 Denote the linear mapping operation of the second linear mapping layer. norm 2 Denote the normalization operation of the second normalization layer.
[0036] In the image modality, set the input at the 0-th position of the first Transformer encoder to have the same dimension as v iEqual all-zero vectors. The global coarse-grained features obtained by the network during training fuse the fine-grained features of the input images at the remaining positions of the Transformer encoder by means of the self-attention mechanism inside the Transformer encoder, and at the same time perform internal multi-granularity feature fusion interaction through the self-attention mechanism. The input at the 0th position of the second Transformer encoder is the output at the 0th position of the first Transformer encoder.
[0037] In the text modality, set the input word id value at the 0th position of the BERT network to 0, fuse the fine-grained features of the input images at the remaining positions by means of the self-attention mechanism inside it, and at the same time perform internal multi-granularity feature fusion interaction through the self-attention mechanism; the input at the 0th position of the third Transformer encoder is the output at the 0th position of the BERT network.
[0038] By normalizing after summing the inputs and outputs of the Transformer encoder, the overfitting phenomenon of the cross-modal retrieval model can be effectively alleviated, and the retrieval accuracy can be improved.
[0039] Preferably, in the step S5, the method for calculating the local fine-grained similarity and the fine-grained total similarity according to the visual modality features and the text modality features is as follows:
[0040] Use the cosine similarity formula to calculate the local fine-grained similarity between the visual modality features and the text modality features. The formula is:
[0041]
[0042] In the formula, represents the local fine-grained similarity between the visual modality features and the text modality features; represents the transpose of, ||*|| represents the modulo operation;
[0043] According to the local fine-grained similarity between the visual modality features and the text modality features, calculate the fine-grained total similarity between the visual modality features and the text modality features. The formula is:
[0044]
[0045] In the formula, represents the fine-grained total similarity between the visual modality features and the text modality features.
[0046] Preferably, in the step S6, the specific method for constructing the final loss function according to the local fine-grained similarity and the fine-grained total similarity is as follows:
[0047] According to the fine-grained total similarity between the visual modality features and the text modality features Filter out cross-modal positive sample pairs and cross-modal negative sample pairs, and construct a triplet loss function:
[0048]
[0049] In the formula, B represents the number of image-text pairs in a mini-batch of data; θ represents the margin parameter, which is used to balance the influence of positive and negative sample pairs on the triplet loss function; represents the cross-modal positive sample pair, represents the cross-modal text negative sample pair, represents the cross-modal image negative sample pair;
[0050] Construct a global feature loss function:
[0051]
[0052] In the formula, k represents the adjustment parameter; V g , S g respectively represent and the 0th feature of, that is and This feature is located at the class-token position of the transformer encoder, and can weightedly aggregate the fine-grained features of the remaining input transformer encoder to obtain a global feature representation; represents the global similarity;
[0053] Construct the final loss function according to the triplet loss function and the global feature loss function:
[0054] Loss = L T + λL G
[0055] In the formula, λ represents the weight parameter.
[0056] By calculating the triplet loss function and the global feature loss function respectively, and then summing them to obtain the final loss function to train the cross-modal retrieval model, the convergence speed of the model is accelerated, and the retrieval accuracy is greatly improved.
[0057] Preferably, in the step S7, the specific method for obtaining the retrieval result is:
[0058] Input the image data or text data to be retrieved into the trained cross-modal retrieval model, calculate the local fine-grained similarity scores between the image data to be retrieved and text instances, and take the text instance with the highest local fine-grained similarity score as the most relevant text instance for the image data; or calculate the local fine-grained similarity scores between the text data to be retrieved and image instances, and take the image instance with the highest local fine-grained similarity score as the most relevant image instance for the text data.
[0059] The present invention also provides a cross-modal retrieval system based on multi-granularity feature fusion for implementing the above-mentioned cross-modal retrieval method based on multi-granularity feature fusion, including:
[0060] An image acquisition module for acquiring a cross-modal data set, where the cross-modal data set contains corresponding image data and text data;
[0061] A raw feature extraction module for extracting the image fine-grained features and position fine-grained features of the image data, and extracting the word fine-grained features of the text data;
[0062] A feature embedding module for embedding the position fine-grained features into the image fine-grained features to obtain regional fine-grained features;
[0063] A modal feature extraction module for inputting the regional fine-grained features and word fine-grained features into the constructed cross-modal retrieval model to extract visual modal features and text modal features;
[0064] A similarity calculation module for calculating local fine-grained similarities and fine-grained total similarities according to the visual modal features and text modal features;
[0065] A model training module for constructing a final loss function according to the local fine-grained similarities and fine-grained total similarities, optimizing with the goal of minimizing the final loss function, and obtaining a trained cross-modal retrieval model;
[0066] A cross-modal retrieval module for inputting the image data or text data to be retrieved into the trained cross-modal retrieval model to perform cross-modal retrieval and obtain retrieval results.
[0067] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0068] The present invention first obtains a cross-modal dataset, extracts the image fine-grained features and location fine-grained features of the image data, and the word fine-grained features of the text data; then embeds the location fine-grained features into the image fine-grained features to obtain regional fine-grained features; then inputs the regional fine-grained features and word fine-grained features into the constructed cross-modal retrieval model to extract visual modal features and text modal features; the location fine-grained features have rich image semantic association information, and the semantic information is increased by embedding the location fine-grained features; the self-attention mechanism in the cross-modal retrieval model weights and sums all the fine-grained features into coarse-grained features, and then fuses the fine-grained features with the coarse-grained features. Finally, the obtained visual modal features and text modal features have the characteristics of multi-granularity feature fusion and sufficient semantic information; next, calculate the local fine-grained similarity and the fine-grained total similarity, and construct the final loss function based on this. Optimize with the goal of minimizing the final loss function to obtain a trained cross-modal retrieval model, which speeds up the convergence speed of the model; finally, input the image data or text data to be retrieved into the trained cross-modal retrieval model to retrieve the text most relevant to the image data or the image most relevant to the text data. The present invention overcomes the heterogeneous gap existing in the cross-modal retrieval method, takes into account both local region information and global information, and fuses multi-granularity features, greatly improving the accuracy of cross-modal retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 FIG. is a flowchart of a cross-modal retrieval method based on multi-granularity feature fusion according to Embodiment 1;
[0070] Figure 2 FIG. is a flowchart of a cross-modal retrieval method based on multi-granularity feature fusion according to Embodiment 2;
[0071] Figure 3 FIG. is a schematic diagram of the retrieval result of image query text according to Embodiment 2;
[0072] Figure 4 FIG. is a schematic diagram of the retrieval result of text query image according to Embodiment 2;
[0073] Figure 5 FIG. is a schematic structural diagram of a cross-modal retrieval system based on multi-granularity feature fusion according to Embodiment 3. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0075] For better illustration of this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the dimensions of the actual product;
[0076] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0077] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.
[0078] Embodiment 1
[0079] This embodiment provides a cross-modal retrieval method based on multi-granularity feature fusion, as Figure 1 shown, including:
[0080] S1: Obtain a cross-modal data set, where the cross-modal data set contains corresponding image data and text data;
[0081] S2: Extract the image fine-grained features and location fine-grained features of the image data, and extract the word fine-grained features of the text data;
[0082] S3: Embed the location fine-grained features into the image fine-grained features to obtain regional fine-grained features;
[0083] S4: Input the regional fine-grained features and word fine-grained features into the constructed cross-modal retrieval model to extract visual modal features and text modal features;
[0084] S5: Calculate the local fine-grained similarity and the fine-grained total similarity according to the visual modal features and the text modal features;
[0085] S6: Construct a final loss function according to the local fine-grained similarity and the fine-grained total similarity, and optimize with the goal of minimizing the final loss function to obtain a trained cross-modal retrieval model;
[0086] S7: Input the image data or text data to be retrieved into the trained cross-modal retrieval model for cross-modal retrieval to obtain retrieval results.
[0087] In the specific implementation process, the method provided in this embodiment first obtains a cross-modal dataset, extracts the image fine-grained features and location fine-grained features of the image data, and the word fine-grained features of the text data; then embeds the location fine-grained features into the image fine-grained features to obtain region fine-grained features; then inputs the region fine-grained features and word fine-grained features into the constructed cross-modal retrieval model to extract visual modal features and text modal features; the location fine-grained features have rich image semantic association information, and the semantic information is increased by embedding the location fine-grained features; the self-attention mechanism in the cross-modal retrieval model weights and sums all the fine-grained features into coarse-grained features, and then fuses the fine-grained features with the coarse-grained features. Finally, the obtained visual modal features and text modal features have the characteristics of multi-granularity feature fusion and sufficient semantic information; next, calculate the local fine-grained similarity and the fine-grained total similarity, and construct the final loss function based on this. Optimize with the goal of minimizing the final loss function to obtain a trained cross-modal retrieval model, which speeds up the convergence rate of the model; finally, input the image data or text data to be retrieved into the trained cross-modal retrieval model to retrieve the text most relevant to the image data or the image most relevant to the text data. The present invention overcomes the heterogeneous gap existing in the cross-modal retrieval method, takes into account both local region information and global information, and fuses multi-granularity features, greatly improving the accuracy of cross-modal retrieval.
[0088] Embodiment 2
[0089] This embodiment provides a cross-modal retrieval method based on multi-granularity feature fusion, including:
[0090] S1: Obtain a cross-modal dataset, where the cross-modal dataset contains corresponding image data and text data;
[0091] S2: Extract the image fine-grained features and location fine-grained features of the image data, and extract the word fine-grained features of the text data; specifically:
[0092] The method for extracting the image fine-grained features and location fine-grained features of the image data is:
[0093] Input the image data into an existing object detection network to extract the target region position in the image data and the coordinate information of the target box in the image data; obtain the image fine-grained feature I according to the target region position i ;
[0094] Calculate the location fine-grained feature O according to the coordinate information of the target box in the image data i ; First, calculate the overlap degree between two target boxes in the image data, and the calculation formula is:
[0095]
[0096] In the formula, represents the overlap degree between the i-th target box and the j-th target box in the image, where i, j ∈ [1, r], and r represents the number of target boxes in the image data; col ij represents the vertical overlap length of the overlapping area between the i-th target box and the j-th target box, and row ij represents the horizontal overlap length of the overlapping area between the i-th target box and the j-th target box; and represent the coordinates of the upper left corner and the lower right corner of the i-th target box respectively; and represent the coordinates of the upper left corner and the lower right corner of the j-th target box respectively;
[0097] Denote the set of overlap degrees between the i-th target box and all target boxes as the position fine-grained feature O of this target box i , then where O i represents the position fine-grained feature of the i-th target box in the image data, and sum represents the total overlap degree between the i-th target box in the image data and all the remaining target boxes;
[0098] The method for extracting the word fine-grained feature of the text data is as follows:
[0099] Input the text data into the existing BERT network to extract the word fine-grained feature t containing context semantic associations m , where m ∈ [1, M], and M represents the number of words in the text data.
[0100] S3: Embed the position fine-grained feature into the image fine-grained feature to obtain the region fine-grained feature; the specific method is as follows:
[0101] Concatenate the image fine-grained feature I i and the position fine-grained feature O i to obtain the region fine-grained feature:
[0102] v i = linear([I i , O i ; θ f )
[0103] In the formula, v i represents the region fine-grained feature of the i-th target region, I i represents the image fine-grained feature of the i-th target region, linear represents the linear mapping operation, and θ f represents the linear mapping parameter.
[0104] S4: Input the region fine-grained features and word fine-grained features into the constructed cross-modal retrieval model to extract visual modal features and text modal features;
[0105] As Figure 2 shown, the constructed cross-modal retrieval model includes a parallel visual modal feature extraction unit and a text modal feature extraction unit; the visual modal feature extraction unit includes a first transformer encoder, a first linear mapping layer, a second transformer encoder, a first summing point, and a first normalization layer connected in sequence; the text modal feature extraction unit includes a second linear mapping layer, a third transformer encoder, a second summing point, and a second normalization layer connected in sequence;
[0106] The first transformer encoder, the second transformer encoder, and the third transformer encoder are each composed of a number of transformer encoding layers, and the structure of each transformer encoding layer is the same; in this embodiment, the first transformer encoder includes 4 transformer encoding layers connected in sequence, and the second transformer encoder and the third transformer encoder include 4 transformer encoding layers connected in sequence.
[0107] For the image modality, combine the region fine-grained features and input them into the first transformer encoder. The processed output is mapped to a first intermediate fine-grained feature through the first linear mapping layer; input the first intermediate fine-grained feature into the second transformer encoder to obtain a second intermediate fine-grained feature after processing; sum the first intermediate fine-grained feature and the second intermediate fine-grained feature at the first summing point and then input them into the first normalization layer for normalization processing to obtain the feature representation of the image modality in the common subspace, that is, the visual modal feature. The formula is:
[0108]
[0109] In the formula, represents the visual modal feature of the i-th target region, TE1 and TE2 respectively represent the processing operations of the first transformer encoder and the second transformer encoder, and linear 1 represents the linear mapping operation of the first linear mapping layer, and norm 1 represents the normalization operation of the first normalization layer;
[0110] The coordinate information of the target region position and the target bounding box in the image data is a fine-grained feature. The regional fine-grained feature contains the coordinate information of the target bounding box in the image data, has rich image semantic associations, and makes the semantic information more sufficient in the text-image matching process. By using the self-attention mechanism inside the Transformer encoder, a network combining coarse-fine fluency is formed. All fine-grained features are weighted and summed to obtain the coarse-grained feature, which contains the information of all fine-grained features. The fine-grained features are extracted from the single local target region of the image data. The self-attention mechanism learns the fine-grained and coarse-grained features to obtain the visually modal features of multi-granularity fusion, improving the accuracy of subsequent retrieval.
[0111] For the text modality, the word fine-grained features are mapped by the second linear mapping layer and then input into the third Transformer encoder. After processing, the third intermediate fine-grained features are output. After summing the input and output of the third Transformer encoder at the second summing point, they are input into the second normalization layer for normalization processing to obtain the feature representation of the text modality in the common subspace, that is, the text modality feature. The formula is:
[0112]
[0113] In the formula, represents the m-th text modality feature, TE3 represents the processing operation of the third Transformer encoder, and t m represents the m-th word fine-grained feature extracted from the text data through the BERT network, and linear 2 represents the linear mapping operation of the second linear mapping layer, and norm 2 represents the normalization operation of the second normalization layer.
[0114] By summing the input and output of the Transformer encoder and then performing normalization processing, the overfitting phenomenon of the cross-modal retrieval model can be effectively alleviated, and the accuracy of retrieval can be improved.
[0115] S5: Calculate the local fine-grained similarity and the fine-grained total similarity according to the visual modal features and the text modal features;
[0116] The cosine similarity formula is used to calculate the local fine-grained similarity between the visual modal features and the text modal features. The formula is:
[0117]
[0118] In the formula, represents the local fine-grained similarity between the visual modal features and the text modal features; represents the transpose of, and ||*|| represents the modulus operation;
[0119] Calculate the fine-grained total similarity of the visual modality features and the text modality features according to the local fine-grained similarity between the visual modality features and the text modality features. The formula is as follows:
[0120]
[0121] In the formula, represents the fine-grained total similarity of the visual modality features and the text modality features.
[0122] S6: Construct a final loss function based on the local fine-grained similarity and the fine-grained total similarity, and optimize it with the goal of minimizing the final loss function to obtain a trained cross-modal retrieval model;
[0123] According to the fine-grained total similarity of the visual modality features and the text modality features Filter out cross-modal positive sample pairs and cross-modal negative sample pairs, and construct a triplet loss function:
[0124]
[0125] In the formula, B represents the number of image-text pairs in a small batch of data; θ represents the margin parameter, which is used to balance the influence of the positive and negative sample pair triplet loss functions; represents a cross-modal positive sample pair, represents a cross-modal text negative sample pair, represents a cross-modal image negative sample pair;
[0126] Construct a global feature loss function:
[0127]
[0128] In the formula, k represents a tuning parameter; V g ,S g respectively represent and The 0th feature of, that is, and This feature is located at the class-token position of the transformer encoder and can weighted-aggregate the fine-grained features of the remaining input to the transformer encoder to obtain a global feature representation; represents the global similarity;
[0129] Construct a final loss function according to the triplet loss function and the global feature loss function:
[0130] Loss = L T + λL G
[0131] Where λ represents the weight parameter.
[0132] By separately calculating the triplet loss function and the global feature loss function, and then summing them to obtain the final loss function to train the cross-modal retrieval model, the convergence speed of the model is accelerated, and the retrieval accuracy is greatly improved.
[0133] S7: Input the image data or text data to be retrieved into the trained cross-modal retrieval model for cross-modal retrieval to obtain the retrieval result.
[0134] Input the image data or text data to be retrieved into the trained cross-modal retrieval model, calculate the local fine-grained similarity score between the image data to be retrieved and the text instance, and take the text instance with the highest local fine-grained similarity score as the most relevant text instance for this image data; or calculate the local fine-grained similarity score between the text data to be retrieved and the image instance, and take the image instance with the highest local fine-grained similarity score as the most relevant image instance for this text data.
[0135] In the specific implementation process, the method proposed in this embodiment (Ours) was evaluated using the MS-COCO and Flickr30k datasets; the MS-COCO and Flickr30k datasets are cross-modal datasets. The MS-COCO dataset is a large dataset with 123,287 images, and each image has five manually annotated text annotations to explain the content shown in the image; the segmentation of the MS-COCO dataset follows the following operations: 113,287 images in the dataset are divided into the training set, 5,000 images are used as the validation set, and 5,000 images are used as the test set. When testing the MS-COCO dataset, the results of 5,000 images and 1,000 images were tested at the same time. The test results of the 1,000 images were obtained by splitting the 5,000-image test set into five 1,000-image test subsets and then averaging the five test results. Compared with the MS-COCO dataset, the Flickr30K dataset has much less data. The Flickr30K dataset contains 31,000 images and 158,915 English annotations. Similar to the MS-COCO dataset, each image in the Flickr30K dataset is annotated with at least five sentences. The segmentation of the dataset follows the following operations: 29,000 images in the dataset are divided into the training set, 1,000 images are used as the validation set, and 1,000 images are used as the test set; the algorithms for effect comparison are fine-grained cross-modal retrieval methods, specifically: SCAN algorithm, CASC algorithm, PFAN algorithm, PFAN++ algorithm, MMCA algorithm, AAMEL algorithm, SMAN algorithm, M3A-Net algorithm, IMRAN algorithm, TERAN algorithm, MLMN algorithm, and RRTC algorithm. The retrieval methods are divided into image query text (Sentence Retrieval) and text query image (Image Retireval). The evaluation result index is the recall rate K (R@K), which represents the proportion of correct results among the K closest results retrieved, that is, the percentage of correct query results among the top K results; three recall rate indicators are set, R@1, R@5, and R@10, which represent the percentages of correct query results among the top 1, 5, and 10 results respectively; and the mean recall rate mR and the total recall rate R@sum are set. The mean recall rate mR reflects the average of R@1, R@5, and R@10; the total recall rate R@sum reflects the sum of R@1, R@5, and R@10 in both the image query text and text query image methods; the test results are shown in the following table:
[0136]
[0137] As can be seen from the above table, whether the method provided in this embodiment queries text by image or queries image by text, the recall rate indicators R@1, R@5, and R@10 are the highest among the comparison algorithms, indicating that the method provided in this embodiment has the best retrieval effect. Specifically, as Figure 3 shown, it is a schematic diagram of the retrieval result for querying text by image. The image data to be retrieved is Figure 3 the left picture in. Among the 10 text instances of the retrieval result, the text instance ranked first is "A woman sitting on a bench next to someone wearing a hat and sitting in a wheelchair", and this text instance perfectly describes the image data; as Figure 4 shown, it is a schematic diagram of the retrieval result for querying image by text. The text data to be retrieved is "Two men on a rooftop while another man stands atop a ladder watching them". Among the 10 image instances of the retrieval result, the image instance ranked first is two men standing on the rooftop and another man standing on the ladder watching them, which completely matches the text to be retrieved; from the above experimental results, it can be seen that the cross-modal retrieval method based on multi-granularity feature fusion proposed in this embodiment effectively improves the accuracy of cross-modal retrieval.
[0138] Embodiment 3
[0139] This embodiment provides a cross-modal retrieval system based on multi-granularity feature fusion for implementing the cross-modal retrieval method based on multi-granularity feature fusion described in Embodiment 1 or 2, as Figure 5 shown, including:
[0140] An image acquisition module for acquiring a cross-modal data set, where the cross-modal data set contains corresponding image data and text data;
[0141] A raw feature extraction module for extracting the image fine-grained feature and the location fine-grained feature of the image data, and extracting the word fine-grained feature of the text data;
[0142] A feature embedding module for embedding the location fine-grained feature into the image fine-grained feature to obtain a region fine-grained feature;
[0143] A modal feature extraction module for inputting the region fine-grained feature and the word fine-grained feature into a constructed cross-modal retrieval model to extract the visual modal feature and the text modal feature;
[0144] A similarity calculation module for calculating the local fine-grained similarity and the fine-grained total similarity according to the visual modal feature and the text modal feature;
[0145] A model training module, configured to construct a final loss function according to the local fine-grained similarity and the fine-grained total similarity, optimize with the goal of minimizing the final loss function, and obtain a trained cross-modal retrieval model;
[0146] A cross-modal retrieval module, configured to input the image data or text data to be retrieved into the trained cross-modal retrieval model for cross-modal retrieval to obtain retrieval results.
[0147] The same or similar reference numerals correspond to the same or similar components;
[0148] The terms describing the positional relationship in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent;
[0149] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A cross-modal retrieval method based on multi-granularity feature fusion, characterized in that, it includes the steps of: S1: Obtain a cross-modal dataset, where the cross-modal dataset contains corresponding image data and text data; S2: Extract the image fine-grained features and location fine-grained features of the image data, and extract the word fine-grained features of the text data; The method for extracting the image fine-grained features and location fine-grained features of the image data is: Input the image data into an existing object detection network to extract the location of the target area in the image data and the coordinate information of the target box in the image data; Obtain the fine-grained feature I of the image according to the target region location i ; Calculate the position fine-grained feature O according to the coordinate information of the target box in the image data i ; First, calculate the overlap degree between two target boxes in the image data. The calculation formula is as follows: Wherein, represents the overlap degree between the i-th target box and the j-th target box in the image, where i, j ∈ [1, r], and r represents the number of target boxes in the image data; col ij represents the vertical overlap length of the overlapping area between the i-th target box and the j-th target box, and row ij represents the horizontal overlap length of the overlapping area between the i-th target box and the j-th target box; and respectively represent the coordinates of the upper left corner and the lower right corner of the i-th target box; and respectively represent the coordinates of the upper left corner and the lower right corner of the j-th target box; Let the set of overlap degrees between the $i$-th target box and all target boxes be denoted as the position fine-grained feature $O$ of this target box i , then where $O$ i represents the position fine-grained feature of the $i$-th target box in the image data, and sum represents the total overlap degree between the $i$-th target box in the image data and all the remaining target boxes; S3: Embed the location fine-grained features into the image fine-grained features to obtain regional fine-grained features; S4: Input the regional fine-grained features and word fine-grained features into the constructed cross-modal retrieval model to extract visual modal features and text modal features; S5: Calculate the local fine-grained similarity and the fine-grained total similarity according to the visual modal features and text modal features; S6: Construct a final loss function according to the local fine-grained similarity and the fine-grained total similarity, and optimize with the goal of minimizing the final loss function to obtain a trained cross-modal retrieval model; S7: Input the image data or text data to be retrieved into the trained cross-modal retrieval model for cross-modal retrieval to obtain retrieval results.
2. The cross-modal retrieval method based on multi-granularity feature fusion according to claim 1, characterized in that, in the step S2, the method for extracting the word fine-grained features of the text data is: Input the text data into the existing BERT network to extract the fine-grained features t of words containing context semantic associations m , where m ∈ [1, M] and M represents the number of words in the text data.
3. The cross-modal retrieval method based on multi-granularity feature fusion according to claim 1, characterized in that, in the step S3, the specific method for obtaining the regional fine-grained features is: Concatenate the image fine-grained feature I i and the location fine-grained feature O i in series to obtain the region fine-grained feature: v i = linear([I i ,O i ; θ f ) where, v i represents the regional fine-grained feature of the i-th target region, I i represents the image fine-grained feature of the i-th target region, linear represents the linear mapping operation, θ f represents the linear mapping parameter.
4. The cross-modal retrieval method based on multi-granularity feature fusion according to claim 2 or 3, characterized in that, in the step S4, the constructed cross-modal retrieval model includes a parallel visual modal feature extraction unit and a text modal feature extraction unit; the visual modal feature extraction unit includes a first transformer encoder, a first linear mapping layer, a second transformer encoder, a first addition point, and a first normalization layer connected in sequence; the text modal feature extraction unit includes a second linear mapping layer, a third transformer encoder, a second addition point, and a second normalization layer connected in sequence; The first transformer encoder, the second transformer encoder, and the third transformer encoder are all composed of several transformer encoding layers, and the structure of each transformer encoding layer is the same.
5. The cross-modal retrieval method based on multi-granularity feature fusion according to claim 4, characterized in that, in the step S4, the specific method for extracting the visual modal features and text modal features is: For the image modality, the region fine-grained features are input into the first Transformer encoder, and the processed output is mapped to the first intermediate fine-grained features through the first linear mapping layer; the first intermediate fine-grained features are input into the second Transformer encoder, and the second intermediate fine-grained features are obtained after processing; After summing the first intermediate fine-grained features and the second intermediate fine-grained features at the first summation point, they are input into the first normalization layer for normalization processing to obtain the feature representation of the image modality in the common subspace, that is, the visual modality features. The formula is: In the formula, represents the visual modal feature of the i-th target region. TE1 and TE2 respectively represent the processing operations of the first transformer encoder and the second transformer encoder. linear 1 represents the linear mapping operation of the first linear mapping layer. norm 1 represents the normalization operation of the first normalization layer. v i represents the region fine-grained feature of the i-th target region; For the text modality, the word fine-grained features are mapped through the second linear mapping layer and then input into the third Transformer encoder. After processing, the third intermediate fine-grained features are output; after summing the input and output of the third Transformer encoder at the second summation point, they are input into the second normalization layer for normalization processing to obtain the feature representation of the text modality in the common subspace, that is, the text modality features. The formula is: In the formula, represents the m-th text modal feature, TE3 represents the processing operation of the third Transformer encoder, and t m represents the m-th word fine-grained feature extracted from the text data through the BERT network, and linear 2 represents the linear mapping operation of the second linear mapping layer, and norm 2 represents the normalization operation of the second normalization layer.
6. The cross-modal retrieval method based on multi-granularity feature fusion according to claim 4, wherein, in the step S5, the method for calculating the local fine-grained similarity and the fine-grained total similarity according to the visual modality features and the text modality features is: The local fine-grained similarity between the visual modality features and the text modality features is calculated using the cosine similarity formula. The formula is: In the formula, represents the local fine-grained similarity between the visual modality feature and the text modality feature; represents the transpose of, ||*|| represents the modulus operation, represents the visual modality feature of the i-th target region, represents the m-th text modality feature; According to the local fine-grained similarity between the visual modality features and the text modality features, the fine-grained total similarity between the visual modality features and the text modality features is calculated. The formula is: In the formula, represents the fine-grained total similarity between the visual modality features and the text modality features, M represents the number of words in the text data, and r represents the number of target boxes in the image data.
7. The cross-modal retrieval method based on multi-granularity feature fusion according to claim 5, wherein, in the step S6, the specific method for constructing the final loss function according to the local fine-grained similarity and the fine-grained total similarity is: Fine-grained total similarity based on visual modality features and text modality features Filter out cross-modal positive sample pairs and cross-modal negative sample pairs, and construct a triplet loss function: Wherein, B represents a cross-modal data set of a training batch, θ represents a marginal parameter; (I, S) represents a cross-modal positive sample pair, represents a cross-modal text negative sample pair, represents a cross-modal image negative sample pair; Construct a global feature loss function: where k represents a regulation parameter; V g , S g respectively represent and the 0th feature of, l(V g , S g ) represents the global similarity; Construct a final loss function according to the triplet loss function and the global feature loss function: Loss=L T +λL G In the formula, λ represents a weight parameter.
8. The cross-modal retrieval method based on multi-granularity feature fusion according to claim 1, wherein, in the step S7, the specific method for obtaining the retrieval result is: Input the image data or text data to be retrieved into the trained cross-modal retrieval model, calculate the local fine-grained similarity score between the image data to be retrieved and the text instances, and take the text instance with the highest local fine-grained similarity score as the text instance most relevant to the image data; or calculate the local fine-grained similarity score between the text data to be retrieved and the image instances, and take the image instance with the highest local fine-grained similarity score as the image instance most relevant to the text data.
9. A cross-modal retrieval system based on multi-granularity feature fusion, wherein, for implementing the cross-modal retrieval method based on multi-granularity feature fusion according to any one of claims 1-8, including: An image acquisition module for acquiring a cross-modal data set, where the cross-modal data set contains corresponding image data and text data; A raw feature extraction module for extracting the image fine-grained features and location fine-grained features of the image data, and extracting the word fine-grained features of the text data; Extract the image fine-grained features and location fine-grained features of the image data, including: Input the image data into an existing object detection network to extract the position of the target region in the image data and the coordinate information of the target box in the image data; obtain the fine-grained image feature I based on the position of the target region i ; Calculate the position fine-grained feature O according to the coordinate information of the target box in the image data i ; First, calculate the overlap degree between two target boxes in the image data, and the calculation formula is: Wherein, represents the overlap degree between the i-th target box and the j-th target box in the image, where i, j ∈ [1, r], and r represents the number of target boxes in the image data; col ij represents the vertical overlap length of the overlapping region between the i-th target box and the j-th target box, and row ij represents the horizontal overlap length of the overlapping region between the i-th target box and the j-th target box; and respectively represent the coordinates of the upper left corner and the lower right corner of the i-th target box; and respectively represent the coordinates of the upper left corner and the lower right corner of the j-th target box; Let the set of overlap degrees between the $i$-th target box and all target boxes be denoted as the position fine-grained feature $O$ of this target box i , then where $O$ i represents the position fine-grained feature of the $i$-th target box in the image data, and sum represents the total overlap degree between the $i$-th target box in the image data and all the remaining target boxes; A feature embedding module for embedding the location fine-grained features into the image fine-grained features to obtain region fine-grained features; A modality feature extraction module for inputting the region fine-grained features and word fine-grained features into a constructed cross-modal retrieval model to extract visual modality features and text modality features; A similarity calculation module for calculating the local fine-grained similarity and the fine-grained total similarity according to the visual modality features and the text modality features; A model training module for constructing a final loss function according to the local fine-grained similarity and the fine-grained total similarity, optimizing with the goal of minimizing the final loss function, and obtaining a trained cross-modal retrieval model; A cross-modal retrieval module for inputting the image data or text data to be retrieved into the trained cross-modal retrieval model for cross-modal retrieval to obtain retrieval results.
Citation Information
Patent Citations
Cross-mode pedestrian re-identification method and system based on a heterogeneous hierarchical attention mechanism
CN109829430A
Image-text matching method based on cross-modal mutual attention mechanism
CN114492646A