Combined image retrieval method and system based on adaptive intermediate granularity aggregation network
Through the adaptive intermediate particle size aggregation network and the target-guided semantic alignment module, the problems of intermediate particle size feature extraction and cross-modal semantic correspondence construction in the prior art are solved, and more efficient combined image retrieval accuracy is achieved.
Patent Information
- Application Number
- CN202510274983.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively extract and aggregate intermediate particle size features, and it is difficult to construct a cross-modal semantic correspondence between reference images and modified text, resulting in insufficient retrieval accuracy of combined images.
A combined image retrieval method based on an adaptive intermediate particle size aggregation network is proposed. The intermediate particle size extraction module generates supervision signals and aggregates intermediate particle size features, combines the target guide semantic alignment module to establish the semantic correspondence between the image and text, and finally generates the final combined features through the multi-grain size combination module.
Improve the accuracy of combined image retrieval, enables more efficient understanding of users' personalized needs, and find more accurate target images in the database.
Smart Images

Figure CN120104822A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a hierarchical semantic understanding and similarity retrieval method and system for e-commerce platform commodity image data and multimodal retrieval requirements, in particular to a combined image retrieval method and system based on an adaptive intermediate granularity aggregation network, belonging to the technical field of multimodal information retrieval. Background Art
[0002] With the rapid development of e-commerce, the effective combination of product images and text descriptions has become particularly important. Users often want to make personalized modifications based on existing reference images, such as adjusting colors, adding or removing elements, or combining different products. However, traditional image retrieval cannot meet these personalized needs of users. Therefore, combined image retrieval has emerged as an emerging task, aiming to integrate multimodal information (images and text) to meet users' complex personalized needs in a flexible way, which has attracted extensive research attention.
[0003] Specifically, the combined image retrieval task retrieves the target image required by the user from the database based on the reference image and modified text input by the user. The key to this task is how to accurately align the modification requirements in the modified text with the corresponding area of the reference image. Although some researchers have tried to effectively retrieve target images based on the perspective of local and global feature fusion, they have failed to fully consider the intermediate granularity features that exist between local and global. At the same time, some researchers have explored the feasibility of multi-granularity feature fusion, but they have ignored the semantic correspondence of mining intermediate granularity multimodal features. At present, there are two main challenges in realizing intermediate granularity feature extraction and mining the semantic correspondence of intermediate granularity multimodal features to build an effective combined image retrieval method:
[0004] (1) Intermediate-granularity aggregation lacks supervisory signals. Since intermediate-granularity features cannot be directly extracted and need to be obtained by aggregation, it is very challenging to construct a supervisory signal that can achieve intermediate-granularity aggregation.
[0005] (2) It is difficult to establish cross-modal semantic correspondence. Since the modified text in the combined image retrieval task expresses the modification requirements of the reference image, the intermediate granularity features of the modified text and the intermediate granularity features of the reference image cannot be completely consistent in semantics. Therefore, establishing the correspondence between the two is also challenging. Summary of the invention
[0006] In view of the shortcomings of the prior art, the present invention proposes a combined image retrieval method based on an adaptive intermediate granularity aggregation network to achieve effective retrieval of user target images.
[0007] The present invention also proposes a combined image retrieval system based on an adaptive intermediate granularity aggregation network.
[0008] To this end, the present invention firstly proposes an intermediate granularity extraction module, generates a supervision signal based on discrete Fourier transform, and constructs a graph attention network to adaptively aggregate intermediate granularity features; secondly, the present invention proposes a target-guided semantic alignment module, which can establish a semantic correspondence of intermediate granularity between the reference image and the modified text; finally, the present invention proposes a multi-granularity combination module, which can realize the "local-intermediate-global" multi-granularity feature combination and achieve effective retrieval of user target images.
[0009] Terminology explanation:
[0010] 1. CLIP is a deep learning model that aims to combine text and image information through contrastive learning methods to achieve cross-modal understanding. CLIP is pre-trained on a large-scale image-text pair dataset and can be used for a variety of tasks such as image retrieval, text generation, image classification, etc. Its cross-modal feature learning ability enables it to perform well in processing visual and language information.
[0011] 2. Attention mechanism is a computational method used in deep learning models that aims to dynamically focus on specific parts of the input data to improve the efficiency and effectiveness of information processing.
[0012] 3. Multilayer Perceptron (MLP) is a feedforward neural network consisting of at least three layers of nodes: input layer, hidden layer and output layer. Each layer consists of multiple neurons, and the neurons are connected by weights. It is widely used in tasks such as classification, regression, and feature extraction, and has shown superior performance in many fields such as pattern recognition, image processing, and natural language processing.
[0013] 4. Softmax function is an activation function widely used in the output layer of multi-classification problems. Its main function is to convert the input real-valued vector into a probability distribution so that each output value ranges between 0 and 1 and the sum of all output values is 1.
[0014] 5. Average Pooling is a downsampling technique widely used in convolutional neural networks. It aims to reduce the spatial dimension of the data and extract important features by performing regional averaging calculations on the input feature map.
[0015] 6. Batch normalization (BN) is a commonly used regularization technique in deep learning. BN normalizes the input of each layer of the network (that is, normalizes the data to a distribution with a mean of 0 and a standard deviation of 1) and introduces two learnable parameters (scaling parameter γ and offset parameter β), so that the network can still learn more flexible feature representations while maintaining the normalization effect.
[0016] 7. ReLU is one of the commonly used nonlinear activation functions in deep learning. Its Chinese name is rectified linear unit. ReLU activation function is widely used in neural networks to introduce nonlinearity so that the network can fit complex nonlinear mapping relationships.
[0017] 8. KL divergence (Kullback-Leibler Divergence) is an asymmetric measure of the difference between two probability distributions. Specifically, KL divergence is used to evaluate the information loss or relative entropy of one probability distribution relative to another probability distribution.
[0018] 9. Multi-head attention mechanism (MHA) is one of the core mechanisms of the Transformer architecture in deep learning, and is mainly used to process sequence data (such as text, speech or time series). It extracts feature representations of different subspaces from the input data through a parallel attention mechanism, significantly improving the model's ability to capture global context. MHA expands the attention mechanism into multiple "heads", each of which independently calculates the attention score and extracts specific features, and then concatenates the outputs of all heads and performs a linear transformation to generate the final feature representation.
[0019] 10. Discrete Fourier Transform (DFT) is a signal processing tool used to convert discrete signals from the time domain or space domain to the frequency domain. It can analyze the frequency components of the signal and is widely used in digital signal processing, image processing, communications, scientific computing and other fields.
[0020] 11. Graph Attention Network is a graph neural network based on the attention mechanism, which is mainly used to process graph structure data. It effectively models the importance between nodes by introducing an adaptive attention weight mechanism, thereby improving the expressiveness of traditional graph neural networks in the process of information aggregation.
[0021] The technical solution of the present invention is as follows:
[0022] A combined image retrieval method based on adaptive intermediate granularity aggregation network (MEDIAN), including:
[0023] The training set data is read in batches, and global and local features are extracted from the training set data. The local features are transformed into supervisory signals through discrete Fourier transform, and the intermediate granularity features are adaptively aggregated based on the generated supervisory signals and the graph attention network.
[0024] For the intermediate granularity features of adaptively aggregated images and texts, we deeply mine and establish their semantic correspondences;
[0025] For the features of different granularities, the final combined features are obtained through multi-granularity combination;
[0026] Calculate the dot product of the combined feature and each image in the gallery to get the similarity score, and sort these similarity scores in descending order;
[0027] The top K candidate target images with similarity scores are selected to form a formal retrieval result set, completing the combined image retrieval.
[0028] As a further preferred solution, the training set data is read in batches, and global features and local features of the training set data are extracted; including:
[0029] The training set data is read in batches; each set of data in each batch is in the form of a triplet of <reference image, modified text, target image>; the reference image refers to the original image input by the user; the modified text refers to the text about the user's modification requirements; the target image refers to the image expected to be obtained after modifying the reference image according to the modified text; through CLIP, the reference image x is extracted respectively r The global feature P G ∈R D and local features P L ∈R N×D , expressed as:
[0030]
[0031] in, and denote the last and penultimate layers of the CLIP image encoder, respectively. D is the CLIP global embedding dimension. FC I Align the local embedding dimension to D; R refers to the real number domain, indicating that all elements in the feature are real numbers;
[0032] Get the global feature T of the target image G ∈R D and local features T L ∈R N×D , and modify the global feature Q of the text G ∈R D and local feature Q L ∈RM×D , where N and M represent the number of channels of image and text respectively;
[0033] A multi-layer perceptron is used to unify the number of local feature channels of the reference image, modified text, and target image, as shown below:
[0034]
[0035] in, They represent the local features of the reference image, modified text, and target image after being mapped to a unified number of channels, respectively. Z is the number of feature channels, and MLP(·) is a multi-layer perceptron. The unified local features are briefly referred to as local features.
[0036] As a further preferred solution, the local features are transformed into a supervisory signal by discrete Fourier transform; including:
[0037] Generate supervision signals based on local features through discrete Fourier transform;
[0038] Based on local features of reference image The corresponding supervision signal is obtained by discrete Fourier transform Among them, P F The feature of the nth channel in Obtained through discrete Fourier transform, it is expressed as follows:
[0039]
[0040] in, Represents the local features of the reference image The feature of the mth channel in , i is an imaginary number;
[0041] Based on local features of modified text Calculate the supervisory signal for modifying the text Among them, Q F The feature of the nth channel in Through discrete Fourier transform, the formula is as follows:
[0042]
[0043] in, Indicates modification of local features of text The feature of the mth channel in , i is an imaginary number;
[0044] Based on the local features of the target image Get the corresponding supervision signal Among them, T F The feature of the nth channel in Obtained through discrete Fourier transform, expressed as:
[0045]
[0046] in, Represents the local features of the target image The feature of the mth channel in , i is an imaginary number.
[0047] As a further preferred solution, according to the generated supervision signal, based on the graph attention network, the intermediate granularity features are adaptively aggregated; including:
[0048] For the reference image, first, the generated supervisory signal P F and local features Splice them together to get the local features guided by the signal K=2Z; then, P U The feature of each channel is regarded as a node in the graph attention network, then the i-th and j-th channel features The edge between The calculation is as follows:
[0049]
[0050] Among them, E p ∈R 1×D , ⊙ represents the multiplication of elements at corresponding positions of the matrix;
[0051] Then, define the aggregation graph in, Represents P U Each channel feature consists of a node set, ε P Represents P U The set of edges between each channel feature; in order to adaptively aggregate the intermediate granularity features, a multi-head attention mechanism is implemented to calculate the attention coefficient of each node, the formula is as follows:
[0052]
[0053]
[0054] Where H represents the number of attention heads, d = D / H, E op ∈R D×D , is a learnable weight matrix, Indicates that feature nodes are calculated through a multi-head attention mechanism and The relationship between head h Refers to the output of the hth head in the multi-head attention mechanism;
[0055] Finally, the intermediate granularity features obtained by adaptive aggregation are expressed as For the characteristics of any channel All are based on the attention coefficient. Corresponding node and located in The nodes of the relevant neighborhood are aggregated and the formula is as follows:
[0056]
[0057] Among them, BN is batch normalization, ReLU is the activation function, N k Represents the kth node The neighbor set.
[0058] Similarly, the intermediate granularity features of the modified text and the target image are obtained, which are represented as Q I ∈R K×D and T I ∈R K×D .
[0059] As a further preferred solution, for the intermediate granularity features of the adaptively aggregated image and text, in-depth mining and establishment of their semantic correspondences are performed; including:
[0060] Assuming that each batch has B training triplets, we first calculate the similarity distribution between the intermediate granularity features of the i-th reference image and the intermediate granularity features of all target images in the same batch, expressed as in, represents the similarity between the intermediate granularity features of the i-th reference image and the intermediate granularity features of the b-th target image, which is formulated as follows:
[0061]
[0062] Among them, cos(·) represents cosine similarity, represents the intermediate granularity features of the reference image after the i-th average pooling, represents the intermediate granularity feature of the target image after the jth average pooling, and τ is the temperature coefficient;
[0063] Similarly, the similarity distribution between the intermediate granularity features of the i-th modified text and the intermediate granularity features of all target images in the same batch is obtained, which is expressed as
[0064] Finally, calculate the KL divergence loss function The similarity distribution is made consistent to achieve cross-modal semantic alignment, which is formulated as follows:
[0065]
[0066] in, is the KL divergence loss function, D KL (·) is the KL divergence, represents the similarity between the intermediate granularity features of the i-th reference image and the intermediate granularity features of the j-th target image, It represents the similarity between the intermediate granularity features of the i-th modified text and the intermediate granularity features of the j-th target image.
[0067] Finally, a preserved semantic correspondence is gradually established between the reference image and the target image, i.e., the part not involved in modification in the user's demand; a modified semantic correspondence is gradually established between the modified text and the target image, i.e., the part involved in modification in the user's demand.
[0068] As a further preferred solution, for the obtained features of different particle sizes, a final combined feature is obtained by combining multiple particle sizes; including:
[0069] Integrate multi-granular features of reference image and modified text respectively;
[0070] Based on the obtained multi-granularity features, multi-granularity interaction is performed through a multi-layer perceptron to learn the respective combined weights of the reference image and the modified text;
[0071] Based on the learned combination weights, the final combination features are obtained.
[0072] As a further preferred solution, for the obtained features of different particle sizes, a final combined feature is obtained by combining multiple particle sizes; including:
[0073] The corresponding global features, intermediate granularity features, and local features are concatenated; for the reference image, the multi-granularity features are expressed as S = 1 + K + Z; Similarly, the multi-granularity features of the modified text are obtained And the multi-granular features of the target image
[0074] Subsequently, the nonlinear capability of the multi-layer perceptron is used to learn the combined weights w between the multi-granular features of the reference image and the modified text, which can be expressed as follows:
[0075] w = MLP([P,Q]);
[0076] Where w∈R S×2D , MLP(·) represents multi-layer perceptron;
[0077] The reference image and the modified text are assigned combined weights through block operations, and each weight is weightedly aggregated to obtain the final multimodal synthetic feature C, which is calculated as follows:
[0078] C=w p ⊙P+w q⊙Q;
[0079] Among them, w p ,w q ∈R S×D is the output of the block operation, ⊙ represents the corresponding multiplication of the elements in the matrix;
[0080] The batch-based classification loss is used to promote the multi-granularity combination features to approach the corresponding target features; the batch-based classification loss is formulated as follows:
[0081]
[0082] Among them, cos(·) represents cosine similarity, C i represents the multi-granularity combined features corresponding to the i-th triple, T i ,T j They represent the target image features corresponding to the i-th triplet and the j-th triplet, τ is the temperature coefficient, and B is the number of triplets in a batch;
[0083] The final objective function is as follows:
[0084]
[0085] Among them, Θ * represents the parameters to be optimized for the combined image retrieval method, and μ is a trade-off hyperparameter.
[0086] The final combined feature is feature C obtained by assigning combined weights to the reference image and the modified text through block operation and performing weighted aggregation on each weight.
[0087] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a combined image retrieval method based on an adaptive intermediate granularity aggregation network are implemented.
[0088] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a combined image retrieval method based on an adaptive intermediate granularity aggregation network.
[0089] A combined image retrieval system based on an adaptive intermediate granularity aggregation network, including:
[0090] The intermediate granularity feature extraction module is configured to: read the training set data in batches, and extract global features and local features from the training set data; generate supervision signals from the local features through discrete Fourier transform, and adaptively aggregate the intermediate granularity features based on the generated supervision signals and the graph attention network;
[0091] The target-guided semantic alignment module is configured to: deeply mine and establish semantic correspondences between the adaptively aggregated intermediate granularity features of images and texts;
[0092] The multi-granularity combination module is configured to: obtain the final combined feature through multi-granularity combination for the obtained features of different granularities;
[0093] The image retrieval module is configured to: use the combined features to screen the images in the candidate target image set to obtain a formal retrieval result set, that is, the retrieval final result.
[0094] Compared with the prior art, the present invention has the following beneficial effects:
[0095] 1. The present invention proposes an intermediate granularity extraction module, generates a supervision signal based on discrete Fourier transform, and constructs a graph attention network to adaptively realize the aggregation of intermediate granularity features, thereby promoting further improvement of the combined image retrieval accuracy.
[0096] 2. The present invention proposes a target-guided semantic alignment module for effectively constructing the semantic correspondence between the reference image and the intermediate granularity of the modified text, thereby promoting further improvement in the accuracy of combined image retrieval.
[0097] 3. The present invention proposes a multi-granularity combination module, which generates the final combined features by effectively fusing the multi-granularity features of the reference image and the modified text. With the assistance of the intermediate granularity features and combined with cross-modal semantic alignment, the module combines more complete and accurate combined features, thereby further improving the accuracy of combined image retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] Figure 1 It is a schematic flow chart of a combined image retrieval method based on an adaptive intermediate granularity aggregation network of the present invention; DETAILED DESCRIPTION
[0099] The present invention will be further defined below in conjunction with the accompanying drawings and embodiments, but is not limited thereto.
[0100] Example 1
[0101] Combined image retrieval method based on adaptive intermediate granularity aggregation network (MEDIAN), such as Figure 1 As shown, including:
[0102] The training set data is read in batches, and global and local features are extracted from the training set data. The local features are transformed into supervisory signals through discrete Fourier transform, and the intermediate granularity features are adaptively aggregated based on the generated supervisory signals and the graph attention network.
[0103] For the intermediate granularity features of adaptively aggregated images and texts, we deeply mine and establish their semantic correspondences;
[0104] For the features of different granularities, the final combined features are obtained through multi-granularity combination;
[0105] The present invention combines multimodal queries for the input reference image and modified text through intermediate granularity extraction, cross-modal semantic alignment, and multi-granularity feature combination, thereby generating combined features. The dot product of the combined features and each image in the gallery is calculated to obtain similarity scores, and these similarity scores are sorted in descending order;
[0106] According to actual needs (for example, retrieving the K images that best meet the query conditions), select the top K candidate target images with the highest similarity scores to form a formal retrieval result set, thus completing the combined image retrieval.
[0107] Example 2
[0108] The combined image retrieval method based on adaptive intermediate granularity aggregation network (MEDIAN) described in Example 1 is different in that:
[0109] Read the training set data in batches and extract global and local features from the training set data; including:
[0110] The training set data is read in batches; each set of data in each batch is in the form of a triplet of <reference image, modified text, target image>; the reference image refers to the original image input by the user, which contains the basic information to be processed; the modified text refers to the text about the user's modification requirements; the target image refers to the image expected to be obtained after modifying the reference image according to the modified text; through CLIP, the reference image x r The global feature P G ∈R D and local features P L ∈R N×D , expressed as:
[0111]
[0112]
[0113] in, and denote the last and penultimate layers of the CLIP image encoder, respectively. D is the CLIP global embedding dimension. FC I Align the local embedding dimension to D; R refers to the real number domain, indicating that all elements in the feature are real numbers;
[0114] Similarly, the global feature T of the target image is obtained G ∈RD and local features T L ∈R N×D , and modify the global feature Q of the text G ∈R D and local feature Q L ∈R M×D , where N and M represent the number of channels of image and text respectively;
[0115] In order to make the number of channels consistent, a multi-layer perceptron is used to unify the number of local feature channels of the reference image, modified text, and target image, as shown below:
[0116]
[0117] in, They represent the local features of the reference image, modified text, and target image after being mapped to a unified number of channels, respectively. Z is the number of feature channels, and MLP(·) is a multi-layer perceptron. For ease of description, the unified local features are referred to as local features.
[0118] Generate supervision signals from local features through discrete Fourier transform; including:
[0119] Generate supervision signals based on local features through discrete Fourier transform;
[0120] Based on local features of reference image The corresponding supervision signal is obtained by discrete Fourier transform Among them, P F The feature of the nth channel in Obtained through discrete Fourier transform, it is expressed as follows:
[0121]
[0122] in, Represents the local features of the reference image The feature of the mth channel in , i is an imaginary number;
[0123] Similarly, based on the local features of the modified text Calculate the supervisory signal for modifying the text Among them, Q F The feature of the nth channel in Through discrete Fourier transform, the formula is as follows:
[0124]
[0125] in, Indicates modification of local features of text The feature of the mth channel in , i is an imaginary number;
[0126] Similarly, based on the local features of the target image Get the corresponding supervision signal Among them, T F The feature of the nth channel in Obtained through discrete Fourier transform, expressed as:
[0127]
[0128] in, Represents the local features of the target image The feature of the mth channel in , i is an imaginary number.
[0129] According to the generated supervision signal, based on the graph attention network, the intermediate granularity features are adaptively aggregated; including:
[0130] To adaptively aggregate intermediate granularity features, a graph attention network is adopted and the previously generated supervisory signals are utilized as guidance.
[0131] For the reference image, first, the generated supervisory signal P F and local features Splice them together to get the local features guided by the signal K=2Z; then, P U The feature of each channel is regarded as a node in the graph attention network, then the i-th and j-th channel features The edge between The calculation is as follows:
[0132]
[0133] Among them, E P ∈R 1×D , ⊙ represents the multiplication of elements at corresponding positions of the matrix;
[0134] Then, define the aggregation graph in, Represents P U Each channel feature consists of a node set, ε P Represents P U The set of edges between each channel feature; in order to adaptively aggregate the intermediate granularity features, a multi-head attention mechanism is implemented to calculate the attention coefficient of each node, the formula is as follows:
[0135]
[0136]
[0137] Where H represents the number of attention heads, d = D / H, E op ∈RD×D , is a learnable weight matrix, Indicates that feature nodes are calculated through a multi-head attention mechanism and The relationship between head h Refers to the output of the hth head in the multi-head attention mechanism;
[0138] Finally, the intermediate granularity features obtained by adaptive aggregation are expressed as For the characteristics of any channel All are based on the attention coefficient. Corresponding node and located in The nodes of the relevant neighborhood (i.e., neighbor nodes) are aggregated and the formula is as follows:
[0139]
[0140] Among them, BN is batch normalization, ReLU is a commonly used activation function, and N k Represents the kth node The neighbor set of (i.e., other channel features).
[0141] Similarly, the intermediate granularity features of the modified text and the target image are obtained, which are represented as Q I ∈R K×D and T I ∈R K×D .
[0142] Get the intermediate granularity features of the modified text, denoted as Q I ∈R K×D ;include:
[0143] First, the generated supervisory signal Q F and local features Splice them together to get the local features guided by the signal Subsequently, the graph structure construction process is the same as that of the reference image, each channel feature is regarded as a node of the text feature, and the edges between the nodes are calculated; the i-th and j-th channel features The edge between The calculation is as follows:
[0144]
[0145] Among them, E Q ∈R 1×D , ⊙ represents the multiplication of elements at corresponding positions of the matrix;
[0146] Then, define the aggregation graph G U (QU ,ε Q ), where Q U Representative Q U Each channel feature consists of a node set, ε Q Representative Q U The set of edges between each channel feature; continue to use the multi-head attention mechanism to calculate the attention coefficient of each node, expressed as follows:
[0147]
[0148]
[0149] Where H represents the number of attention heads, d = D / H, E oq ∈R D×D , is the learnable weight matrix, Indicates that feature nodes are calculated through a multi-head attention mechanism and The relationship between head h Refers to the output of the h-th head in the multi-head attention mechanism.
[0150] Finally, the intermediate granularity features obtained by adaptive aggregation are expressed as For the characteristics of any channel All are based on the attention coefficient. Corresponding node and located in The nodes of the relevant neighborhood (i.e., neighbor nodes) are aggregated and formulated as follows:
[0151]
[0152] Among them, BN is batch normalization, ReLU is a commonly used activation function, and N k Represents the kth node The neighbor set of (i.e., other channel features).
[0153] Get the intermediate granularity features of the target image, represented by T I ∈R K×D ;include:
[0154] First, the supervisory signal T generated by the target image F and local features of the target image Splicing together to get signal-guided local features Then, T U Each channel feature is regarded as a node, and the graph G is constructed U (T U ,ε T ), where TU Represents T U Each channel feature consists of a node set, ε T Represents T U The set of edges between each channel feature; ε T The edge Indicates T U The i-th and j-th channel features in The edge between them is calculated as follows:
[0155]
[0156] Among them, E T ∈R 1×D , ⊙ represents the multiplication of elements at corresponding positions of the matrix;
[0157] Then, the attention coefficient of each node is calculated through the multi-head attention mechanism, which is expressed as follows:
[0158]
[0159]
[0160] Where H represents the number of attention heads, d = D / H, W ot ∈R D×D , is the learnable weight matrix, Indicates that feature nodes are calculated through a multi-head attention mechanism and The relationship between head h Refers to the output of the hth head in the multi-head attention mechanism;
[0161] Finally, the intermediate granularity features obtained by adaptive aggregation are expressed as For the characteristics of any channel All are based on the attention coefficient. Corresponding node and located in The nodes of the relevant neighborhood (i.e., neighbor nodes) are aggregated and expressed as follows:
[0162]
[0163] Among them, BN is batch normalization, ReLU is a commonly used activation function, and N k Represents the kth node The neighbor set of (i.e. other channel features).
[0164] For the intermediate granularity features of the adaptively aggregated images and texts, we deeply explore and establish their semantic correspondences, including:
[0165] First, the reference image and modified text include preserved semantics and modified semantics respectively, and it is difficult to establish a direct correspondence based on similarity. Given that the target image in the triple contains both preserved semantics and modified semantics, it can be used as a link between the reference image and the modified text. Specifically, assuming that each batch has B training triplets, we first calculate the similarity distribution between the intermediate granularity features of the i-th reference image and the intermediate granularity features of all target images in the same batch, expressed as in, represents the similarity between the intermediate granularity features of the i-th reference image and the intermediate granularity features of the b-th target image, which is formulated as follows:
[0166]
[0167] Among them, cos(·) represents cosine similarity, represents the intermediate granularity features of the reference image after the i-th average pooling, represents the intermediate granularity feature of the target image after the jth average pooling, and τ is the temperature coefficient;
[0168] Similarly, the similarity distribution between the intermediate granularity features of the i-th modified text and the intermediate granularity features of all target images in the same batch is obtained, which is expressed as
[0169] Finally, calculate the KL divergence loss function The similarity distribution is made consistent to achieve cross-modal semantic alignment, which is formulated as follows:
[0170]
[0171] in, is the KL divergence loss function, D KL (·) is the KL divergence, represents the similarity between the intermediate granularity features of the i-th reference image and the intermediate granularity features of the j-th target image, It represents the similarity between the intermediate granularity features of the i-th modified text and the intermediate granularity features of the j-th target image.
[0172] Finally, a preserved semantic correspondence is gradually established between the reference image and the target image, i.e., the part not involved in modification in the user's demand; a modified semantic correspondence is gradually established between the modified text and the target image, i.e., the part involved in modification in the user's demand.
[0173] For the features of different granularities, the final combined features are obtained through multi-granularity combination, including:
[0174] Integrate multi-granular features of reference image and modified text respectively;
[0175] Based on the obtained multi-granularity features, multi-granularity interaction is performed through a multi-layer perceptron to learn the combined weights of the reference image and the modified text;
[0176] Based on the learned combination weights, the final combination features are obtained.
[0177] For the features of different granularities, the final combined features are obtained through multi-granularity combination, including:
[0178] The corresponding global features, intermediate granularity features, and local features are concatenated; for the reference image, the multi-granularity features are expressed as S = 1 + K + Z; Similarly, the multi-granularity features of the modified text are obtained And the multi-granular features of the target image
[0179] Subsequently, the nonlinear capability of the multi-layer perceptron is used to learn the combined weights w between the multi-granular features of the reference image and the modified text, which can be expressed as follows:
[0180] w = MLP([P,Q]);
[0181] Where w∈R S×2D , MLP(·) represents multi-layer perceptron;
[0182] The reference image and the modified text are assigned combined weights through block operations, and each weight is weightedly aggregated to obtain the final multimodal synthetic feature C, which is calculated as follows:
[0183] C=w p ⊙P+w q ⊙Q;
[0184] Among them, w p ,w q ∈R S×D is the output of the block operation, ⊙ represents the corresponding multiplication of the elements in the matrix;
[0185] The batch-based classification loss is used to promote the multi-granularity combination features to approach the corresponding target features; the batch-based classification loss is formulated as follows:
[0186]
[0187] Among them, cos(·) represents cosine similarity, C i represents the multi-granularity combined features corresponding to the i-th triple, T i ,T j They represent the target image features corresponding to the i-th triplet and the j-th triplet, τ is the temperature coefficient, and B is the number of triplets in a batch;
[0188] The final objective function is as follows:
[0189]
[0190] Among them, Θ * represents the parameters to be optimized for the combined image retrieval method, and μ is a trade-off hyperparameter.
[0191] The final combined feature is feature C obtained by assigning combined weights to the reference image and the modified text through block operation and performing weighted aggregation on each weight.
[0192] Table 1 is a schematic diagram of the retrieval accuracy comparison of the present invention on the FashionIQ dataset; here, Dresses, Shirts, Tops&Tees represent different subsets of the FashionIQ dataset; Avg represents the average performance when searching on these three subsets; R@10 and R@50 refer to the recall rates calculated by returning the first 10 and first 50 results in the search, that is, the ratio of the number of retrieved relevant results to the number of all relevant results in the library, which aims to evaluate the recall rate of the retrieval system.
[0193] Table 1
[0194]
[0195] Table 2 is a schematic diagram of the comparison of the retrieval accuracy of the present invention on the Shoes dataset; here, R@1, R@10 and R@50 refer to the recall rate calculated by returning the first 1, 10 and 50 results in the retrieval, that is, the ratio of the number of relevant results retrieved to the number of all relevant results in the library, which is intended to evaluate the recall rate of the retrieval system. Avg is the average recall rate, which is the arithmetic mean of R@1, R@10 and R@50.
[0196] Table 2
[0197]
[0198] Table 3 is a schematic diagram of the comparison of the retrieval accuracy of the present invention on the CIRR dataset; here, R@k represents the recall rate calculated by returning the first k results in the retrieval, and Rsubset@k represents the recall rate calculated by returning the first k results when the system performs retrieval on the subset divided in the CIRR dataset. The recall rate is the ratio of the number of relevant results retrieved to the number of all relevant results in the database, which is intended to evaluate the recall rate of the retrieval system.
[0199] Table 3
[0200]
[0201] As shown in Tables 1 to 3, the query efficiency and retrieval accuracy are compared with the international leading similar methods and the results are displayed. The results show that compared with similar combined image retrieval methods, the model MEDIAN of the present invention has higher accuracy on three widely used benchmark datasets.
[0202] Example 3
[0203] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the combined image retrieval method based on an adaptive intermediate granularity aggregation network described in embodiment 1 or 2 are implemented.
[0204] Example 4
[0205] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the combined image retrieval method based on an adaptive intermediate granularity aggregation network described in Example 1 or 2.
[0206] Example 5
[0207] A combined image retrieval system based on an adaptive intermediate granularity aggregation network, including:
[0208] The intermediate granularity feature extraction module is configured to: read the training set data in batches, and extract global features and local features from the training set data; generate supervision signals from the local features through discrete Fourier transform, and adaptively aggregate the intermediate granularity features based on the generated supervision signals and the graph attention network;
[0209] The target-guided semantic alignment module is configured to: deeply mine and establish semantic correspondences between the adaptively aggregated intermediate granularity features of images and texts;
[0210] The multi-granularity combination module is configured to: obtain the final combined feature through multi-granularity combination for the obtained features of different granularities;
[0211] The image retrieval module is configured to: use the combined features to screen the images in the candidate target image set to obtain a formal retrieval result set, that is, the retrieval final result.
[0212] The specific architecture of the adaptive intermediate granularity aggregation network of the present invention includes an intermediate granularity extraction module, a target-guided semantic alignment module and a multi-granularity combination module. The intermediate granularity extraction module includes two main steps: generation of intermediate granularity supervision signals and aggregation of adaptive features.
Claims
1. A combined image retrieval method based on an adaptive intermediate granularity aggregation network, characterized in that: include: The training set data is read in batches, and global and local features are extracted from the training set data. The local features are transformed into supervisory signals through discrete Fourier transform, and the intermediate granularity features are adaptively aggregated based on the generated supervisory signals and the graph attention network. For the intermediate granularity features of adaptively aggregated images and texts, we deeply mine and establish their semantic correspondences; For the features of different granularities, the final combined features are obtained through multi-granularity combination; Calculate the dot product of the combined feature and each image in the gallery to get the similarity score, and sort these similarity scores in descending order; The top K candidate target images with similarity scores are selected to form a formal retrieval result set, completing the combined image retrieval.
2. The combined image retrieval method based on the adaptive intermediate granularity aggregation network according to claim 1 is characterized in that: Read the training set data in batches and extract global and local features from the training set data; including: The training set data is read in batches; each set of data in each batch is in the form of a triplet of <reference image, modified text, target image>; the reference image refers to the original image input by the user; the modified text refers to the text about the user's modification requirements; the target image refers to the image expected to be obtained after modifying the reference image according to the modified text; through CLIP, the reference image x is extracted respectively r The global feature P G ∈R D and local features P L ∈R N×D , expressed as: in, and denote the last and penultimate layers of the CLIP image encoder, respectively. D is the CLIP global embedding dimension. FC I Align the local embedding dimension to D; R refers to the real number domain, indicating that all elements in the feature are real numbers; Get the global feature T of the target image G ∈R D and local features T L ∈R N×D , and modify the global feature Q of the text G ∈R D and local feature Q L ∈R M×D , where N and M represent the number of channels of image and text respectively; A multi-layer perceptron is used to unify the number of local feature channels of the reference image, modified text, and target image, as shown below: in, They represent the local features of the reference image, modified text, and target image after being mapped to a unified number of channels, respectively. Z is the number of feature channels, and MLP(·) is a multi-layer perceptron. The unified local features are briefly referred to as local features.
3. The combined image retrieval method based on the adaptive intermediate granularity aggregation network according to claim 1 is characterized in that: Generate supervision signals from local features through discrete Fourier transform; include: Generate supervision signals based on local features through discrete Fourier transform; Based on local features of reference image The corresponding supervision signal is obtained by discrete Fourier transform Among them, P F The feature of the nth channel in Obtained through discrete Fourier transform, expressed as follows: in, Represents the local features of the reference image The feature of the mth channel in , i is an imaginary number; Based on local features of modified text Calculate the supervisory signal for modifying the text Among them, Q F The feature of the nth channel in Through discrete Fourier transform, the formula is as follows: in, Indicates modification of local features of text The feature of the mth channel in , i is an imaginary number; Based on local features of target image Get the corresponding supervision signal Among them, T F The feature of the nth channel in Obtained through discrete Fourier transform, expressed as: in, Represents the local features of the target image The feature of the mth channel in , i is an imaginary number.
4. The combined image retrieval method based on the adaptive intermediate granularity aggregation network according to claim 1 is characterized in that: According to the generated supervision signal, based on the graph attention network, the intermediate granularity features are adaptively aggregated; including: For the reference image, first, the generated supervisory signal P F and local features Splice them together to get the local features guided by the signal Subsequently, the feature of each channel of P1 is regarded as a node in the graph attention network, and the i-th and j-th channel features The edge between The calculation is as follows: Among them, E < ∈R" ×D , ⊙ represents the multiplication of elements at corresponding positions of the matrix; Then, define the aggregation graph in, represents the node set composed of each channel feature of P1, ε < Represents the set of edges between each channel feature of P1; in order to adaptively aggregate the intermediate granularity features, a multi-head attention mechanism is implemented to calculate the attention coefficient of each node, the formula is as follows: Where H represents the number of attention heads, d = D / H, W FG ∈R D×D , is a learnable weight matrix, Indicates that feature nodes are calculated through a multi-head attention mechanism and The relationship between head h Refers to the output of the hth head in the multi-head attention mechanism; Finally, the intermediate granularity features obtained by adaptive aggregation are expressed as For the characteristics of any channel All are based on the attention coefficient. Corresponding node and located in The nodes of the relevant neighborhood are aggregated and the formula is as follows: Among them, BN is batch normalization, ReLU is the activation function, N ] Represents the kth node The neighbor set of Similarly, the intermediate granularity features of the modified text and the target image are obtained, which are represented as Q I ∈R 4×D and T I ∈R 4×D .
5. The combined image retrieval method based on the adaptive intermediate granularity aggregation network according to claim 1 is characterized in that: For the intermediate granularity features of the adaptively aggregated images and texts, we deeply explore and establish their semantic correspondences, including: Assuming that each batch has B training triplets, we first calculate the similarity distribution between the intermediate granularity features of the i-th reference image and the intermediate granularity features of all target images in the same batch, expressed as in, represents the similarity between the intermediate granularity features of the i-th reference image and the intermediate granularity features of the b-th target image, which is formulated as follows: Among them, cos(·) represents cosine similarity, represents the intermediate granularity features of the reference image after the i-th average pooling, represents the intermediate granularity feature of the target image after the jth average pooling, and τ is the temperature coefficient; Similarly, the similarity distribution between the intermediate granularity features of the i-th modified text and the intermediate granularity features of all target images in the same batch is obtained, which is expressed as Finally, calculate the KL divergence loss function The similarity distribution is made consistent to achieve cross-modal semantic alignment, which is formulated as follows: in, is the KL divergence loss function, D 4L (·) is the KL divergence, represents the similarity between the intermediate granularity features of the i-th reference image and the intermediate granularity features of the j-th target image, Represents the similarity between the intermediate granularity features of the i-th modified text and the intermediate granularity features of the j-th target image; Finally, a preserved semantic correspondence is gradually established between the reference image and the target image, i.e., the part not involved in modification in the user's demand; a modified semantic correspondence is gradually established between the modified text and the target image, i.e., the part involved in modification in the user's demand.
6. The combined image retrieval method based on the adaptive intermediate granularity aggregation network according to claim 1 is characterized in that: For the features of different granularities, the final combined features are obtained through multi-granularity combination, including: Integrate multi-granular features of reference image and modified text respectively; Based on the obtained multi-granularity features, multi-granularity interaction is performed through a multi-layer perceptron to learn the respective combined weights of the reference image and the modified text; Based on the learned combination weights, the final combination features are obtained.
7. The combined image retrieval method based on an adaptive intermediate granularity aggregation network according to any one of claims 1 to 6, characterized in that: For the features of different granularities, the final combined features are obtained through multi-granularity combination, including: The corresponding global features, intermediate granularity features, and local features are concatenated; for the reference image, the multi-granularity features are expressed as S = 1 + K + Z; Similarly, the multi-granularity features of the modified text are obtained And the multi-granular features of the target image Subsequently, the nonlinear capability of the multi-layer perceptron is used to learn the combined weights w between the multi-granular features of the reference image and the modified text, which can be expressed as follows: w = MLP([P,Q]); Where w∈R S×2D , MLP(·) represents multi-layer perceptron; The reference image and the modified text are assigned combined weights through block operations, and each weight is weightedly aggregated to obtain the final multimodal synthetic feature C, which is calculated as follows: C=w p ⊙P+w q ⊙Q; Among them, w p ,w q ∈R S×D is the output of the block operation, ⊙ represents the corresponding multiplication of the elements in the matrix; The batch-based classification loss is used to promote the multi-granularity combination features to approach the corresponding target features; the batch-based classification loss is formulated as follows: Among them, cos(·) represents cosine similarity, C i represents the multi-granularity combined features corresponding to the i-th triple, T i ,T j They represent the target image features corresponding to the i-th triplet and the j-th triplet, τ is the temperature coefficient, and B is the number of triplets in a batch; The final objective function is as follows: Among them, Θ * represents the parameters to be optimized for the combined image retrieval method, and μ is a trade-off hyperparameter; The final combined feature is feature C obtained by assigning combined weights to the reference image and the modified text through block operation and performing weighted aggregation on each weight.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the combined image retrieval method based on an adaptive intermediate granularity aggregation network described in any one of claims 1-7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the combined image retrieval method based on an adaptive intermediate granularity aggregation network described in any one of claims 1 to 7 are implemented.
10. A combined image retrieval system based on an adaptive intermediate granularity aggregation network, characterized in that: include: The intermediate granularity feature extraction module is configured to: read the training set data in batches, and extract global features and local features from the training set data; generate supervision signals from the local features through discrete Fourier transform, and adaptively aggregate the intermediate granularity features based on the generated supervision signals and the graph attention network; The target-guided semantic alignment module is configured to: deeply mine and establish semantic correspondences between the adaptively aggregated intermediate granularity features of images and texts; The multi-granularity combination module is configured to: obtain the final combined features through multi-granularity combination for the obtained features of different granularities; The image retrieval module is configured to: use the combined features to screen the images in the candidate target image set to obtain a formal retrieval result set, that is, the retrieval final result.