Combined image retrieval method and system based on entity mining and modification relation binding

By using entity mining and modifying relationship binding methods in combined image retrieval, the problems of irrelevant factor perturbation, fuzzy semantic boundaries and implicitly modifying relationships in the prior art are solved, and a more efficient combined image retrieval effect is achieved.

CN120067365AActive Publication Date: 2025-05-30SHANDONG UNIV

Patent Information

Application Number
CN202411903224.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-30
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

In the retrieval of combined image, it is difficult to effectively identify visual and text semantics related to modification behavior in the prior art. There are problems of irrelevant factor perturbation, fuzzy semantic boundaries and implicit modification relationships, resulting in insufficient retrieval performance.

Method used

A combined image retrieval method based on entity mining and modification relationship binding is proposed. Through the latent factor filtering module, entity-action binding module and multi-scale combination module, the visual and text potential factors related to modifying semantics are identified, the semantic relationship is deeply explored and entities and actions are bound, and multi-modal combination features are generated to achieve effective retrieval.

Benefits of technology

By effectively filtering irrelevant factors, probing semantic boundaries and binding modification relationships, the accuracy and efficiency of combined image retrieval is improved to meet the complex personalized needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067365A_ABST
    Figure CN120067365A_ABST
Patent Text Reader

Abstract

The invention relates to a combined image retrieval method and system based on entity mining and modification relation binding, and the method comprises the steps: reading training set data in batches, and carrying out the extraction of global features and local features of the training set data; filtering the local features through potential factors, and further extracting visual and text potential factor features related to the modified semantics; for potential factor characteristics, combining entity-action binding, deeply mining a semantic relationship in a reference image and a modification text, detecting a semantic boundary, and respectively aggregating potential factors into a visual entity and a modification action; for the obtained features of different scales, performing multi-scale combination to obtain final combined features; and dot products of the combined features and different images in the image library are calculated respectively to serve as similarity scores, the similarity scores are arranged in a descending order, target images with the similarity scores ranked in the first several positions are selected, and combined image retrieval is completed. According to the method, the target image of the user is effectively retrieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for query understanding and similarity retrieval of e-commerce image data, and particularly to a combined image retrieval method and system based on entity mining and modified relationship binding, belonging to the technical field of multimodal information retrieval. Background Art

[0002] In the context of the rapid development of e-commerce, the effective combination of product images and text descriptions has become increasingly important. Users often hope to make modifications based on existing reference images, such as adjusting colors, adding or deleting certain elements, or combining different products. In this case, traditional image retrieval cannot meet the personalized needs of users. Therefore, the combined image retrieval task has emerged, aiming to flexibly meet the complex personalized needs of users by integrating multimodal information (images and text), and has received extensive attention from researchers.

[0003] Specifically, the combined image retrieval task retrieves the target image required by the user in the database according to the reference image and modification text input by the user. The key to this task is to accurately identify the modification requirements and locate the corresponding entities that need to be modified in the reference image. For the convenience of description, the visual entities in the reference image and the modification actions in the modification text are hereinafter referred to as "entities" and "actions" respectively. Although some researchers have tried to effectively retrieve the target image based on cross-modal semantic alignment, they have not fully considered the inherent semantic asymmetry between the reference image and the modification text, that is, the entities included in the action do not necessarily have a direct semantic correspondence in the reference image, resulting in room for improvement in retrieval performance. Therefore, in order to construct an effective combined image retrieval method, not only the semantic alignment between visual data and text data needs to be considered, but also the binding of modification relationships needs to be considered. Currently, there are mainly the following three challenges in realizing the binding of the modification relationship between entities and actions to construct an effective combined image retrieval method:

[0004] (1) Irrelevant factor perturbation. Before binding the modification relationship, it is first necessary to mine entities and actions. However, not all words in the text and all regions in the image are directly related to the modification requirements, and the irrelevant words and visual regions will affect the mining of entities and behaviors. Therefore, it is very challenging to identify the visual and text semantics related to the modification behavior and exclude the interference of irrelevant factors.

[0005] (2) Fuzzy semantic boundaries. Actions are often hidden in different combinations of words in the modification text, and there is no clear supervision signal to achieve accurate division. In addition, due to the irregular shape of the entity, its boundary is also difficult to identify. In view of this, it is also challenging to identify semantic boundaries to mine entities and actions.

[0006] (3) Implicit modification relationship. Since there may be no semantic match between entities and actions, it is challenging to measure the corresponding relationship of the modification relationship only by feature similarity, and there is a lack of direct supervision signals. Therefore, it is very challenging to identify the modification relationship and bind the entity to the corresponding operation. Summary of the Invention

[0007] In view of the deficiencies of the prior art, the present invention proposes a combined image retrieval method based on entity mining and modification relationship binding to achieve effective retrieval of user target images.

[0008] The present invention also proposes a combined image retrieval system based on entity mining and modification relationship binding.

[0009] To this end, in the present invention, first, a potential factor filtering module is proposed to calculate cross-modal semantic relevance and filter out visual and text potential factors related to modification semantics; second, the present invention proposes an entity-action binding module to study the semantic boundaries of visual and text potential factors, aggregate them into visual entities and modification actions respectively, and simultaneously learn the implicit modification relationship between them to achieve entity-action binding; finally, the present invention proposes a multi-scale combination module to construct multi-modal query features at multiple scales under the guidance of the entity-action binding relationship to achieve effective retrieval of user target images.

[0010] Term Explanation:

[0011] 1. CLIP is a deep learning model designed to combine text and image information through contrastive learning methods to achieve cross-modal understanding. CLIP is pre-trained on a large-scale image-text pair dataset and can be used for various tasks, such as image retrieval, text generation, image classification, etc. Its cross-modal feature learning ability makes it perform excellently in processing visual and language information.

[0012] 2. The attention mechanism is a computational method used in deep learning models, aiming to dynamically focus on specific parts of the input data to improve the efficiency and effect of information processing.

[0013] 3. Multilayer Perceptron (MLP) is a feedforward neural network composed of at least three layers of nodes: an input layer, a hidden layer, and an output layer. Each layer consists of multiple neurons, and the neurons are connected by weights. It is widely used in tasks such as classification, regression, and feature extraction, and shows excellent performance in multiple fields such as pattern recognition, image processing, and natural language processing.

[0014] 4. The Softmax function is an activation function widely used in the output layer of multi-classification problems. Its main function is to convert the input real-valued vector into a probability distribution, making each output value range from 0 to 1, and the sum of all output values equal to 1.

[0015] 5. Average Pooling is a downsampling technique widely used in convolutional neural networks, aiming to reduce the spatial dimension of data and extract important features by calculating the regional average of the input feature map.

[0016] 6. The identity matrix is a special square matrix, whose main feature is that the elements on the diagonal are all 1, and the elements off the diagonal are all 0, usually denoted as I n , where n represents the dimension of the matrix.

[0017] 7. The Frobenius norm is a mathematical metric used to measure the size or "length" of a matrix, and the definition of this norm is the square root of the sum of the squares of all elements in the matrix.

[0018] 8. The KL divergence (Kullback-Leibler Divergence) is an asymmetric metric for measuring the difference between two probability distributions. Specifically, the KL divergence is used to evaluate the information loss or relative entropy of one probability distribution relative to another probability distribution.

[0019] The technical solution of the present invention is as follows:

[0020] A combined image retrieval method based on entity mining and modified relation binding, including:

[0021] Read the training set data in batches, and extract the global features and local features of the training set data; filter the local features through latent factors, and further extract the latent factor features of vision and text related to modified semantics;

[0022] For the latent factor features, combined with entity-action binding, deeply mine the semantic relationships in the reference image and the modified text, detect the semantic boundaries, and aggregate the latent factors into visual entities and modified actions respectively;

[0023] For the obtained features of different scales, through multi-scale combination, obtain the final combined features;

[0024] Calculate the dot product of the combined features and different images in the image library respectively as the similarity score, sort the similarity scores in descending order, and select the target images with the top several similarity scores to complete the combined image retrieval.

[0025] As a further preferred solution, the local features are filtered by latent factors to further extract the latent factor features of vision and text related to the modified semantics, including:

[0026] Based on the local features, calculate the cross-relationship scores between each visual factor and text factor through cross-modal cross-attention and intra-modal self-attention; at the same time, calculate the intra-modal self-attention scores.

[0027] Combine the self-attention scores and cross-relationship scores to distinguish the latent factors and shelved factors of vision and text.

[0028] Use a multi-layer perceptron to obtain the adaptive weights of the latent factors to enhance the weights of the factors more relevant to the modified semantics.

[0029] Further preferably, based on the local features, calculate the cross-relationship scores between each visual factor and text factor through cross-modal cross-attention and intra-modal self-attention; at the same time, calculate the intra-modal self-attention scores, including:

[0030] Based on the local features of the reference image, calculate the cross-modal cross-relationship score w c and the intra-modal self-attention score w s , which are expressed as follows:

[0031]

[0032] where w c , w s ∈R C ; Inter-CA is cross-modal cross-attention, Intra-SA is intra-modal self-attention, Q and K are feature vectors for calculating attention weights, V is a vector representing the input features, refers to the local features of the modified text; refers to the local features of the reference image x r .

[0033] Further preferably, combine the self-attention scores and cross-relationship scores to distinguish the latent factors and shelved factors of vision and text, including:

[0034] Combine the cross-relationship score w c and the self-attention score w s , and the factors exceeding the specified threshold σ are retained as the latent factors for subsequent entity-action binding. At the same time, for the factors below the specified threshold σ, they are represented as shelved factors.

[0035] Further preferably, use a multi-layer perceptron to obtain the adaptive weights of the latent factors, including:

[0036] Through the multi - layer perceptron and the Softmax function, the adaptive weights of the latent factors are obtained. In the reference image, the enhanced latent factors The final representation is as follows:

[0037]

[0038] Among them, w # ∈R 0×2 represents the factor weight, P represents the number of latent factors to be learned, refers to the latent factors of the modified text, refers to the shelved factors of the modified text, Softmax(·) is the activation function, MLP(·) is the multi - layer perceptron. In addition, average pooling is also used for the shelved factors; similarly, the final latent factor features of the modified text and the final latent factor features of the target image

[0039] As a further preferred solution, for the latent factor features, combined with entity - action binding, deeply explore the semantic relationships in the reference image and the modified text, detect the semantic boundaries, and aggregate the latent factors into visual entities and modification actions respectively; including:

[0040] Train a learnable relationship query shared by modalities to explore the semantic relationships in the reference image and the modified text, detect the semantic boundaries, and aggregate their latent factors into visual entities and modification actions respectively;

[0041] Through the learnable relationship query, learn the implicit modification relationship between the visual entity and the modification action, and use it as a medium to assist entity - action binding.

[0042] Further preferably, for the latent factor features, combined with entity - action binding, deeply explore the semantic relationships in the reference image and the modified text, detect the semantic boundaries, and aggregate the latent factors into visual entities and modification actions respectively; including:

[0043] Initialize the learnable relationship query R∈R 8×3 , where E is the number of queries and D is the global embedding dimension of CLIP;

[0044] First, perform an interaction between the learnable relationship queries to calculate the self - attention weights e s , and also calculate the cross - attention weights between the learnable relationship query and the latent factors of different modalities and The representation is as follows:

[0045]

[0046] Among them, e s∈R 8×8 , is the cross-attention weight calculated for the potential factors of the learnable relationship query and the reference image, is the cross-attention weight calculated for the potential factors of the learnable relationship query and the modified text;

[0047] Based on e s , the adaptive attention weights of the learnable relationship query for the potential factors of the reference image and the potential factors of the modified text are respectively and

[0048] Assign e r and e m to the corresponding potential factors respectively, and use a multi-layer perceptron to adaptively learn the weights, so as to map the potential factors to the corresponding tokens, where each token matches the corresponding query of the learnable relationship query, as shown below:

[0049]

[0050] Among them, and correspond to the entity token and the action token respectively;

[0051] Design a binding orthogonality loss function for the learnable relationship query as shown below:

[0052]

[0053] where I ∈ R 8×8 , is the Frobenius norm of the matrix;

[0054] Calculate the similarity distributions of visual entities and modified actions using the learnable relationship query respectively; let represent the similarity distribution between the i-th token in the entity token and the learnable relationship query, where is shown as follows:

[0055]

[0056] where s(·) represents the cosine similarity, E rL , R Q represents the i-th entity token and the j-th query in the learnable relationship query, τ is the temperature coefficient, E represents the number of triples, and R A represents the e-th query in the learnable relationship query; similarly, obtain the similarity distribution between each token in the action token and each query of the learnable relationship query, denoted as Among them, Denote the similarity between the $i$-th action token and the $j$-th learnable relation query;

[0057] Calculate the KL divergence loss function between two similarity distributions To bind the corresponding entity with the action, it is expressed as follows:

[0058]

[0059] where, is the KL divergence loss function, $D 2` (·)$ is the KL divergence, Denote the similarity between the $i$-th target image token and the $j$-th learnable relation query, Denote the similarity between the $i$-th action token and each learnable relation query, Denote the similarity between the $i$-th entity token and each learnable relation query.

[0060] As a further preferred solution, for the features of different scales obtained, through multi-scale combination, the final combined feature is obtained; including:

[0061] Integrate the multi-scale features of the reference image and the modified text respectively;

[0062] Based on the obtained multi-scale features, perform multi-scale interaction via a multi-layer perceptron to learn the modification weights of the reference image and the modified text respectively;

[0063] Based on the learned modification weights, obtain the final combined feature.

[0064] As a further preferred solution, for the features of different scales obtained, through multi-scale combination, the final combined feature is obtained; including:

[0065] Concatenate the corresponding global feature, local feature, latent factor and shelving factor; for the reference image, the multi-scale feature is expressed as $Q = 1 + P + E + 1$; similarly, obtain the multi-scale feature of the modified text and the multi-scale feature of the target image $Q m = 1 + P + 1$;

[0066] Use a multi-layer perceptron to perform multi-scale interaction on $E r and $E m to learn the modification weights of the reference image and the modified text respectively, which is expressed as follows:

[0067] $W = MLP([E r , E m )

[0068] where, $W \in \mathbb{R}$i×J3 , split W into W r ∈R i×3 and W m ∈R i×3 , as the weights of the reference image and the modified text respectively;

[0069] Aggregate the modification weights with the multi-scale features E r 、E m of the corresponding reference image and modified text to obtain the aggregated reference image feature W r E r and the modified text feature W m E m , then sum W r E r and W m E m to obtain the final multi-modal combined feature E c , expressed as follows:

[0070] E c =W r E r +W m E m ,

[0071] where E c ∈Q×D, E r is the multi-scale feature of the reference image, and E m is the multi-scale weight of the modified text;

[0072] Use the batch-based classification loss to make the combined feature approach the target image feature. The batch-based classification loss is expressed as follows:

[0073]

[0074] where represents the result of average pooling of E c and E 6 corresponding to the i-th triple, B represents the batch size, and s(·) represents the cosine similarity;

[0075] Let represent the similarity distribution of the i-th combined feature, where the similarity with the j-th target image is calculated as follows:

[0076]

[0077] Similarly, obtain the similarity distribution between the i-th target image feature and other target images in the batch, denoted as Subsequently, the KL divergence is used to converge these two similar distributions to optimize the combined feature space, which is expressed as follows:

[0078]

[0079] The following final objective function is obtained:

[0080]

[0081] where Θ * is the parameter to be optimized for the model ENCODER, and κ, τ, μ are trade-off hyperparameters.

[0082] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the combined image retrieval method based on entity mining and modified relationship binding are implemented.

[0083] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the combined image retrieval method based on entity mining and modified relationship binding are implemented.

[0084] A combined image retrieval system based on entity mining and modified relationship binding includes:

[0085] A latent factor filtering module configured to: read the training set data in batches and extract global features and local features from the training set data; filter the local features through latent factors to further extract visual and text latent factor features related to modified semantics;

[0086] An entity-action binding module configured to: for the latent factor features, combine entity-action binding, deeply mine the semantic relationships in the reference image and the modified text, detect the semantic boundaries, and aggregate the latent factors into visual entities and modified actions respectively;

[0087] A multi-scale combination module configured to: for the obtained features of different scales, obtain the final combined features through multi-scale combination;

[0088] An image retrieval module configured to: calculate the dot product of the combined features with different images in the image library respectively as the similarity scores, sort the similarity scores in descending order, and select the target images with the top several similarity scores to complete the combined image retrieval.

[0089] Compared with the prior art, the beneficial effects of the present invention are:

[0090] 1. The present invention proposes a potential factor filtering module, which helps to filter the potential factors of vision and text, and as much as possible, excludes the interference of irrelevant factors in the process of identifying and modifying the visual and text semantics related to the behavior.

[0091] 2. The present invention proposes an entity-action binding module, which aims to deeply explore the relationship between entities and actions and detect semantic boundaries. Under the challenge of implicit modification relationships, this module realizes the effective binding of entities and actions, further improving the accuracy of combined image retrieval.

[0092] 3. The present invention proposes a multi-scale combination module, which generates the final combined features by integrating the multi-scale features of the reference image and the modified text. Under the guidance of entity-action binding, this module successfully enhances the multi-scale semantic perception of multi-modal combined features, thereby further improving the accuracy of combined image retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] Figure 1 is a schematic flow chart of the combined image retrieval method based on entity mining and modification relationship binding of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0094] The present invention will be further limited below in conjunction with the accompanying drawings of the specification and embodiments, but not limited thereto.

[0095] Embodiment 1

[0096] The combined image retrieval method based on entity mining and modification relationship binding, as Figure 1 shown, realizes the combined image retrieval of multi-modal queries by constructing a combined image retrieval method ENCODER based on entity mining and modification relationship binding. This method first performs potential factor filtering based on the multi-modal query semantics, and then based on the potential factors, studies the semantic boundaries of visual potential factors and text potential factors, aggregates them into visual entities and modification actions respectively, and simultaneously learns the implicit modification relationship between them. Finally, according to the entity-action binding relationship, multi-modal combined features are constructed at multiple scales, and the candidate target image set is screened through the combined features to obtain the final result set, realizing combined image retrieval. Specifically, it includes:

[0097] S1: Potential factor filtering: Batch-read the training set data, and extract the global features and local features of the training set data; filter the local features through potential factors to further extract the potential factor features of vision and text related to the modification semantics;

[0098] S2: Entity-action relationship binding: For the potential factor features, in combination with entity-action binding, deeply explore the semantic relationship in the reference image and the modified text, detect the semantic boundary, and aggregate the potential factors into visual entities and modification actions respectively;

[0099] S3: Multi-scale feature combination: For the features of different scales obtained, through multi-scale combination, the final combined features are obtained;

[0100] Finally, in the inference stage, the model ENCODER of the present invention (i.e., the combined image retrieval method of the present invention) passes through latent factor filtering, entity-action relationship binding, and multi-scale feature combination, performs multi-modal query combination on the input reference image and modified text, and obtains combined features. Calculate the dot product of the combined features with different images in the image library respectively as the similarity score, sort the similarity scores in descending order, and select the target images (i.e., the formal result set) with the top several (K) similarity scores according to the needs (such as retrieving the K images that best meet the query requirements) to complete the combined image retrieval.

[0101] Example 2

[0102] The combined image retrieval method based on entity mining and modification relationship binding according to Example 1 is characterized in that:

[0103] In this method, the training set data is read in batches, and the global features and local features of the training set data are extracted; including:

[0104] Read the training set data in batches; the form of each group of data in each batch is a triple of <reference image, modified text, target image>; through CLIP, the global feature of the reference image x r is extracted and the local feature which are expressed as:

[0105]

[0106] wherein, and respectively represent the last layer and the penultimate layer of the image encoder of CLIP, D is the CLIP global embedding dimension, FC I aligns the local embedding dimension to D; similarly, the global feature and the local feature of the target image and the global feature and the local feature of the modified text are obtained, where C and S respectively represent the number of channels of the image and the text.

[0107] In this method, the local features are filtered through latent factors to further extract the latent factor features of vision and text related to the modified semantics; including:

[0108] Based on local features, calculate the cross-relationship scores between each visual factor and text factor through cross-modal cross-attention and intra-modal self-attention; at the same time, also calculate the intra-modal self-attention scores;

[0109] Combine the self-attention scores and cross-relationship scores to distinguish the potential factors and shelved factors of vision and text;

[0110] Use a multi-layer perceptron to obtain the adaptive weights of the potential factors to enhance the weights of factors more relevant to modifying semantics.

[0111] To make full use of cross-modal semantic associations, based on local features, calculate the cross-relationship scores between each visual factor and text factor through cross-modal cross-attention and intra-modal self-attention; at the same time, also calculate the intra-modal self-attention scores; including:

[0112] Taking the reference image as an example, mainly based on the local features of the reference image and supplemented by modifying the local features of the text, calculate the inter-modal cross-relationship score w c and the intra-modal self-attention score w s , which are expressed as follows:

[0113]

[0114] where w c , w s ∈R C ; Inter-CA is cross-modal cross-attention, Intra-SA is intra-modal self-attention, Q and K are feature vectors for calculating attention weights, and V is a vector representing input features, all of which are obtained from the input features, refers to modifying the local features of the text; refers to the local features of the reference image x r .

[0115] For the modified text, mainly based on the local features of the modified text and supplemented by the local features of the reference image. For the target image, use its local features for calculation.

[0116] Combine the self-attention scores and cross-relationship scores to distinguish the potential factors and shelved factors of vision and text; including:

[0117] Combine the cross-relationship score w c and the self-attention score w s, the present invention implements a threshold gating mechanism. Factors exceeding the specified threshold σ (based on experimental experience, σ = 0.6 is taken) are retained as potential factors for subsequent entity-action binding. At the same time, for factors below the specified threshold σ, since they may appear as reserved regions in the target image, they are retained and represented as shelved factors. Taking the reference image as an example, its potential factors and shelved factors are formulated as follows:

[0118]

[0119] Similarly, the potential factors of the modified text are obtained and shelved factors The potential factors of the target image and shelved factors

[0120] The adaptive weights of the potential factors are obtained using a multi-layer perceptron; including:

[0121] Through the multi-layer perceptron and the Softmax function, the adaptive weights of the potential factors are obtained, enhancing the weights of the factors more relevant to the modified semantics, alleviating the strict boundaries of the potential factors, and preventing some shelved factors from being misclassified as potential factors. In the reference image, the enhanced potential factors are finally represented as follows:

[0122]

[0123] where, w # ∈R 0×2 , represents the factor weight, P represents the number of potential factors to be learned, refers to the potential factors of the modified text, refers to the shelved factors of the modified text, Softmax(·) is the activation function, MLP(·) is the multi-layer perceptron. In addition, in order to retain the semantic information of the shelved factors without distracting the attention from the potential factors, average pooling is also used for the shelved factors; similarly, the final potential factor features of the modified text and the final potential factor features of the target image

[0124] For the potential factor features, in combination with entity-action binding, the semantic relationships in the reference image and the modified text are deeply mined, the semantic boundaries are detected, and the potential factors are respectively aggregated into visual entities and modification actions; including:

[0125] Train a learnable relation query shared by modalities to mine the semantic relationships in the reference image and the modified text, detect the semantic boundaries, and aggregate their potential factors into visual entities and modification actions respectively;

[0126] Through learnable relation query, learn the implicit modification relationship between visual entities and modification actions to assist entity-action binding as a medium.

[0127] For latent factor features, combined with entity-action binding, deeply mine the semantic relationship in the reference image and the modification text, detect the semantic boundary, and aggregate the latent factors into visual entities and modification actions respectively; including:

[0128] Initialize the learnable relation query R ∈ ℝ 8×3 , where E is the number of queries and D is the global embedding dimension of CLIP;

[0129] First, perform an interaction between the learnable relation queries to calculate the self-attention weights e s , and also calculate the cross-attention weights between the learnable relation queries and the latent factors of different modalities and are expressed as follows:

[0130]

[0131] where e s ∈ ℝ 8×8 , is the cross-attention weight calculated between the learnable relation query and the latent factor of the reference image, is the cross-attention weight calculated between the learnable relation query and the latent factor of the modification text;

[0132] Based on e s , the adaptive attention weights of the learnable relation query for the latent factors of the reference image and the latent factors of the modification text are respectively and

[0133] Assign e r and e m to the corresponding latent factors respectively, and use a multi-layer perceptron to adaptively learn the weights, so as to map the latent factors to the corresponding tokens, where each token matches the corresponding query of the learnable relation query, and is expressed as follows:

[0134]

[0135] where, and correspond to the entity token and the action token respectively;

[0136] To ensure semantic independence and act as a medium for entity-action implicit modification relationship binding, design a binding orthogonality loss function for the learnable relation query is expressed as follows:

[0137]

[0138] where \(I\in R\) 8×8 , is the Frobenius norm of the matrix;

[0139] Calculate the similarity distributions of visual entities and modification actions using learnable relation queries respectively; let represent the similarity distribution between the \(i\)-th token in the entity token and the learnable relation query, where is expressed as follows:

[0140]

[0141] where \(s(\cdot)\) represents the cosine similarity, \(E\) rL , \(R\) Q represents the \(j\)-th query in the \(i\)-th entity token and the learnable relation query, \(\tau\) is the temperature coefficient, \(E\) represents the number of triples, \(R\) A represents the \(e\)-th query in the learnable relation query; similarly, obtain the similarity distribution between each token in the action token and each query in the learnable relation query, denoted as where represents the similarity between the \(i\)-th action token and the \(j\)-th learnable relation query;

[0142] Calculate the KL divergence loss function between the two similarity distributions To bind the corresponding entities and actions, it is expressed as follows:

[0143]

[0144] where is the KL divergence loss function, \(D\) 2` (\cdot)\) is the KL divergence, represents the similarity between the \(i\)-th target image token and the \(j\)-th learnable relation query, represents the similarity between the \(i\)-th action token and each learnable relation query, represents the similarity between the \(i\)-th entity token and each learnable relation query.

[0145] In this method, for the obtained features of different scales, through multi-scale combination, the final combined feature is obtained; including:

[0146] Integrate the multi-scale features of the reference image and the modified text respectively;

[0147] Based on the obtained multi-scale features, perform multi-scale interaction via a multi-layer perceptron to learn the modification weights of the reference image and the modified text respectively;

[0148] Based on the learned modification weights, the final combined feature is obtained.

[0149] For the features of different scales obtained, through multi-scale combination, the final combined feature is obtained, including:

[0150] Concatenate the corresponding global feature, local feature, latent factor, and shelved factor; for the reference image, the multi-scale feature is expressed as Q = 1 + P + E + 1; similarly, the multi-scale feature of the modified text is obtained as well as the multi-scale feature of the target image Q m = 1 + P + 1;

[0151] Use a multi-layer perceptron to perform multi-scale interaction on E r and E m to learn the modification weights of the reference image and the modified text respectively, expressed as follows:

[0152] W = MLP([E r , E m )

[0153] where W ∈ R i×J3 , and use block operation to divide W into W r ∈ R i×3 and W m ∈ R i×3 , as the weights of the reference image and the modified text respectively;

[0154] Aggregate the modification weights with the multi-scale features E r , E m of the corresponding reference image and modified text respectively to obtain the aggregated reference image feature W r E r and the modified text feature W m E m , then sum W r E r and W m E m to obtain the final multi-modal combined feature E c , expressed as follows:

[0155] E c = W r E r + W m E m ,

[0156] where E c ∈ Q × D, E r is the multi-scale feature of the reference image, E mTo modify the multi-scale weights of the text;

[0157] Using the batch-based classification loss, the combined features are made to approach the target image features to promote the correctness of the combination. The batch-based classification loss is expressed as follows:

[0158]

[0159] Where, represents E corresponding to the i-th triple c and E 6 after average pooling, B represents the batch size, and s(·) represents the cosine similarity;

[0160] Let be expressed as the similarity distribution of the i-th combined feature, where the similarity with the j-th target image is calculated as follows:

[0161]

[0162] Similarly, the similarity distribution between the i-th target image feature and other target images in the batch is obtained, denoted as Subsequently, the KL divergence is used to converge these two similarity distributions to optimize the combined feature space, expressed as follows:

[0163]

[0164] The final objective function is obtained as follows:

[0165]

[0166] Where, Θ * are the parameters to be optimized of the model ENCODER, and κ, τ, μ are the trade-off hyperparameters.

[0167] Table 1 is the comparison table of the retrieval accuracy of the present invention on the FashionIQ dataset;

[0168] Table 1

[0169]

[0170] Table 2 is the comparison table of the retrieval accuracy of the present invention on the Shoes dataset;

[0171] Table 2

[0172]

[0173] Table 3 is the comparison table of the retrieval accuracy of the present invention on the CIRR dataset;

[0174] Table 3

[0175]

[0176] Table 4 is a comparison table of the retrieval accuracy of the present invention on the Fashion200K dataset;

[0177] Table 4

[0178]

[0179] As shown in Tables 1-4, a comparison of query efficiency and retrieval accuracy with internationally leading similar methods is carried out and the results are shown. The results show that, compared with similar combined image retrieval methods, the ENCODER model of the present invention has higher accuracy on four widely used benchmark datasets.

[0180] Embodiment 3

[0181] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the combined image retrieval method based on entity mining and modified relationship binding described in Embodiment 1 or 2 are implemented.

[0182] Embodiment 4

[0183] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the combined image retrieval method based on entity mining and modified relationship binding described in Embodiment 1 or 2 are implemented.

[0184] Embodiment 5

[0185] A combined image retrieval system based on entity mining and modified relationship binding includes:

[0186] A latent factor filtering module is configured to: batch-read training set data, and extract global features and local features from the training set data; filter the local features through latent factors to further extract visual and text latent factor features related to modified semantics;

[0187] An entity-action binding module is configured to: for the latent factor features, in combination with entity-action binding, deeply mine the semantic relationships in the reference image and the modified text, detect the semantic boundaries, and aggregate the latent factors into visual entities and modified actions respectively;

[0188] A multi-scale combination module is configured to: for the obtained features of different scales, obtain the final combined features through multi-scale combination;

[0189] The image retrieval module is configured to: calculate the dot product of the combined features and different images in the image library respectively as the similarity scores, sort the similarity scores in descending order, and select the target images with the top several similarity scores to complete the combined image retrieval.

Claims

1. A combined image retrieval method based on entity mining and modified relationship binding, characterized in that: include: The training set data is read in batches, and global and local features are extracted from the training set data; the local features are filtered through latent factors, and the latent factor features of vision and text related to the modified semantics are further extracted; For latent factor features, we combine entity-action binding to deeply explore the semantic relationship between reference images and modified texts, detect semantic boundaries, and aggregate latent factors into visual entities and modification actions respectively; For the features of different scales, the final combined features are obtained through multi-scale combination; The dot products of the combined features and different images in the gallery are calculated as similarity scores, and the similarity scores are sorted in descending order. The target images with the top similarity scores are selected to complete the combined image retrieval.

2. The combined image retrieval method based on entity mining and modified relationship binding according to claim 1 is characterized in that: The local features are filtered through latent factors to further extract the latent factor features of vision and text related to the modified semantics; including: Based on local features, the cross-relationship score between each visual factor and text factor is calculated through cross-modal cross-attention and intra-modal self-attention; the intra-modal self-attention score is also calculated; Combining self-attention scores and cross-relation scores to distinguish latent factors and shelving factors from visual and textual ones; The adaptive weights of latent factors are obtained by using multi-layer perceptron to enhance the weights of factors that are more relevant to the modified semantics.

3. The combined image retrieval method based on entity mining and modified relationship binding according to claim 2 is characterized in that: Based on local features, the cross-relationship score between each visual factor and text factor is calculated through cross-modal cross-attention and intra-modal self-attention. The intra-modal self-attention score is also calculated. This includes: Based on the local features of the reference image, the inter-modal cross-correlation score w is calculated c and the intra-modal self-attention score w s , which is expressed as follows: Among them, w c , w s ∈R C ; Inter-CA is inter-modality cross attention, Intra-SA is intra-modality self-attention, Q and K are feature vectors for calculating attention weights, and V is a vector representing input features. It refers to modifying local features of the text; refers to the reference image x r Local features of Further preferably, the self-attention score and the cross-relation score are combined to distinguish the latent factors and the shelving factors of vision and text; including: Combined cross-relation score w c and the self-attention score w s , factors exceeding the specified threshold σ are retained as potential factors for subsequent entity action binding, while factors below the specified threshold σ are represented as shelved factors; Further preferably, a multi-layer perceptron is used to obtain the adaptive weights of the potential factors; including: Through the multi-layer perceptron and Softmax function, the adaptive weights of the latent factors are obtained. In the reference image, the enhanced latent factors The final expression is as follows: Among them, w f ∈R P×K , represents the factor weight, P represents the number of potential factors to be learned, refers to the potential factors that modify the text, refers to the shelving factor of the modified text, Softmax(·) is the activation function, MLP(·) is the multi-layer perceptron, and average pooling is also used for the shelving factor; similarly, the final latent factor feature of the modified text is obtained and the final latent factor features of the target image 4. The combined image retrieval method based on entity mining and modified relationship binding according to claim 1, characterized in that: For latent factor features, combined with entity-action binding, the semantic relationship between reference images and modified texts is deeply explored, semantic boundaries are detected, and latent factors are aggregated into visual entities and modification actions respectively; including: Train modality-shared learnable relation queries to mine semantic relations in reference images and modified texts, detect semantic boundaries, and aggregate their latent factors into visual entities and modification actions, respectively; Through learnable relation query, the implicit modification relation between visual entities and modification actions is learned as a medium to assist entity-action binding.

5. The combined image retrieval method based on entity mining and modified relationship binding according to claim 3 is characterized in that: For latent factor features, combined with entity-action binding, the semantic relationship between reference images and modified texts is deeply explored, semantic boundaries are detected, and latent factors are aggregated into visual entities and modification actions respectively; including: Initialize the learnable relation query R∈R E×D , where E is the number of queries and D is the global embedding dimension of CLIP; First interact with the learnable relationship queries and calculate the self-attention weight e s , and also calculates the cross-attention weights between the learnable relation query and the latent factors of different modalities and It is expressed as follows: Among them, e s ∈R E×E , Cross-attention weights calculated for latent factors that can learn the relationship between the query and the reference image, Cross-attention weights calculated for latent factors that can learn the relationship between query and modified text; Based on e s , The adaptive attention weights of the learnable relation query for the latent factors of the reference image and the latent factors of the modified text are and E r and e m The corresponding latent factors are assigned to them respectively, and the weights are adaptively learned using a multi-layer perceptron to map the latent factors to the corresponding tokens, where each token matches the corresponding query of the learnable relation query, as shown below: in, and They correspond to entity tokens and action tokens respectively; Designing bound orthogonalized loss functions for learnable relational queries It is expressed as follows: Where I∈R E×E , is the Frobenius norm of the matrix; Use learnable relation queries to calculate the similarity distribution of visual entities and modification actions respectively; let represents the similarity distribution between the ith token in the entity token and the learnable relation query, where It is expressed as follows: Among them, s(·) represents the cosine similarity, E ri , R j represents the jth query in the i-th entity token and learnable relation query, τ is the temperature coefficient, E represents the number of triples, R e represents the e-th query in the learnable relation query; similarly, the similarity distribution between each token in the action token and each query in the learnable relation query is obtained, denoted as in, represents the similarity between the i-th action token and the j-th learnable relation query; Calculate the KL divergence loss function between two similarity distributions To bind the corresponding entities and actions, as shown below: in, is the KL divergence loss function, D KL (·) is the KL divergence, represents the similarity between the i-th target image token and the j-th learnable relation query, represents the similarity between the ith action token and each learnable relation query, represents the similarity between the i-th entity token and each learnable relation query.

6. The combined image retrieval method based on entity mining and modified relationship binding according to claim 1, characterized in that: For the features of different scales, the final combined features are obtained through multi-scale combination, including: Integrate multi-scale features of reference image and modified text respectively; Based on the obtained multi-scale features, multi-scale interaction is performed via a multi-layer perceptron to learn the modification weights of the reference image and the modified text; Based on the learned modification weights, the final combined features are obtained.

7. The combined image retrieval method based on entity mining and modified relationship binding according to claim 5 is characterized in that: For the features of different scales, the final combined features are obtained through multi-scale combination, including: The corresponding global features, local features, latent factors and shelving factors are concatenated; for the reference image, the multi-scale features are expressed as Q = 1 + P + E + 1; Similarly, the multi-scale features of the modified text are obtained And the multi-scale features of the target image Q′=1+P+1; Using multi-layer perceptron in E r and E m Multi-scale interaction is performed on the reference image and the modified text to learn the modification weights of each, which is expressed as follows: In=MLP([E r ,E m ]) Where W∈R Q×2D , use block operation to split W into W r ∈R Q×D and W m ∈R Q×D , as the respective weights of the reference image and the modified text; The modification weights are respectively compared with the corresponding multi-scale features E of the reference image and the modified text. r 、E m Aggregate and obtain the aggregated reference image feature W r E r and modify the text feature W m E m , and then to W r E r and W m E m Sum and get the final multimodal combination feature E c , which is expressed as follows: HAVE BEEN c =W r HAVE BEEN r +W m HAVE BEEN m , Among them, E c ∈Q×D,E r is the multi-scale feature of the reference image, E m To modify the multi-scale weights of the text; Using batch-based classification loss, the combined features are approached to the target image features. The batch-based classification loss is expressed as follows: in, Indicates E corresponding to the i-th triple c and E t The result after average pooling, B represents the batch size, s(·) represents the cosine similarity; make It is represented as the similarity distribution of the i-th combined feature, where the similarity with the j-th target image is The calculation is as follows: Similarly, the similarity distribution between the i-th target image feature and other target images in the batch is obtained, which is recorded as Then KL divergence is used to converge these two similar distributions and optimize the combined feature space, which is expressed as follows: The final objective function is as follows: Among them, Θ * are the parameters to be optimized of the ENCODER model, and κ, τ, and μ are the trade-off hyperparameters.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the combined image retrieval method based on entity mining and modified relationship binding are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a combined image retrieval method based on entity mining and modified relationship binding are implemented.

10. A combined image retrieval system based on entity mining and modified relationship binding, characterized in that: include: The latent factor filtering module is configured to: read the training set data in batches, and extract global features and local features from the training set data; filter the local features through latent factors, and further extract the latent factor features of vision and text related to the modified semantics; The entity-action binding module is configured to: for latent factor features, combine entity-action binding, deeply mine the semantic relationship between reference image and modified text, detect semantic boundaries, and aggregate latent factors into visual entities and modification actions respectively; The multi-scale combination module is configured to: obtain the final combined features through multi-scale combination for the obtained features of different scales; The image retrieval module is configured to: use the combined features to screen images in the candidate target image set to obtain a formal result set, that is, a final result of the combined image retrieval.

Citation Information

Patent Citations

  • Methods for labeling and searching advanced semantics of imagse based on network hot topics and device

    CN102902821A

  • Multi-label image retrieval method and equipment based on direct-push type zero sample hash

    CN110795590A

  • Image retrieval method and system based on multi-modal query, medium and equipment

    CN113239219A

  • Search device and method for biological system information using keyword hierarchy

    KR1020180106677A

  • System for providing machine learning based textile product searching service including advanced matching algorithm

    KR102099561B1

Cited By

  • Multi-modal combined image retrieval method fusing fine-grained semantic positioning and optimization generation features

    CN121958587A