An Image Retrieval Method Based on the Synthesis of Hard Negative Sample Representations

By building a network of difficult-negative samples representation generation, combining the image representation extraction module, the global association learning module between samples and the channel diversity interpolation module of association perception, the problem of lack of global category distribution considerations in the synthesis of difficult-negative samples representation in the existing technology is solved, and the efficient and accurate classification of the category boundaries of the image retrieval model is achieved.

CN119397047BActive Publication Date: 2025-06-27SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411521888.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-06-27
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

The prior art lacks consideration for global category distribution when synthesis of difficult negative sample characterization and synthesis, resulting in inappropriate characterization difficulty and insufficient category correlation, which affects the accuracy of the image retrieval model.

Method used

A hard-negative sample representation generation network is constructed that includes image representation extraction module, global correlation learning module between samples and channel diversity interpolation modules that are related to the channel. Through the category balanced sampling strategy and global correlation learning, difficult-negative sample representations with appropriate synthesis and category-relatedness are characterized.

Benefits of technology

It effectively improves the distinction and accuracy of the image retrieval model in the measurement space, and improves the efficiency and accuracy of image retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119397047B_ABST
    Figure CN119397047B_ABST
Patent Text Reader

Abstract

The present invention discloses an image retrieval method based on the synthesis of hard negative sample representations, comprising the following steps: constructing a hard negative sample representation generation network capable of synthesizing information-rich representations; using a batch of images constructed with a class-balanced sampling strategy as the network input, extracting a batch of image representations through an image representation extraction module, then inputting the batch of image representations into a global inter-sample correlation learning module for learning, and inputting pairs of samples that are negative to each other into a correlation-aware channel diversity interpolation module to synthesize hard negative sample representations; training the hard negative sample representation generation network capable of synthesizing information-rich representations, and jointly training the image representation extraction module using the synthesized hard negative sample representations and real sample representations; using the trained image representation extraction module for image retrieval; by combining the global inter-sample correlation learning ability, synthesizing more informative hard negative sample representations to guide the image representation extraction module to extract more discriminative image representations to enhance the retrieval performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing and artificial intelligence, and particularly relates to an image retrieval method based on the synthesis of difficult negative sample representations. Background Art

[0002] In the digital age, image retrieval technology combines traditional visual understanding with modern deep learning methods, providing an efficient way to quickly retrieve target images similar to a query image from a large image database. The core of the image retrieval task lies in designing a metric method that can accurately quantify the similarity of samples in a metric space, such that similar images are close to each other in this space, while dissimilar images are far apart. To improve the discrimination ability of the image retrieval model, this task aims to synthesize information-rich difficult negative sample representations for enhancing the discrimination at the class boundary in the metric space. This plays an important role in achieving efficient and accurate image retrieval in complex class distribution scenarios.

[0003] Current research on the synthesis of difficult negative sample representations usually fuses a small number of sample pairs or triplet samples through a generative adversarial network (GAN) or an auto-encoder (Auto-Encoder) to synthesize difficult negative sample representations. These methods mainly focus on the local correlation between a few selected samples, so there may be a problem of inappropriate difficulty level in the generated negative sample representations. This is because they lack consideration of the potential influence of other classes. In this case, the method of synthesizing difficult negative sample representations that only considers local sample correlation may synthesize negative sample representations that do not match the target class distribution, thus interfering with the training process of the image retrieval model and resulting in limited retrieval performance. This highlights the necessity of considering a wider range of negative sample class correlations in the process of synthesizing difficult negative sample representations, in order to more accurately grasp the similarity between different classes, thereby generating negative sample representations with appropriate difficulty and class relevance to promote the retrieval model to distinguish class boundaries.

[0004] The main challenge of the image retrieval task based on the synthesis of difficult negative sample representations lies in how to learn and utilize the association information of a set of sample pairs in the global class distribution, so as to synthesize information-rich difficult negative sample representations and use them to enhance the accuracy of the image retrieval model. Summary of the Invention

[0005] In view of this, in order to solve the above technical problems, the present invention provides an image retrieval method based on the synthesis of difficult negative sample representations, including the following steps:

[0006] Step 1, construct a difficult negative sample representation generation network capable of synthesizing information-rich difficult negative sample representations, including an image representation extraction module, a global association learning module between samples, and an association-aware channel diversity interpolation module;

[0007] Step 2: Use the batch of images constructed by the class-balanced sampling strategy as the network input. Extract the batch of image representations through the image representation extraction module. Then, input the batch of image representations into the inter-sample global correlation learning module to learn the correlation between each anchor sample and all other negative samples. Input the sample pairs that are negative classes to the correlation-aware channel diversity interpolation module, and interpolate and synthesize the hard negative sample representations according to the correlation of the sample pairs.

[0008] Step 3: Train the network capable of synthesizing rich information hard negative sample representation generation network, and jointly train the image representation extraction module with the synthesized hard negative sample representations and real sample representations.

[0009] Step 4: Use the trained image representation extraction module for image retrieval.

[0010] Furthermore, the image representation extraction module includes a convolutional neural network and a fully connected layer. The specific steps for extracting the batch of image representations are as follows:

[0011] Step 20201: After inputting the batch of images into the convolutional neural network, compress them through two-dimensional max pooling to obtain a one-dimensional tensor.

[0012] Step 20202: The one-dimensional tensor is processed through batch normalization, a fully connected layer, and L2 norm normalization to obtain the batch of image representations.

[0013] Among them, the batch of image representations contains the representations of all samples in the batch of images.

[0014] Furthermore, the inter-sample global correlation learning module includes a node message propagation network and an edge message propagation network. The specific steps for building the inter-sample global correlation learning module are as follows:

[0015] Step 20301: The inter-sample global correlation learning module constructs the batch of image representations into a graph structure, which includes several nodes and edges. Each node is the representation of a single sample, and each edge is the dot product of the representation points of the sample pairs that are negative classes.

[0016] Step 20302: The node message propagation network acts on all the nodes in the entire graph structure. Among them, each node fuses the adjacent nodes and edges through the Transformer mechanism, so that the node learns the global sample correlation information.

[0017] Step 20303: The edge message propagation network acts on each pair of negative sample pairs to integrate the global sample correlation information learned by the nodes into the edges.

[0018] Step 20304: After two iterations of the node message propagation network and the edge message propagation network, the construction of the global association learning module between samples is completed.

[0019] Furthermore, the association-aware channel diversity interpolation module includes an association-aware sample pair interpolation module and an intra-class diversity interpolation module. The steps for building the association-aware channel diversity interpolation module are as follows:

[0020] Step 20401: The graph structure obtains the diversity interpolation weight of each negative class sample pair, the training loss of the image representation extraction module, and the distance relationship interpolation between samples in the batch image through the association-aware sample pair interpolation module, and synthesizes the difficult negative sample representation between the negative class sample pairs;

[0021] Step 20402: The intra-class diversity interpolation module randomly weights all the hard negative sample representations of the same negative class corresponding to each anchor sample to mine potential hard negative sample representations.

[0022] Furthermore, the specific steps of step 3 are as follows:

[0023] Step 20501: Use a training function to train the network capable of synthesizing information-rich difficult negative sample representation generation network;

[0024] Step 20502: Optimize the image representation extraction module using an optimization function;

[0025] Step 20503: During the training of the network capable of synthesizing information-rich hard negative sample representation generation, overlapping training is performed on step 20501 and step 20502.

[0026] Further, the training function includes classification loss, similarity loss and diversity loss;

[0027] The classification loss ensures that the corresponding negative class distribution is maintained between the synthesized representations of the hard negative samples through a classification head network;

[0028] The similarity loss makes the representation of the hard negative sample similar to the representation of the anchor sample under the premise that the classification loss maintains the corresponding negative class distribution between the representations of the synthetic hard negative samples;

[0029] The diversity loss improves the standard deviation of the interpolation weights to expand the search space for synthesizing the hard negative sample representation within a reasonable constraint.

[0030] Further, the optimization function includes global association learning loss, synthetic hard negative sample representation optimization loss and real sample representation optimization loss;

[0031] The global correlation learning loss enables the nodes to maintain category semantic information by imposing category constraints on the nodes;

[0032] The synthetic hard negative sample representation optimization loss optimizes the metric space by combining the similarity relationship of sample pairs of the real sample representation and the hard negative sample representation;

[0033] The real sample representation optimization loss optimizes the metric space by combining the similarity relationship of sample pairs of the real sample representation.

[0034] Furthermore, the specific steps of step 4 are as follows:

[0035] The image representation extraction module extracts database image representations from all images in the database and saves them. After inputting the query image to the image representation extraction module to extract the query image representation, the cosine similarity between the query image representation and the database image representations is calculated, and the top K results in the descending order of similarity are selected as the output.

[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0037] The present invention effectively considers the correlation between samples in the global category distribution, can synthesize hard negative sample representations with appropriate difficulty and category relevance, helps the image retrieval model to efficiently distinguish category boundaries, and improves the accuracy of image retrieval;

[0038] The present invention also extracts image representations through the image representation extraction module. The proposed global correlation learning module between samples can learn the correlation between each sample and all other negative samples in the batch in the global category distribution; the proposed correlation-aware channel diversity interpolation module interpolates and synthesizes hard negative sample representations according to the learned correlation between samples; alternately trains the hard negative sample representation generation network that can synthesize rich information hard negative sample representations and jointly trains the image representation extraction module with the synthesized hard negative sample representations and the real sample representations to guide the image representation extraction module to efficiently distinguish category boundaries and continuously synthesize more difficult hard negative sample representations to enhance the discriminability of the model in the metric space, and achieve accurate image retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 Shows a schematic flow diagram of the implementation method of the present invention;

[0040] Figure 2 Shows a schematic structural diagram of the operation of each module of the embodiment of the present invention;

[0041] Figure 3 Shows a schematic structural diagram of the category-balanced sampling strategy. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0043] For the sake of citation and clarity, the technical terms, abbreviations, or acronyms used hereinafter are summarized and explained as follows:

[0044] Sigmoid / ReLU: Non-linear activation function.

[0045] ResNet50: A convolutional neural network belonging to the family of residual networks, consisting of 50 layers, including convolutional layers, batch normalization layers, ReLU activation functions, and fully connected layers.

[0046] Transformer: A neural network architecture for processing sequence data.

[0047] Batch Normalization: A method for accelerating the training of neural networks and improving the stability of the model. Its core idea is to normalize the batch image representations of each batch of images to alleviate the problems of vanishing gradients and exploding gradients, and to make the training of the network more stable and converge faster.

[0048] Metric space: A metric space refers to the geometric space in which samples are embedded as feature representations, and the distance between samples can be measured by a certain metric function (such as Euclidean distance, cosine similarity, etc.).

[0049] Anchor sample: An anchor sample refers to a reference sample used to learn the distance or similarity between samples. Its main role is to compare with other samples and calculate the similarity difference between positive and negative samples relative to it.

[0050] Positive sample: A positive sample refers to a sample that belongs to the same class as the anchor sample. The model expects the distance between them and the anchor sample in the metric space to be as close as possible, that is, the similarity is high.

[0051] Negative sample: A negative sample refers to a sample that does not belong to the same class as the anchor sample. The model expects the distance between them and the anchor sample in the metric space to be as far as possible, that is, the similarity is low.

[0052] Positive sample pair: A positive sample pair refers to a pair of samples of the same class, usually including an anchor sample and a positive sample.

[0053] Negative sample pair: A negative sample pair refers to a combination of samples of two different classes, usually including an anchor sample and a negative sample.

[0054] Hard negative samples: Hard negative samples refer to those negative samples that come from different categories than anchor samples but are close in the metric space. Such samples are challenging to train the model because the model needs to distinguish them as much as possible rather than mistakenly believe that they belong to the same category.

[0055] L2 norm normalization: L2 norm normalization refers to scaling a vector by its L2 norm so that the L2 norm of the normalized vector is equal to 1. This is often used in machine learning to make data more stable or easier to calculate, especially when dealing with feature vectors.

[0056] The invention discloses an image retrieval method based on hard negative sample representation synthesis to solve many problems existing in the prior art.

[0057] Figure 1 A flow chart of an embodiment of the present invention is shown, which is an image retrieval method based on hard negative sample representation synthesis, comprising the following steps:

[0058] Step 1: construct a network capable of synthesizing information-rich hard negative sample representation generation, including an image representation extraction module, a global association learning module between samples, and a correlation-aware channel diversity interpolation module;

[0059] Step 2: The batch images constructed by the class-balanced sampling strategy are used as network input, and the batch image representations are extracted by the image representation extraction module. The batch image representations are then input into the global association learning module between samples to learn the correlation between each anchor sample and all other negative samples. The sample pairs that are mutually negative are input into the association-aware channel diversity interpolation module and the hard negative sample representations are interpolated and synthesized according to the correlation of the sample pairs.

[0060] Step 3, training the network capable of synthesizing information-rich hard negative sample representation generation, and using the synthesized hard negative sample representation and the real sample representation to jointly train the image representation extraction module;

[0061] Step 4: Use the trained image representation extraction module to perform image retrieval.

[0062] Specifically, Figure 2As shown, the inter-sample global correlation learning module and the correlation-aware channel diversity interpolation module in step 2 are only used to synthesize hard negative sample representations during the model training process. The hard negative sample representations are then combined with the real sample representations to jointly train the image representation extraction module for image retrieval, enabling the image representation extraction module to extract more discriminative real sample representations from the input batch of images and improving its retrieval accuracy in the metric space. During the testing process, only the image representation extraction module is needed to extract image representations from the query image (input image) and the database images respectively for similarity measurement to achieve image retrieval.

[0063] Furthermore, the specific operation steps of step 2 in this embodiment are as follows;

[0064] Step 201. Use the batch of images constructed with the class-balanced sampling strategy as the network input;

[0065] The steps of the class-balanced sampling strategy are as follows:

[0066] Each time a batch of images is constructed, 27 classes are sampled from the entire dataset, with 3 samples sampled from each class, for a total of 81 samples; among them, the first 27 samples in the batch of images are 1 of the samples sampled from each class, and the subsequent samples after these 27 samples are also arranged in the same class order; as Figure 3 shown, sample classes, with m samples sampled from each class. Therefore, the first samples in the batch of images are x 1,1 to The subsequent second group samples are x 2,1 to The subsequent mth group samples are x m,1 to A total of samples are obtained; the role of the class-balanced sampling strategy is to avoid lack of attention to classes with a small number of samples caused by random sampling;

[0067] Next, first randomly crop and distort the batch of images to a resolution of 224*224, and then apply the batch of images after random horizontal flip augmentation as the network input.

[0068] Step 202. Construct an image representation extraction module, including a ResNet50 convolutional neural network and a fully connected layer; among them, the specific steps for extracting the batch of image representations are as follows:

[0069] Step 20201: After inputting the batch of images into the ResNet50 convolutional neural network and downsampling by 32 times to obtain a tensor of dimension 7×7×2048, compress it into a one-dimensional tensor of dimension 2048 through two-dimensional max pooling;

[0070] Step 20202: After performing Batch Normalization on the one-dimensional tensor, input it into a fully connected layer and compress it into a one-dimensional tensor of dimension 512, and finally obtain the batch image representation z through L2 norm normalization; wherein, the batch image representation z contains the representations of all samples within the batch of images;

[0071] Step 203. Build a global inter-sample correlation learning module, and the global inter-sample correlation learning module includes a node message propagation network and an edge message propagation network;

[0072] Step 20301: The global inter-sample correlation learning module constructs the batch image representation into a graph structure, and the representation of each sample in the batch image representation is connected to the representations of all mutually negative samples. The graph structure includes several nodes and edges. Each node is defined as the representation of a single sample, and each edge is defined as the dot product of the representations of the mutually negative sample pairs;

[0073] Step 20302: The node message propagation network acts on all nodes in the entire graph structure so that each sample in the batch of images can learn a more accurate similarity relationship with all negative samples within the batch of images from a global perspective; wherein, each node fuses the adjacent nodes and edges through the Transformer mechanism, thereby effectively learning global sample correlation information to enhance the node's perception of global correlation; the mathematical expression of the node message propagation network is as follows:

[0074]

[0075] represents the representation of the i-th node after the k-th action of the node message propagation network, SA(V k ) i represents taking out the representation of the i-th node after acting on all nodes in the k-th time using the self-attention mechanism, represents the number of samples in the batch of images, LN represents layer normalization, and FFN represents the feed-forward network; represents the representation of the adjacent edge between the i-th node and the j-th node after the k-th action of the edge message propagation network, represents the intermediate variable for node update of the i-th node in the (k + 1)-th action of the node message propagation network;

[0076] Step 20303. The edge message propagation network acts on each pair of negative sample pairs to incorporate the global sample association information learned by the nodes into the edges, so as to help the association-aware channel diversity interpolation module generate the hard negative sample representation. The mathematical expression of the edge message propagation network is as follows:

[0077]

[0078] represents the representation of the adjacent edge between the i-th node and the j-th node after the k-th edge message propagation network operation. CA represents the cross-attention mechanism; represents the representation of the i-th node after the (k + 1)-th node message propagation network operation, represents the intermediate variable of edge update in the (k + 1)-th edge message propagation network operation for the adjacent edge between the i-th node and the j-th node;

[0079] Step 20304. After the iterative operations of the node message propagation network and the edge message propagation network are performed twice respectively, the global association learning between samples is completed, so that each edge in the graph structure can effectively perceive the correlation of global samples and the construction of the global association learning module between samples is completed;

[0080] Step 204. Build an association-aware channel diversity interpolation module, and the association-aware channel diversity interpolation module includes an association-aware sample pair interpolation module and an intra-class diversity interpolation module;

[0081] Step 20401. The steps of the association-aware sample pair interpolation module are as follows: First, the representations of all edges in the graph structure are normalized by a sigmoid function after passing through a fully connected layer to provide channel-level diversity interpolation weights for each negative sample pair. Then, based on the interpolation weights, the training loss of the image representation extraction module, and the distance relationship between positive and negative samples within the batch of images, interpolation is performed to synthesize the hard negative sample representation between negative class sample pairs. Its mathematical expression is as follows:

[0082]

[0083]

[0084] Among them, represents the representation of the adjacent edge between the i-th node and the j-th node after the K-th edge message propagation network operation. FC represents a fully connected layer. The sigmoid function is used for numerical normalization. In a negative class sample pair, the anchor sample representation is defined as z i , and the negative sample representation opposite to the anchor is defined as z j , d - is the distance between the anchor sample and the negative sample, d+ is the distance between the anchor point and the positive sample, α is a hyperparameter, J avg Represents the metric loss value of the current image representation extraction module training; λ ij Represents the correlation interpolation of sample pairs, that is, in z i As anchor point samples, z j As the interpolated vector representation of negative samples, Indicates that in z i As anchor point samples, z j As negative samples, they are represented by synthetic hard negative samples, and e represents a natural constant;

[0085] Step 20402: The intra-class diversity interpolation module randomly weights all synthetic hard negative sample representations of the same negative class corresponding to each anchor sample to mine as many potential hard negative sample representations as possible. The mathematical expression is as follows:

[0086]

[0087] Indicates that the category is the nth category and the anchor sample is z i The set of all synthetic hard negative sample representations; l j represents the category of the jth class), RW represents the fusion of two difficult negative sample representations in the set by random weighting to obtain a new interpolation vector representation, and then repeats the above operation between the new interpolation vector representation and the next difficult negative sample representation until all difficult negative sample representations of the set are traversed, thereby obtaining each anchor point sample z i Representation of hard negative samples between the nth category

[0088] Step 205: training the network capable of synthesizing information-rich hard negative sample representation generation, and using the synthesized hard negative sample representation and the real sample representation to jointly train the image representation extraction module;

[0089] Step 20501, training the network capable of synthesizing information-rich difficult negative sample representation generation, including three parts of the loss function: classification loss, through a classification head network to ensure that the synthesized difficult negative sample representation accurately maintains the negative class distribution corresponding to the difficult negative sample representation, that is, to guide the network capable of synthesizing information-rich difficult negative sample representation generation to synthesize another difficult negative sample representation as close as possible to the negative class distribution corresponding to a difficult negative sample representation; similarity loss, under the premise that the classification loss maintains the category relevance of the synthesized difficult negative sample representation, that is, the corresponding negative class distribution, the difficult negative sample representation is made as similar as possible to the anchor sample representation; diversity loss, which is used to increase the standard deviation of the interpolation weight to expand the search space of the synthesized difficult negative sample representation within a reasonable constraint range; the mathematical expression of the training function composed of these three together is as follows:

[0090]

[0091] J div = 1 - σ(λ i. )

[0092]

[0093] where C z is a classification head network, CE represents cross-entropy loss, l' n represents the label of class n, σ represents the standard deviation function, γ s is the weight of the similarity loss, γ d is the weight of the diversity loss, represents the number of classes sampled within the batch of images, G represents the global inter-sample association learning module; J ce is the classification loss, J sim is the similarity loss, J div is the diversity loss, λ i. represents the set of interpolation vector characterizations of the anchor sample z i and all negative samples, l i represents the class of the i-th class, represents the number of sample sets in the batch of images, means minimizing the optimization function to optimize the global inter-sample association learning module G, J gen is the loss generated by optimizing the hard negative sample characterization by the training function.

[0094] The classification head network C z also needs to optimize the weights of the classification head network through the real sample characterization, and the mathematical expression is as follows:

[0095]

[0096] J cz represents the classification loss of the real sample characterization;

[0097] Step 20502, jointly optimize the image characterization extraction module using the synthesized hard negative sample characterization and the real sample characterization, including three parts of the loss function: global association learning loss, by imposing class constraints on the nodes, so that the nodes can still maintain the class semantic information of the nodes after perceiving the global association in step 203; synthesized hard negative sample characterization optimization loss, optimizing the metric space by combining the similarity relationship of the sample pairs of the real sample characterization and the synthesized hard negative sample characterization; real sample characterization optimization loss, optimizing the metric space by combining the similarity relationship of the real sample characterization sample pairs; these three together form the optimization function, and the mathematical expression is as follows:

[0098]

[0099] represents the representation of the \(i\)-th node after passing through the node message propagation network \(K\) times, \(C\) v represents another classification head, the \(C\) v and \(C\) z Although they have the same structure but different weights, their objects of action are also different; represents a positive sample corresponding to the anchor sample \(z\) in the batch of images, \(m\) represents the number of samples sampled for each category in the batch of images, i \(n\) represents the number of categories sampled in the batch of images, \(F\) represents the weights of the image representation extraction module, \(\beta\) is a hyperparameter, \(J\) represents the global correlation learning loss, \(J\) gca represents the synthetic hard negative sample representation optimization loss, \(J\) syn represents the real sample representation optimization loss \(J\) r (\(z\)), \(J\) m is the loss of the optimization function, represents the inner product of the positive sample pair, represents the inner product of the negative sample pair, \(q\) represents the \(q\)-th class, is that is, the number of negative samples, means minimizing the optimization function to optimize the image representation extraction module \(F\), the global correlation learning module \(G\) between samples, and the other classification head \(C\) v ;

[0100] The optimization process of the optimization function is a process of gradually introducing the synthetic hard negative sample representation into the optimization of the image representation extraction module as the \(J\) gen loss decreases;

[0101] Step 20503, during the process of training the network capable of generating synthetic informative hard negative sample representations, alternately train Step 20501 and Step 20502;

[0102] Step 206. Use the trained image representation extraction module for image retrieval; the image representation extraction module extracts and saves the database image representations of all images in the candidate database, and after inputting the query image to the image representation extraction module to extract the query image representation, calculate the cosine similarity between the query image representation and the database image representations, and select the top \(K\), that is, the top \(K\) results in the similarity descending order as the output.

[0103] In the present invention, the AdamW optimizer is used during the training process, with a weight decay factor of 0.0001, a learning rate of 0.00015 for the network layer of the image feature extraction module, a learning rate of 0.0003 for the network layer of the global inter-sample correlation learning module, a learning rate of 0.001 for the network layer of the classification head, hyperparameters α of 5, β of 2, and γ s is 1, and γ d is 0.01. The batch size is set to 81, and it is trained for 50 epochs on 1 RTX3090 GPU.

[0104] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0105] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0106] The technical features of the above embodiments can be combined arbitrarily. At the same time, the labels of each step are not used to restrict the sequence between steps. As long as there is no strict sequence constraint between steps, their sequence is allowed to be adjusted and transformed. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not conflict, it should be considered to be within the scope described in this specification.

[0107] The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An image retrieval method based on hard negative sample representation synthesis, characterized in that: The following steps are involved: Step 1: construct a network capable of synthesizing information-rich hard negative sample representation generation, including an image representation extraction module, a global association learning module between samples, and a correlation-aware channel diversity interpolation module; Step 2: The batch images constructed by the class-balanced sampling strategy are used as network input, and the batch image representations are extracted by the image representation extraction module. The batch image representations are then input into the global association learning module between samples to learn the correlation between each anchor sample and all other negative samples. The sample pairs that are mutually negative are input into the association-aware channel diversity interpolation module, and the hard negative sample representations are interpolated and synthesized according to the correlation of the sample pairs. Step 3, training the network capable of synthesizing information-rich hard negative sample representation generation, and using the synthesized hard negative sample representation and the real sample representation to jointly train the image representation extraction module; Step 4, use the trained image representation extraction module to perform image retrieval; The image representation extraction module includes a convolutional neural network and a fully connected layer; the batch image representation extraction steps are as follows: Step 20201: After the batch images are input into the convolutional neural network, they are compressed by two-dimensional maximum pooling to obtain a one-dimensional tensor; Step 20202: The one-dimensional tensor is processed by batch normalization, a fully connected layer, and L2 norm normalization to obtain the batch image representation; Wherein, the batch image representation includes the representation of all samples in the batch image; The inter-sample global association learning module includes a node message propagation network and an edge message propagation network. The specific steps of building the inter-sample global association learning module are as follows: Step 20301: The sample global association learning module constructs the batch image representation into a graph structure, wherein the graph structure includes a plurality of nodes and edges, each of the nodes is a representation of a single sample, and each of the edges is a dot product of the representations of the mutually negative sample pairs; Step 20302: the node message propagation network acts on all the nodes in the entire graph structure; wherein each node fuses the adjacent nodes and edges through the Tranformer mechanism, so that the node learns the global sample association information; Step 20303: the edge message propagation network acts between each pair of negative class samples to integrate the global sample association information learned by the node into the edge; Step 20304: After two iterations of the node message propagation network and the edge message propagation network, the construction of the global association learning module between samples is completed.

2. The image retrieval method based on hard negative sample representation synthesis according to claim 1, characterized in that: The association-aware channel diversity interpolation module includes an association-aware sample pair interpolation module and an intra-class diversity interpolation module. The steps for building the association-aware channel diversity interpolation module are as follows: Step 20401: The graph structure obtains the diversity interpolation weight of each negative class sample pair, the training loss of the image representation extraction module, and the distance relationship interpolation between samples in the batch image through the association-aware sample pair interpolation module, and synthesizes the difficult negative sample representation between the negative class sample pairs; Step 20402: The intra-class diversity interpolation module randomly weights all the hard negative sample representations of the same negative class corresponding to each anchor sample to mine potential hard negative sample representations.

3. The image retrieval method based on hard negative sample representation synthesis according to claim 2, characterized in that: The specific steps of step 3 are as follows: Step 20501: Use a training function to train the network capable of synthesizing information-rich difficult negative sample representation generation network; Step 20502: Optimize the image representation extraction module using an optimization function; Step 20503: During the training of the network capable of synthesizing information-rich hard negative sample representation generation, overlapping training is performed on step 20501 and step 20502.

4. The image retrieval method based on hard negative sample representation synthesis according to claim 3, characterized in that: The training function includes classification loss, similarity loss and diversity loss; The classification loss ensures that the corresponding negative class distribution is maintained between the synthesized representations of the hard negative samples through a classification head network; The similarity loss makes the representation of the hard negative sample similar to the representation of the anchor sample under the premise that the classification loss maintains the corresponding negative class distribution between the representations of the synthetic hard negative samples; The diversity loss improves the standard deviation of the interpolation weights to expand the search space for synthesizing the hard negative sample representation within a reasonable constraint.

5. The image retrieval method based on hard negative sample representation synthesis according to claim 3, characterized in that: The optimization function includes global association learning loss, synthetic hard negative sample representation optimization loss and real sample representation optimization loss; The global association learning loss enables the node to maintain category semantic information by performing category constraints on the node; The synthetic hard negative sample representation optimizes the loss, and optimizes the metric space by combining the similarity relationship between the sample pairs of the real sample representation and the hard negative sample representation; The real sample representation optimizes the loss, and the similarity relationship between sample pairs represented by the real samples is combined to optimize the metric space.

6. The image retrieval method based on hard negative sample representation synthesis according to claim 1, characterized in that: The specific steps of step 4 are as follows: The image representation extraction module extracts database image representations from all images in the database and saves them. After the query image is input into the image representation extraction module to extract the query image representation, the cosine similarity between the query image representation and the database image representation is calculated, and the top K results in descending order of similarity are selected as output.

Citation Information

Patent Citations

  • Target accurate retrieval method and system based on difficult sample generation

    CN110674692A

  • Comparative learning enhanced two-stream model recommendation system and algorithm

    WO2023108324A1