Deep Image Retrieval Method, System and Storage Medium Based on Metric Learning
By introducing optimization strategies and hash codecs based on metric learning in the depth image search method, the problems of insufficient feature representation capabilities and information loss in the prior art are solved, and more efficient and accurate image retrieval is achieved.
Patent Information
- Application Number
- CN202510353245.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing image hashing method fails to effectively utilize the advantages of metric learning, resulting in insufficient feature representation ability and low retrieval accuracy, and the critical information will be lost during hash encoding, affecting the retrieval performance.
The deep image retrieval method based on metric learning is adopted, and the similarity between the feature vector and the agent is optimized by constructing the image training set and hash center, so that the feature vector converges to the same type of agent. Combining the metric learning optimization strategy based on proxy and paired samples, a hash codec is built to achieve the reversality and measurement invariance of hash encoding.
It improves the expression ability and retrieval accuracy of hash encoding, reduces the loss of information during the encoding process, enhances the discrimination ability and consistency of hash code, and optimizes the stability and distinction of hash mapping.
Smart Images

Figure CN119862292B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image retrieval, and in particular to a deep image retrieval method, system and storage medium based on metric learning. Background Art
[0002] 1) Image retrieval systems based on image hashing;
[0003] With the rapid development of information technology and the popularization of the Internet, a vast amount of image data is continuously generated and stored, making efficient image retrieval technology an important research direction in the fields of computer vision and artificial intelligence. In recent years, due to the breakthroughs in deep learning, content-based image retrieval (CBIR) technology has received extensive attention and has been applied in multiple fields. Compared with traditional retrieval methods that rely on text tags, CBIR systems not only use the visual features of images for similarity matching but also can combine cross-modal information to further improve the retrieval accuracy, providing a more intelligent search experience in application scenarios such as e-commerce, medical diagnosis, and security monitoring.
[0004] The core of content-based image retrieval lies in two key steps: the calculation of feature description vectors and feature matching, as Figure 3 shown. Traditional methods based on handcrafted features can capture local features of images, but they have poor robustness in complex scenarios and are difficult to adapt to the diversity of large-scale image sets. In recent years, due to their advantages in learning high-level semantic features, deep neural networks have gradually become the mainstream method, extracting more discriminative features through end-to-end training and optimizing the retrieval efficiency by combining index structures. Usually, hashing methods and quantization methods are used to calculate feature description vectors. Hashing methods map continuous image features to a discrete binary space through hashing coding technology, significantly reducing storage and distance calculation costs. The quantization method divides high-dimensional feature vectors into finite subspaces and efficiently calculates feature distances using a lookup table, thereby reducing the computational complexity. With the continuous growth of the scale of image databases and the increasing complexity of image content, how to optimize the computational efficiency while ensuring the retrieval accuracy remains a key problem to be solved urgently.
[0005] In image hashing, the hash function maps image features to compact binary hash codes to achieve efficient feature matching and similarity calculation. The core idea is to convert similar images into similar hash codes, thereby accelerating the indexing and comparison processes. Image hashing methods can be divided into traditional hashing methods and deep learning hashing methods. Traditional hashing methods mainly rely on mathematical transformations and statistical characteristics to reduce the dimensionality and binarize image features, including Locality-Sensitive Hashing (LSH), Spectral Hashing (SH), and hashing methods based on Principal Component Analysis (PCA), etc. These methods have a relatively low computational cost and are suitable for scenarios with relatively regular data distributions. However, when dealing with complex image data, their feature expression ability is relatively limited, and it is difficult to ensure high retrieval accuracy.
[0006] With the rapid development of deep learning, hashing methods based on neural networks have been widely studied and have achieved excellent performance on multiple tasks. Deep hashing methods usually use convolutional neural networks to extract image features and combine specific loss functions to optimize the hash coding, so as to improve the feature expression ability and retrieval accuracy. Usually, these methods optimize feature learning and hash coding in an end-to-end manner, thereby improving the discriminative ability of the hash codes. For example, Deep Neural Network Hashing (DNNH) uses triplet ranking loss for similarity learning to optimize the discrimination ability of the hash codes; HashNet adopts weighted maximum likelihood estimation and introduces weights in the pairwise loss function to alleviate the problem of data imbalance and improve the robustness of the model. Compared with traditional hashing methods, deep learning hashing methods can adapt to different data distributions and automatically learn the optimal hash representation, thus maintaining high retrieval accuracy and generalization ability in complex data environments. Deep hashing methods have been widely applied in fields such as image retrieval, face recognition, and anomaly detection and have achieved good results.
[0007] 2) Metric learning;
[0008] Metric Learning is a machine learning method used to optimize the similarity metric between data. Its goal is to construct an optimal distance metric in a specific task, so that similar samples are closer in the feature space, while dissimilar samples maintain a large gap. Compared with traditional methods that rely on manually designed features and predefined distance functions, metric learning can adaptively adjust the structure of the feature space according to the task requirements, thereby improving the performance of tasks such as classification, clustering, and retrieval. By optimizing the metric function from supervised data, this method makes the feature representation more in line with the task requirements and improves the discriminative ability and generalization ability of the model.
[0009] Early metric learning methods mainly relied on linear transformations and optimization strategies to learn the similarity metric between data. Among them, Mahalanobis distance learning optimized the performance of nearest neighbor classification by constructing a transformation matrix to make the data more separable in the new feature space. With the development of deep learning, deep metric learning methods have become a research hotspot and achieved remarkable progress in multiple tasks. Compared with traditional linear metric learning methods, deep metric learning relies on the powerful feature extraction ability of neural networks and can learn more expressive non-linear metric functions. Current research mainly focuses on pair-based and proxy-based methods. Pair-based methods, such as contrastive learning and siamese networks, usually combine convolutional neural networks to extract image features and optimize the similarity relationship between samples through contrastive loss or triplet loss. These methods can effectively enhance the discriminative ability of feature embeddings and have been widely applied to tasks such as signature verification, face recognition, and fingerprint matching. However, due to the exponential growth of the number of sample pairs or triplets with the data scale, their sample efficiency is low and the training cost is high. In contrast, proxy-based methods optimize by learning proxies to replace pair samples. For example, Proxy Anchor Loss directly optimizes the distance between samples and their class proxies to improve training efficiency. These methods perform well on large-scale datasets, but due to relying on global proxy information, they may ignore the fine-grained information of samples and affect the characterization of complex data structures. Therefore, how to maintain the fine-grained information of data while improving training efficiency is an important challenge in current deep metric learning research.
[0010] Disadvantage 1 of the prior art: Existing hashing methods fail to effectively utilize the advantages of metric learning. Existing image hashing methods usually directly encode image features into binary hash codes, without fully leveraging the role of metric learning in optimizing feature representations. Introducing metric learning and directly performing binary quantization on its results may introduce quantization errors, causing similar feature vectors to be mapped to distant hash codes in the hash space, thus affecting retrieval performance.
[0011] Disadvantage 2 of the prior art: Key information is lost during the image hashing encoding process. Since hashing encoding needs to map continuous features to a discrete binary space, information loss is inevitable. Especially when dealing with complex images, key details may not be fully retained, and similar images may be encoded as distant hash codes in the hash space, destroying the similarity structure of the original data and affecting retrieval accuracy and stability.
[0012] Disadvantages of the prior art 3: In the existing image hashing encoding process, only the loss based on pairwise samples or the centering loss is usually considered, resulting in limited information expression ability. The hashing method based on pairwise samples relies on pairwise supervision information, with low sample efficiency and difficulty in fully utilizing data structure information. Although the center-based loss can improve the global clustering effect, it may ignore fine-grained features, leading to insufficient characterization of local similarity, thereby affecting the expression ability and retrieval performance of hash codes. Summary of the Invention
[0013] Based on the technical problems existing in the background art, the present invention proposes a deep image retrieval method, system and storage medium based on metric learning to improve the expression ability and retrieval accuracy of hashing encoding.
[0014] The deep image retrieval method based on metric learning proposed by the present invention inputs a query image into an image hashing encoding model for image retrieval to obtain a list of similar images;
[0015] The training process of the image hashing encoding model is as follows:
[0016] Construct an image training set and a hashing center, and each time sample a batch of images from the training set and input them into the image hashing encoding model;
[0017] Extract features from the batch of images and project them into the embedding space to obtain feature vectors, and project the hashing center into the embedding space to obtain a proxy;
[0018] In the embedding space, optimize the similarity between the feature vectors and the proxy as well as between the feature vectors based on metric learning, so that the feature vectors converge to the same-class proxies and the proxies adjust towards the same-class feature vectors;
[0019] Impose a reciprocal constraint between the hashing encoder and decoder, and impose a metric invariance constraint on the hashing encoding process, so that the hash codes converge to the corresponding hashing centers.
[0020] Furthermore, in optimizing the similarity between the feature vectors and the proxy as well as between the feature vectors based on metric learning, an optimization strategy based on proxy-based metric learning is used to optimize the similarity between the feature vectors and the proxy, specifically:
[0021] Optimize the cosine similarity between the feature vectors and the same-class proxies until it exceeds a set first same-class similarity threshold, and the optimization intensity of the cosine similarity increases as the cosine similarity decreases;
[0022] Optimize the cosine similarity between the feature vectors and the different-class proxies until it is lower than a set first different-class similarity threshold, and the optimization intensity of the cosine similarity increases as the cosine similarity increases.
[0023] Further, in optimizing the similarity between the feature vector and the proxy as well as between feature vectors based on metric learning, an optimization strategy based on pairwise sample metric learning is adopted to optimize the similarity between feature vectors. Specifically:
[0024] Optimize the cosine similarity between the feature vectors of samples in the same category until it exceeds the set second same - class similarity threshold, and the optimization intensity of the cosine similarity increases as the cosine similarity decreases;
[0025] Optimize the cosine similarity between the feature vectors of samples in different categories until it is lower than the set second different - class similarity threshold, and the optimization intensity of the cosine similarity increases as the cosine similarity increases.
[0026] Further, in the model training stage, Gaussian distribution modeling is performed on the feature vectors in the embedding space. Each time of optimization, the feature vectors are sampled from the Gaussian distribution, and then the similarity between the sampled feature vectors and the proxy is optimized based on metric learning.
[0027] Further, the hash codec includes a hash encoder and a hash decoder, and the reciprocity of the encoding process of the hash encoder and the decoding process of the hash decoder is constrained.
[0028] Further, it is constrained that the cosine similarity between two hash codes after encoding remains the same as the cosine similarity between the two feature vectors before encoding during the hash encoding process.
[0029] Further, a total loss function is constructed to train the image hash encoding model. The total loss function includes a hash center loss function and a metric learning loss function in the embedding space as well as a hash encoder loss function ;
[0030] The hash center loss function is used to increase the Hamming distance between hash centers of different categories when parameterizing the hash center to meet the requirements of hash encoding;
[0031] The metric learning loss function in the embedding space is used to make the feature vectors converge to the same - class proxies and the proxies adjust towards the same - class feature vectors in the embedding space;
[0032] The hash encoder loss function is used to ensure that the encoding process and the decoding process are reciprocal and the encoding process has metric invariance.
[0033] Further, during the execution of the image hash encoding model, specifically:
[0034] In the offline state, the image hashing encoding model takes an image set as input and obtains an image set hash code. The image set is a collection of images that are pre-collected, stored, and used for image retrieval.
[0035] In the online retrieval state, the query image is input into the image hashing encoding model to obtain a query image hash code.
[0036] Calculate the Hamming distance between the query image hash code and each offline image hash code in the image set hash code, and sort them in ascending order. Take the top K offline images corresponding to the Hamming distances as the similar image list of the query image, where K is an integer.
[0037] A deep image retrieval system based on metric learning inputs the query image into the image hashing encoding model for image retrieval to obtain a similar image list.
[0038] The training process of the image hashing encoding model includes a training set construction module, a feature projection module, a hash center projection module, and a total loss module.
[0039] The training set construction module is used to construct an image training set and a hash center, and each time samples a batch of images from the training set and inputs them into the image hashing encoding model.
[0040] The feature projection module is used to extract features from the batch of images and project them into the embedding space to obtain feature vectors.
[0041] The hash center projection module is used to project the hash center into the embedding space to obtain a proxy.
[0042] The total loss module is used to optimize the similarity between the feature vectors and the proxy in the embedding space based on metric learning, so that the feature vectors converge to the same-class proxies and the proxies adjust to the same-class feature vectors; at the same time, impose a reciprocal constraint between the hash encoder and decoder, and impose a metric invariance constraint on the hashing encoding process, so that the hash codes converge to the corresponding hash centers.
[0043] A computer-readable storage medium, characterized in that a number of classification programs are stored on the computer-readable storage medium, and the number of classification programs are used to be called by a processor and execute the above-mentioned deep image retrieval method.
[0044] The advantages of the deep image retrieval method, system and storage medium based on metric learning provided by the present invention are as follows: introducing proxy-based metric learning in the deep hashing system to enhance the feature representation ability, and combining pairwise-sample-based metric learning to effectively capture fine-grained information; realizing hash encoding processing by constructing a hash codec (hash encoder and hash decoder), and the encoding process of the hash encoder has the invariance of cosine similarity to ensure the discriminative ability and consistency of the hash encoding, and ensure that the hash code converges to the corresponding hash center, thereby optimizing the stability and distinguishability of the hash mapping; realizing the reversible mapping between the feature vector in the embedding space and the hash code in the Hamming space through the hash codec to reduce the loss of information in the encoding process and improve the expression ability and retrieval accuracy of the hash encoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a schematic flow chart of the present invention;
[0046] Figure 2 is a training flow chart of the image hash encoding model;
[0047] Figure 3 is a schematic diagram of the existing content-based image retrieval method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] Next, the technical solution of the present invention will be described in detail through specific embodiments. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0049] As Figures 1 to 3 shown, for the deep image retrieval method based on metric learning proposed by the present invention, the query image is input into the image hash encoding model for image retrieval to obtain a list of similar images;
[0050] The training process of the image hash encoding model is as follows:
[0051] Step 1: Extract features from a batch of images and project them into the embedding space to obtain feature vectors, and project the hash center into the embedding space to obtain a proxy;
[0052] Step 2: In the embedding space, optimize the similarity between the feature vector and the proxy and between the feature vectors based on metric learning, so that the feature vector converges to the same-class proxy, and the proxy adjusts to the same-class feature vectors;
[0053] Step 3: Impose a reciprocal constraint between the hash encoders and decoders, and impose a metric invariance constraint on the hash encoding process, so that the hash codes converge to the corresponding hash centers.
[0054] Through Steps 1 to 3, this embodiment introduces proxy-based metric learning in the deep hash system to enhance the feature representation ability, and combines pairwise sample-based metric learning to effectively capture fine-grained information; realizes hash encoding processing by constructing hash encoders and decoders (hash encoder and hash decoder), and the hash encoder keeps the cosine similarity between two hash codes after encoding the same as the cosine similarity between two feature vectors before encoding during the encoding process to ensure the discriminative ability and consistency of hash encoding, and ensure that the hash codes converge to the corresponding hash centers, thereby optimizing the stability and distinctiveness of hash mapping; realizes the reversible mapping between the feature vectors in the embedding space and the hash codes in the Hamming space through the hash encoders and decoders to reduce the loss of information during the encoding process, and improve the expression ability and retrieval accuracy of hash encoding.
[0055] In one embodiment, a backbone network and a feature encoder are used to extract and project features of an image. Specifically: the backbone network encodes the image into image features represented by vectors, and the feature encoder projects the image features into the embedding space. The backbone network and the feature encoder can adopt existing network structures for image processing, which will not be elaborated in this embodiment.
[0056] In this embodiment, the construction process of the hash centers is as follows:
[0057] In the image hashing task, the hash centers are used to represent the binary hash codes of a certain category of images. There are two ways to construct the hash centers: predefined hash centers and parameterized hash centers.
[0058] The predefined hash centers are designed to make the Hamming distance between different categories as large as possible. When the length of the hash code is a power of 2 and the number of image categories does not exceed twice the length of the hash code, the Hadamard matrix can be used for construction. Specifically, if the length of the hash code is and the number of categories is , the first rows of the Hadamard matrix vertically concatenated with can be selected as the hash centers. If the relationship between the length of the hash code and the number of categories does not meet the above conditions, the hash centers can be generated by random sampling, that is, randomly sample times from the Bernoulli distribution (parameter 0.5) with length , and use the sampling results as the hash centers.
[0059] The parameterized hash center represents the hash center of each category using trainable vectors. During the training process, the image hash coding model optimizes the hash center by minimizing the metric learning loss and quantization loss, so as to increase the Hamming distance between the hash centers of different categories, and at the same time, constraints are imposed on the binarization of the hash center to meet the requirements of hash coding. The parameterized hash center can be dynamically adjusted according to the data distribution, thereby improving the effect of hash learning.
[0060] The metric learning loss of the hash center uses the first Proxy Anchor Loss function :
[0061] ; (1)
[0062] where represents the cosine similarity of two vectors, represents a threshold, is a scaling factor, represents the set of hash centers, represents the set of positive hash centers of the data in the batch image, and the positive hash center is the set of hash centers with the same class label as the batch image, represents the hash code of the batch image, is the hash code of the image, is divided into two sets, is the set of hash codes of the same class as the hash center and is the set of hash codes of different classes from the hash center .
[0063] The quantization loss function of the hash center , where is an index used to distinguish different loss functions and has no actual physical meaning:
[0064] ; (2)
[0065] where represents the hash code of the batch image.
[0066] The total hash center loss function is:
[0067] ; (3)
[0068] where is the weight coefficient used to balance the metric learning loss and quantization loss.
[0069] In one embodiment, before optimizing the similarity between the feature vector and the proxy based on metric learning, the feature vector is modeled as a Gaussian distribution to generate a feature representation in the embedding space, and then the similarity between the feature representation and the proxy is optimized based on metric learning. By modeling the feature vector in the embedding space as a Gaussian distribution to characterize the uncertainty in the image feature representation, the robustness of the hash coding is improved.
[0070] Specifically: First, use the feature encoder to project the image features into the embedding space and model them based on the Gaussian distribution, where the output of the feature encoder includes the mean vector and the log variance vector. In the training stage, sample the feature vectors from the Gaussian distribution for feature coding, and use the reparameterization method to achieve gradient backpropagation, that is, represent the feature vector as the product of the standard deviation vector and the noise sampled from the standard Gaussian distribution plus the mean vector. Model the feature vector as a Gaussian distribution in the embedding space. Each time of optimization, the feature vector is sampled from the Gaussian distribution, and then optimize the similarity between the sampled feature vector and the proxy based on metric learning. At the same time, use the hash decoder to map the hash center to obtain the proxy in the embedding space, so as to optimize the similarity between the feature vector and the proxy based on metric learning. In the testing stage, directly use the mean vector as the feature vector to reduce the influence of randomness and ensure the stability of feature coding.
[0071] In one embodiment, optimize the similarity between the feature vector and the proxy as well as between the feature vectors based on metric learning. Specifically: During the process of optimizing the feature coding in the embedding space, adopt two optimization strategies: proxy-based metric learning and pairwise-sample-based metric learning.
[0072] (a1)Adopt the optimization strategy of proxy-based metric learning to optimize the similarity between the feature vector and the proxy: Proxy-based metric learning constructs proxies and optimizes the similarity between the sample feature vectors and the proxies, making the distances of the same-class samples closer in the embedding space and improving the discrimination between different-class samples. Eventually, the feature vectors converge to the corresponding proxies, achieving global optimization and improving the discrimination and optimization efficiency of the feature representation.
[0073] Second Proxy Anchor Loss function is:
[0074] ; (4)
[0075] where, represents the cosine similarity of two vectors, represents a distance threshold, is a scaling factor, represents the set of proxies, Represents the set of positive proxies of data in the batch images, represents the feature vectors of the batch images, is divided into two sets, for the proxies the set of feature vectors of the same class, for the proxy the set of feature vectors of different classes.
[0076] It should be noted that the proxies in the embedding space are all obtained by the hash decoder from the hash center, and the positive proxies are the set of proxies with the same class labels as the batch images.
[0077] For ease of understanding, this loss function can be rewritten in the following form:
[0078] ; (5)
[0079] where , is a smooth approximation of the activation function ReLU, and the parameter refers to or in the above formula (4), , is the set or the set the total number of feature vectors in, i.e., Log-Sum-Exp, is a smooth approximation of the maximum function, and the parameter refers to or .
[0080] Based on formulas (4) and (5), it directly quantifies optimizing the similarity between the feature vectors and the proxies using the optimization strategy of proxy-based metric learning. Specifically: optimizing the cosine similarity between the feature vectors and the proxies of the same class until it exceeds the set first same-class similarity threshold, and the optimization intensity of the cosine similarity increases as the cosine similarity decreases; optimizing the cosine similarity between the feature vectors and the proxies of different classes until it is lower than the set first different-class similarity threshold, and the optimization intensity of the cosine similarity increases as the cosine similarity increases.
[0081] (a2) Optimize the similarity between the feature vectors using the optimization strategy of pairwise-sample-based metric learning: Pairwise-sample-based metric learning optimizes the similarity relationship of features between sample pairs, increasing the feature similarity between samples of the same class and decreasing the similarity between samples of different classes, thereby enhancing the discriminative ability of feature encoding.
[0082] The pairwise-sample loss function is:
[0083] ; (6)
[0084] Among them, is the feature vector of the batch images, is the set excluding the feature vector of the set, is the positive sample set of the feature vector (not including ), , is the negative sample set of the feature vector . Among them, the positive sample set is the set composed of the feature vectors of the sample images with the same category, and the negative sample set is the set composed of the feature vectors of the sample images with different categories.
[0085] Similarly, formula (6) can also be written as in the form of, and the above pairwise sample loss function can also be replaced with the existing triplet loss to construct a pairwise sample-based metric learning loss function.
[0086] Based on formula (6), directly quantify and optimize the similarity between feature vectors by using the optimization strategy of pairwise sample-based metric learning. Specifically: optimize the cosine similarity between the feature vectors of the same-category samples until it exceeds the set second same-category similarity threshold, and the optimization intensity of the cosine similarity increases as the cosine similarity decreases; optimize the cosine similarity between the feature vectors of different-category samples until it is lower than the set second different-category similarity threshold, and the optimization intensity of the cosine similarity increases as the cosine similarity increases.
[0087] For the optimization strategy of pairwise sample-based metric learning, this optimization strategy depends on a large number of sample pairs, and the optimization efficiency is relatively low, but it can capture more fine-grained feature information and optimize the local structure of the feature space. Combining the above two metric learning (proxy-based metric learning and pairwise sample-based metric learning) optimization strategies can further improve the overall discrimination of the feature space while maintaining the optimization efficiency.
[0088] Therefore, the metric learning loss function in the embedding space is:
[0089] ; (7)
[0090] Among them, is the weight coefficient, which is used to balance the proxy anchor loss and the pairwise sample loss.
[0091] The smooth activation functions in Formulas (4) and (6) can improve the stability of gradient calculation, and the approximate maximum value function therein can better optimize the metric relationship. The optimization process is calculated based on batch data to improve the training efficiency and enhance the feature representation ability. Through the above two optimization strategies, the consistency and discrimination ability of feature mapping are improved, thereby enhancing the accuracy of subsequent hash coding. Compared with the prior art, the optimized hash space feature distribution in this embodiment is more reasonable, making the feature representations of the same-class data more compact, while improving the distinguishability of different-class data.
[0092] In one embodiment, a hash codec is provided to implement the mapping conversion between the feature vectors in the embedding space and the hash codes in the Hamming space. The hash codec includes a hash encoder and a hash decoder.
[0093] The hash encoding and decoding process needs to satisfy the reciprocity of the encoding process and the decoding process, and ensure that the encoding process maintains the invariance of cosine similarity, that is, it is constrained that the cosine similarity between two hash codes after encoding is the same as the cosine similarity between the two feature vectors before encoding. To this end, the hash codec can be implemented based on matrix operations, where the decoding matrix is the transpose of the encoding matrix, the product of the transpose of the encoding matrix and the encoding matrix approximates the identity matrix, and at the same time, the product of the encoding matrix and its transpose also approximates the identity matrix, thereby ensuring the reciprocity of the encoding and decoding processes and making the encoding process maintain the metric relationship of the features in the embedding space.
[0094] Specifically, the hash encoding matrix is , and the hash decoding matrix is , and the loss function of hash encoding :
[0095] ; (8)
[0096] Wherein, represents the identity matrix, is the F-norm (Frobenius norm) of the matrix.
[0097] By setting the hash codec, the proxy corresponds to the hash center one by one. Therefore, in this embodiment, the parameters of the hash encoder are optimized based on the reciprocity of the hash encoder and the hash decoder and the consistency of cosine similarity, ensuring that the hash code has a higher similarity with the same-class hash center, making the feature vectors in the embedding space converge to the corresponding hash center after being encoded by the hash encoder, generating hash codes with good discrimination, thereby enhancing the stability and retrieval performance of hash mapping. Image hashing based on metric learning in the hash code embedding space such as Figure 2As shown, the hash decoder maps the hash center to an agent in the embedding space and makes the feature vectors of images of the same class maintain a small distance from their corresponding agents through metric learning. The hash encoder maintains the consistency of cosine similarity during the encoding process, thereby ensuring that the generated hash code is close enough to its corresponding hash center, improving the accuracy and stability of hash encoding. In addition, the hash codec guarantees the recoverability of the encoding process, effectively reducing information loss, thus enhancing the quality and retrieval performance of hash encoding.
[0098] In this embodiment, a total loss function is constructed to train the image hash encoding model. The total loss function includes a hash center loss function , a metric learning loss function in the embedding space , and a hash encoder loss function ;
[0099] The hash center loss function is used to increase the Hamming distance between hash centers of different classes in the case of parameterizing the hash center to meet the requirements of hash encoding; the metric learning loss function in the embedding space is used to make the feature vectors converge to the agents of the same class and the agents adjust to the feature vectors of the same class in the embedding space; the hash encoder loss function is used to ensure that the encoding process and the decoding process are reciprocal and the encoding process has metric invariance, making the hash code converge to the hash centers of the same class. Based on the total loss function, the model parameters in the image hash encoding model are optimized using the stochastic gradient descent method.
[0100] After the image hash encoding model is trained, it enters the execution phase. The depth image retrieval process of the online image is specifically as follows (b1) to (b3):
[0101] (b1) In the offline state, the image hash encoding model takes an image set as input and obtains the image set hash code. The image set is a set of images that are pre-collected, stored, and used for image retrieval, and usually will not be updated in real time;
[0102] (b2) In the online retrieval state, the query image is input into the image hash encoding model to obtain the query image hash code;
[0103] (b3) Calculate the Hamming distance between the query image hash code and each offline image hash code in the image set hash code and sort them in ascending order. Take the first K offline images corresponding to the Hamming distances as the similar image list of the query image, where K is an integer.
[0104] In (b1) and (b2), both the image set hash code and the query image hash code are in the binary hash code structure, specifically using the sign function Convert the output of the hash encoder into a binary hash code of -1 and 1. By using the binary hash code, the storage overhead and distance calculation overhead of the system can be effectively reduced, and the efficiency of the image retrieval system can be improved.
[0105] This embodiment proposes a deep image retrieval system based on metric learning. The query image is input into the image hash coding model for image retrieval to obtain a list of similar images;
[0106] The training process of the image hash coding model includes a training set construction module, a feature projection module, a hash center projection module, and a total loss module;
[0107] The training set construction module is used to construct an image training set and a hash center, and each time a batch of images is sampled from the training set and input into the image hash coding model;
[0108] The feature projection module is used to extract features from the image and project them into the embedding space to obtain feature vectors;
[0109] The hash center projection module is used to project the hash center into the embedding space to obtain a proxy;
[0110] The total loss module is used to optimize the similarity between the feature vector and the proxy based on metric learning in the embedding space, so that the feature vector converges to the same-class proxy and the proxy adjusts to the same-class feature vector; at the same time, a reciprocal constraint is imposed between the hash encoder and decoder, and a metric invariance constraint is imposed on the hash coding process, so that the hash code converges to the corresponding hash center.
[0111] In addition, a computer-readable storage medium is proposed. A number of classification programs are stored on the computer-readable storage medium, and the number of classification programs is used to be called by a processor and execute the above-mentioned deep image retrieval method.
[0112] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.
[0113] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A deep image retrieval method based on metric learning, characterized in that: Input the query image into the image hash coding model for image retrieval to obtain a list of similar images; The training process of the image hash coding model is as follows: Construct an image training set and hash center, and sample batches of images from the training set each time and input them into the image hash coding model; Extract features from batch images and project them into the embedding space to obtain feature vectors, and project the hash centers into the embedding space to obtain proxies; In the embedding space, the similarity between feature vectors and agents and between feature vectors is optimized based on metric learning, so that feature vectors converge to similar agents and agents adjust to similar feature vectors. Reciprocity constraints are imposed on the hash encoder and decoder, and metric invariance constraints are imposed on the hash encoding process, so that the hash codes converge to the corresponding hash centers.
2. The deep image retrieval method based on metric learning according to claim 1, characterized in that: In optimizing the similarity between feature vectors and agents and between feature vectors based on metric learning, an optimization strategy based on agent-based metric learning is used to optimize the similarity between feature vectors and agents, specifically: Optimizing the cosine similarity between the feature vector and similar agents until it exceeds a set first similarity threshold, and the optimization intensity of the cosine similarity increases as the cosine similarity decreases; The cosine similarity between the feature vector and the agents of different classes is optimized until it is lower than a set first different class similarity threshold, and the optimization intensity of the cosine similarity increases as the cosine similarity increases.
3. The deep image retrieval method based on metric learning according to claim 1, characterized in that: In optimizing the similarity between feature vectors and agents and between feature vectors based on metric learning, an optimization strategy based on metric learning of paired samples is used to optimize the similarity between feature vectors, specifically: Optimize the cosine similarity between feature vectors of samples of the same category until it exceeds the set second similarity threshold of the same category, and the optimization intensity of the cosine similarity increases as the cosine similarity decreases; The cosine similarity between the feature vectors of samples of different categories is optimized until it is lower than the set second different-category similarity threshold, and the optimization intensity of the cosine similarity increases with the increase of the cosine similarity.
4. The deep image retrieval method based on metric learning according to claim 1, characterized in that: During the model training phase, the feature vector is modeled with a Gaussian distribution in the embedding space. Each time the optimization is performed, the feature vector is sampled from the Gaussian distribution, and then the similarity between the sampled feature vector and the agent is optimized based on metric learning.
5. The deep image retrieval method based on metric learning according to claim 1, characterized in that: The hash codec includes a hash encoder and a hash decoder, and the encoding process of the hash encoder and the decoding process of the hash decoder are constrained to be reciprocal.
6. The deep image retrieval method based on metric learning according to claim 1, characterized in that: The constrained hash coding process keeps the cosine similarity between the two hash codes after encoding and the cosine similarity between the two feature vectors before encoding unchanged.
7. The deep image retrieval method based on metric learning according to claim 1, characterized in that: Construct a total loss function to train the image hash coding model. The total loss function includes the hash center loss function. , metric learning loss function in embedding space And the hash encoder loss function ; Hash Center Loss Function Used to increase the Hamming distance between hash centers of different categories in the case of parameterized hash centers to meet the requirements of hash coding; Metric Learning Loss Function in Embedding Space It is used to make the feature vectors converge to similar agents in the embedding space, and the agents adjust to similar feature vectors; Hash Encoder Loss Function It is used to ensure that the encoding process is inverse to the decoding process and that the encoding process is metrically invariant.
8. The deep image retrieval method based on metric learning according to claim 1, characterized in that: During the execution of the image hash coding model, specifically: In an offline state, the image hash coding model takes an image set as input to obtain an image set hash code, wherein the image set is a collection of images that are pre-collected, stored, and used for image retrieval; In the online retrieval state, the query image is input into the image hash coding model to obtain the query image hash code; Calculate the Hamming distance between the query image hash code and each offline image hash code in the image set hash code and arrange them in ascending order. Take the offline images corresponding to the first K Hamming distances as the similar image list of the query image, where K is an integer.
9. A deep image retrieval system based on metric learning, characterized in that: Input the query image into the image hash coding model for image retrieval to obtain a list of similar images; The training process of the image hash coding model includes a training set construction module, a feature projection module, a hash center projection module and a total loss module; The training set construction module is used to construct the image training set and hash center. Each time, a batch of images are sampled from the training set and input into the image hash coding model. The feature projection module is used to extract features from batch images and project them into the embedding space to obtain feature vectors; The hash center projection module is used to project the hash center into the embedding space to obtain the proxy; The total loss module is used to optimize the similarity between feature vectors and agents in the embedding space based on metric learning, so that feature vectors converge to similar agents and agents adjust to similar feature vectors; at the same time, reciprocity constraints are imposed between hash encoders and decoders, and metric invariance constraints are imposed on the hash encoding process, so that the hash code converges to the corresponding hash center.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of classification programs, and the plurality of classification programs are used to be called by a processor and execute the depth image retrieval method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Pedestrian feature Hash coding method based on deep Hash algorithm
CN118038194A
Depth metric learning image retrieval system based on proxy correction and relation constraint
CN118397421A