Double-path unsupervised Hash retrieval method and system based on mutual information maximization
The generation of high-quality hash codes through the dual-encoding network and triple mutual information optimization modules solves the problem of insufficient hash code representation ability in the unsupervised hash method, and improves the accuracy and efficiency of image retrieval.
Patent Information
- Application Number
- CN202510335612.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The existing unsupervised hashing methods have problems in image retrieval of insufficient hash code representation capabilities and poor retrieval performance, especially in unsupervised learning scenarios, which makes it difficult for hash code to accurately distinguish different categories of data.
A dual-path unsupervised hash retrieval method based on mutual information maximization is adopted to generate approximate hash codes and continuous latent variables through a dual-encoded network, and combined with a triple mutual information optimization module and a graph convolution network, the separability and semantic consistency of hash codes are optimized, and a high-quality hash code is generated using the decoder reconstruction.
The characterization and retrieval performance of hash codes is significantly improved, and the generated hash codes have stronger discrimination, compactness and generalization capabilities, improving the accuracy and efficiency of image retrieval.
Smart Images

Figure CN120256659A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of image retrieval, and more particularly, to a dual-path unsupervised hashing retrieval method and system based on maximizing mutual information. Background Art
[0002] The statements in this section merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] With the rapid development of Internet technology and intelligent devices, multimedia data has shown an explosive growth, leading to huge challenges for traditional content-based multimedia retrieval. Especially in the processing of high-dimensional data, traditional approximate nearest neighbor search (ANN) methods are difficult to meet the requirements of efficient retrieval in terms of computational cost and storage requirements. Therefore, hashing methods have become an important means to solve the problem of rapid retrieval of high-dimensional data due to their advantages of fast retrieval speed and low storage cost. Hashing methods convert the original high-dimensional data into low-dimensional binary codes through hashing functions, enabling efficient search in large-scale datasets, significantly improving computational efficiency and reducing memory occupancy. In the development of hashing methods, supervised hashing methods rely on labeled data to train hashing models. Although they can enhance the representational ability of hash codes, the high cost of manual annotation and the difficulty in obtaining high-quality large-scale data limit their popularization in practical applications.
[0004] Unsupervised hashing methods provide a more cost-effective solution, but the existing technologies still face problems such as insufficient representational ability of hash codes and poor performance in image retrieval. In recent years, researchers have tried to use deep unsupervised hashing methods to improve the quality of hash codes. For example, the autoencoder (AE) structure learns an effective representation of data by minimizing the reconstruction error between the input and output. However, since its objective is mainly data reconstruction, it is difficult to ensure the discriminability and separability of hash codes. To further optimize the learning process of hash codes, variational autoencoders (VAEs) have been introduced into unsupervised hashing methods, using their bottleneck layers to directly construct hash codes, thereby enhancing the structure and interpretability of hash codes. However, existing VAE variational methods still suffer from insufficient semantic expression ability of hash codes. In unsupervised learning scenarios, it is difficult to fully capture the underlying patterns of data, resulting in hash codes being unable to accurately distinguish different categories of data, ultimately affecting the accuracy and efficiency of image retrieval. Summary of the Invention
[0005] To solve the above problems, the present disclosure proposes a dual-path unsupervised hashing retrieval method and system based on mutual information maximization, and a deep learning optimization framework of triple mutual information based on the dual-path variational idea. This framework optimizes using the structural and semantic interrelationships among features, approximate hash codes, and continuous latent variables, learns the semantic content of the original images, thereby improving the representational ability of the hash codes and further enhancing the image retrieval performance.
[0006] To achieve the above objectives, the present disclosure adopts the following technical solutions:
[0007] One or more embodiments provide a dual-path unsupervised hashing retrieval method based on mutual information maximization, including the following steps:
[0008] Extract image features for the acquired images;
[0009] Use a dual-encoding network to learn the image features to obtain approximate hash codes and continuous latent variables;
[0010] The dual-encoding network is connected with a triple mutual information optimization module using a discriminator. During the training process, the triple mutual information optimization method is used to optimize the approximate hash codes, continuous latent variables, and image features, and the output of the dual-encoding network is optimized through discriminator adversarial learning;
[0011] Use a graph convolutional network to process the obtained approximate hash codes and continuous latent variables to obtain the final latent variables;
[0012] Input the latent variables into a decoder for reconstruction, and finally obtain the hash codes of the data to be retrieved, and obtain the retrieval results based on the hash codes.
[0013] One or more embodiments provide a dual-path unsupervised hashing retrieval system based on mutual information maximization, including:
[0014] An image feature extraction module configured to extract image features for the acquired images;
[0015] A dual-encoding module configured to use a dual-encoding network to learn the image features to obtain approximate hash codes and continuous latent variables;
[0016] The dual-encoding network is connected with a triple mutual information optimization module using a discriminator. During the training process, the triple mutual information optimization method is used to optimize the approximate hash codes, continuous latent variables, and image features, and the output of the dual-encoding network is optimized through discriminator adversarial learning;
[0017] A graph convolutional module configured to use a graph convolutional network to process the obtained approximate hash codes and continuous latent variables to obtain the final latent variables;
[0018] The reconstruction retrieval module is configured to input the latent variables into the decoder for reconstruction, and finally obtain the hash code of the data to be retrieved, and obtain the retrieval result based on the hash code.
[0019] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the above-mentioned dual-path unsupervised hashing retrieval method based on mutual information maximization are completed.
[0020] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the steps in the above-mentioned dual-path unsupervised hashing retrieval method based on mutual information maximization are completed.
[0021] Compared with the prior art, the beneficial effects of the present disclosure are as follows:
[0022] Through the dual-encoding network, triple mutual information optimization, Graph Convolutional Networks (GCN) processing, and decoder reconstruction, the present disclosure significantly improves the representational ability and retrieval performance of the hash code. First, image features are extracted and an approximate hash code and continuous latent variables are generated using a dual-encoding network, and are optimized and trained through a triple mutual information optimization module to enhance the separability and semantic consistency of the hash code. Subsequently, GCN is used to further optimize the hash code and latent variables, improving the structured expression ability of the hash code in the data distribution, making the hash codes of similar data points closer. Finally, the optimized latent variables are input into the decoder for reconstruction. This method can still generate high-quality hash codes under unsupervised conditions, with stronger discriminability, compactness, and generalization ability, significantly improving the accuracy and efficiency of image retrieval.
[0023] The advantages of the present disclosure and the advantages of additional aspects will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments and descriptions thereof of the present disclosure are used to explain the present disclosure and do not constitute a limitation to the present disclosure.
[0025] Figure 1 It is a block diagram of a network model for realizing unsupervised hashing retrieval in Embodiment 1 of the present disclosure;
[0026] Figure 2 It is a flowchart of the unsupervised hashing retrieval method in Embodiment 1 of the present disclosure; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0028] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.
[0029] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. It should be noted that, without conflict, the various embodiments and features in the present disclosure can be combined with each other. The embodiments will be described in detail below with reference to the drawings.
[0030] Embodiment 1
[0031] In the technical solutions disclosed in one or more embodiments, as Figures 1 to 2 shown, a dual-path unsupervised hashing retrieval method based on mutual information maximization includes the following steps:
[0032] Step 1: Extract image features for the acquired images;
[0033] Step 2: Use a dual-encoding network to learn the image features to obtain an approximate hash code b i and a continuous latent variable c i ;
[0034] The dual-encoding network is connected with a triple mutual information optimization module using a discriminator, and during the training process, a triple mutual information optimization method is used to optimize the approximate hash code b i , the continuous latent variable c i and the image feature z i , and the output of the dual-encoding network is optimized through discriminator adversarial learning;
[0035] Step 3: Use a graph convolutional network to process the obtained approximate hash code b i and the continuous latent variable c i to obtain a final latent variable z' i ;
[0036] Step 4: Input the latent variable z' i into the decoder for reconstruction, and finally obtain the hash code of the data to be retrieved, and obtain the retrieval result based on the hash code.
[0037] In this embodiment, it is set to extract the hash code and the continuous latent variable through two paths respectively, process them through a graph convolutional network to obtain the final latent variable, and then input it into the decoder network for decoding; optimize and refine the approximate hash code, continuous latent variable and image features through triple mutual information optimization, so that their structures and distributions meet the requirements of the hash function; thus, the accuracy of retrieval can be improved.
[0038] The dual-encoding network adopts a two-path (hash code + continuous latent variable) to achieve structural optimization, making the hash code have stronger expressive ability and generalization ability. The two-path strategy reduces the loss of binary information, improves the discriminability and compactness of the hash code, and improves the retrieval accuracy. Combining the triple mutual information maximization strategy, adversarial training, and GCN processing keeps the semantic information of the hash code consistent and improves the image retrieval performance. The hash method designed in this way can still generate high-quality hash codes and effectively improve the image retrieval performance in the case of unsupervised learning.
[0039] In step 1, the image is represented as x i and the corresponding image feature is represented as z i , where i represents the i-th sample;
[0040] In step 1, optionally, a convolutional neural network can be used to extract image features; specifically, a VGG-16 network can be used.
[0041] In step 2, the image feature z i is fed into the dual-encoding network, and an approximate hash code b i and a continuous latent variable c i are obtained through the dual-encoding network. The specific formula is as follows:
[0042]
[0043] where b i and c i represent the approximate hash code and the continuous latent variable respectively, f1(·) and f2(·) represent activation functions, and are normal distributions composed of the mean μ1 and μ2, variances and respectively. ω1 and ω2 are network parameters, K represents the approximate length of the hash code, and L represents the length of the continuous latent variable.
[0044] In this embodiment, a dual - coding network is set up to extract hash codes and continuous latent variables respectively through two paths. First, the representational ability of the hash code is enhanced. Traditional hash methods usually directly learn binary hash codes from data, while this method additionally introduces continuous latent variables to provide richer semantic information, enabling the hash code to more accurately represent the structure and semantics of the input data. The hash code is used for retrieval, and the continuous latent variable serves as an auxiliary variable to optimize the expressive ability of the hash code, ensuring that the hash code can retain more high - dimensional feature information. Second, the binary information loss is solved to improve the discriminability of the hash code. Traditional binary hash codes are limited by a fixed length and may lose some information. The continuous latent variable can capture more fine - grained feature information through Gaussian distribution modeling, compensating for the information loss of the hash code and improving the separability of the data. In the case of unsupervised learning, due to the lack of label supervision, the distribution of the hash code may be too discrete, while the continuous latent variable provides a smooth representation, making the hash code more coherent and improving the discriminative ability of the hash code.
[0045] In a further solution, the dual - coding network includes two parallel variational auto - encoders (VAEs), forming a deep - learning architecture, which are respectively used to learn hash codes and continuous latent variables; the VAE provides a probabilistic modeling method, making the hash code distribution smoother, and at the same time, through KL - divergence optimization, improving the discriminative ability and generalization of the hash code.
[0046] This embodiment sets up a dual - coding network and trains two VAEs simultaneously, and optimizes through the mutual - information maximization strategy, thereby enhancing the expressive ability of the hash code and ultimately improving the performance of unsupervised hash retrieval.
[0047] To optimize the training objective of the variational auto - encoder, the loss function in the training process of the variational auto - encoder includes a Kullback - Leibler (KL) divergence term and a reconstruction error term. The specific loss function is:
[0048]
[0049] Among them, is the prior distribution, which follows an isotropic multivariate Gaussian distribution, and θ is the model distribution parameter calculated from the hash code b i ; p θ (z i |b i ) is a multivariate Gaussian or Bernoulli generative model;
[0050] Since the prior distribution lacks parameters, the objective function constructed in this embodiment transforms the posterior distribution p θ (z i |b i ), which is difficult to estimate, into a tractable variational distribution q φ (b i|z i ), where φ is a variational parameter, and the variational distribution q is used with KL divergence φ (b i |z i ) is used to replace the posterior distribution p θ (z i |b i ). In this way, while minimizing the reconstruction error, the KL divergence between the distribution of the latent space and the prior distribution can be minimized. At the same time, the mapping from the hash code to the continuous latent variable is modeled by a Gaussian and set to a standard normal distribution to ensure good continuity and separability when generating the hash code.
[0051] The dual-encoding network of this embodiment combines the triple mutual information (TMI) optimization strategy to maximize the mutual information among the image features, hash code, and continuous latent variable, ensuring that the generated hash code has good discriminative ability while maintaining the consistency of the data structure. The encoder is optimized through an adversarial training strategy to make the hash code maintain good discriminability while being compact, improving the accuracy of hash retrieval.
[0052] In some embodiments, the triple mutual information optimization method is specifically implemented by constructing a triple mutual information optimization module, including a first optimization branch, a second optimization branch, and a third optimization branch:
[0053] First optimization branch: used for matching optimization of the hash code b i and the image feature z i , including a first fully connected layer and a first discriminator connected in sequence. The input of the first fully connected layer is respectively connected to the output of the convolutional neural network and the dual-encoding network;
[0054] The first optimization branch constructs the objective function of the first optimization branch through the JS divergence term and is trained by the first discriminator to calculate the JS divergence term between the hash code and the image feature, obtaining the mutual information between the hash code b i and the image feature z i as the first mutual information, maximizing the correlation between the hash code b i and the image feature z i so that the hash code can better represent the original image features. Maximize the mutual information between the hash code and the image feature space to ensure that the hash code can retain as much semantic information of the original image features as possible while enhancing separability.
[0055] The objective function L of the first optimization branch MI includes two terms:
[0056] The first item (Jensen-Shannon divergence, JS): Maximize the mutual information between the hash code and the image features;
[0057] The second item (KL divergence): Minimize the deviation between the hash code distribution and the standard normal distribution to ensure the separability and compactness of the hash code;
[0058] Among them, the hyperparameters α and β are used to control the influence of maximizing mutual information and KL divergence regularization.
[0059] Specifically, the objective function of the first optimization branch is as follows:
[0060]
[0061] Among them, the calculation formula of the JS divergence is:
[0062]
[0063] Among them, a discriminant network τ1(T(z i ,b i )) is introduced to measure and maximize the JS divergence, thereby enhancing the mutual information between the encoder input z i and the output b i . q φ (z i ) represents the distribution of the original image features. z i and the corresponding b i , that is, the matching image feature and hash code pair (z i ,b i ) are regarded as positive samples; z i and the randomly sampled that is is the hash code randomly sampled from b i , then it is regarded as a negative sample.
[0064] The second optimization branch: It includes a second fully connected layer and a second discriminator connected in sequence. The input of the first fully connected layer is connected to the output of the double coding network, which is used to construct positive and negative sample pairs of the hash code b i , maximize the JS divergence of different hash codes to make similar hash codes closer and different hash codes farther away, improve the discriminability of the hash code, and serve as the second mutual information optimization; through optimization, the noise influence caused by upsampling can be reduced;
[0065] In the second optimization branch, by randomly shuffling the approximate hash code b i of the input to construct positive and negative sample pairs; the positive and negative sample pairs are trained through the discriminator. Specifically, (b i ,b i ) is used as a positive sample, while As negative samples, the constructed objective function includes:
[0066] The first term: JS divergence (Jensen-Shannon divergence);
[0067] The purpose of this term is to maximize the mutual information between the hash code and the image features to ensure that the hash code can effectively represent the feature information of the original data.
[0068] The second term: KL divergence (Kullback-Leibler divergence). The purpose of this term is to constrain the distribution of the hash code to be close to the standard normal distribution to ensure the compactness and generalization ability of the hash code.
[0069] Specifically, the objective function of the second optimization branch is as follows:
[0070]
[0071] Among them,
[0072]
[0073] Different from the traditional discriminator and generator that follow the optimization path of zero-sum game, in this embodiment, the optimization directions of the discriminator and the dual-path encoder are aligned to maximize the JS divergence. This strategy enables the sharing of structures and collaborative optimization between the encoder and the discriminator, enhancing the representation ability of the feature structure. Therefore, (b i , b i ) is used as positive samples, while is used as negative samples and fed into the discriminator for training. Among them, τ2(T(b i , b i )) is a discriminator. q φ (b i ) = ∫q φ (b i |z i )q φ (z i )dz represents the distribution of the entire B after given q φ (b i |z i ), while q φ (b i ) is the standard normal distribution. p θ (b i ) is the prior distribution, representing an isotropic multivariate Gaussian distribution. The approximate hash code b i follows this prior distribution, manifested as the standard normal distribution. This alignment method promotes a unified coding space and helps to decouple features, thereby improving the efficiency of the subsequent learning process.
[0074] The third optimization branch: It includes a third fully connected layer and a second discriminator connected in sequence. The input of the third fully connected layer is connected to the output of the dual encoding network, which is used to construct positive and negative sample pairs of the continuous latent variable c i to maximize the Jensen-Shannon (JS) divergence of different continuous latent variables, making similar continuous latent variables closer and different continuous latent variables farther apart, improving the discriminability of continuous latent variables, which is used as the third mutual information optimization; through optimization, the noise impact caused by upsampling can be reduced;
[0075] In the third optimization branch, by randomly shuffling the approximate hash code b i of the input, positive and negative sample pairs are constructed; the positive and negative sample pairs are trained through the discriminator. Specifically, (b i , b i ) is used as the positive sample, while is used as the negative sample. The constructed objective function includes:
[0076] The first term: JS divergence (Jensen-Shannon divergence)
[0077] The purpose of this term is to maximize the mutual information between the continuous latent variable and the image feature to ensure that the hash code can ensure that the latent variable retains sufficient semantic information, thereby assisting the hash code to generate better feature representations.
[0078] The second term: KL divergence (Kullback-Leibler divergence). The purpose of this term is to constrain the distribution of the continuous latent variable to be close to the standard normal distribution to ensure the compactness and generalization ability of the hash code.
[0079] Two hyperparameters α and β are used to balance the influence between mutual information optimization (maximizing JS divergence) and distribution constraint (minimizing KL divergence).
[0080] Next, a similar method is adopted for the continuous latent variable. (c i , c i ) is used as the positive sample, is used as the negative sample, as shown below:
[0081]
[0082] Among them,
[0083]
[0084] Among them, τ3(T(c i , c i )) is a discriminant network. q φ (c i ) = ∫q φ (c i |z i)q φ (z i )dz represents the distribution of the entire C after a given q φ (c i |z i ) and q φ (c i ) is the standard normal distribution. p θ (c i ) is the prior distribution, and the continuous latent variable c i follows the prior distribution of the standard normal distribution.
[0085] This embodiment adopts a triple mutual information optimization strategy to jointly optimize and train the approximate hash code b i , the continuous latent variable c i and the image feature z i Through adversarial training, the distribution of the hash code is constrained to force its bits to be independent and uniformly distributed, maximizing the information entropy. Maximizing the mutual information preserves the semantic information of the original features, causing similar samples to cluster in the latent space. A deep discriminator network is constructed through fully connected layers, and hierarchical non-linear feature transformation is achieved by combining ReLU activation functions to maximize the mutual information entropy between the latent variable space and the feature space, ensuring that the hash code distribution is aligned with the target prior distribution. Specifically, a deep discriminator network is constructed by stacking fully connected layers, with a ReLU activation function connected after each layer. For example: FC1 → ReLU → FC2 → ReLU →... → FCn. Introducing non-linear expression ability enables the network to learn complex feature relationships and form hierarchical feature abstractions.
[0086] At the same time, an adversarial training strategy is adopted to synchronously optimize the parameters of the encoder and the discriminator, improving the compactness and discriminability of the hash code. Finally, this method ensures that the hash code is both compact and has good semantic expression ability, while avoiding the information loss problem in unsupervised learning and improving the accuracy and generalization ability of hash retrieval.
[0087] In unsupervised deep learning methods, learning representative features is a key task, where the encoder is trained to convert the original image into a fixed-length vector. Its main goal is to capture the most distinguishable information in the original image, similar to the goal of hashing. For the deep feature set and the approximate hash code set, mutual information is used as a metric to quantify the uniqueness of the extracted features, and the encoder emphasizes that this mutual information should be maximized. In this embodiment, three mutual information optimizations are performed, and for the deep features, it is expressed as τ1(T(z i ,b i))). In actual operation, in order to reduce the noise impact caused by upsampling in the feature optimization module, in this embodiment, a network architecture method similar to the deep feature optimization module is first adopted, and positive and negative sample pairs are respectively constructed by randomly shuffling the approximate hash code and the continuous latent variable. Therefore, for the hash code, it is represented as τ2(T(b i ,b i ))), and for the continuous latent variable, it is represented as τ3(T(c i ,c i ))).
[0088] Specifically, during the training process, the weighted mixing strategy combines the objectives of the variational autoencoder and the triple mutual information. Specifically, in the dual-encoding network structure, minimizing the KL divergence makes the posterior distribution q φ (b i |z i ) consistent with the standard normal distribution pθ(b i ). This alignment ensures the smooth distribution of the latent variables, thereby enhancing generalization. In addition, maximizing the triple mutual information optimizes the features, approximate hash code, and continuous latent variables respectively, increases the distribution separation using the JS divergence, and minimizes the objective functions L MI 、L B and L C , thereby refining the representation learning process. To sum up, the overall objective function is:
[0089] L = γ·L V +(1 - γ)(L MI +L B +L C )
[0090] where γ is a hyperparameter. After the last round is completed, for the data to be retrieved, the hash code to be retrieved is generated based on the final model parameters.
[0091] Furthermore, the generated hash code and the continuous latent variable c i are processed using a graph convolutional network to obtain the final latent variable z′ i . The latent variable z′ i is input into the decoder for reconstruction, and finally the recovered features, that is, the hash code of the data to be retrieved, are obtained. The graph convolutional network GCN optimizes the node representation using the neighbor information through the message passing mechanism, making the hash codes of similar data points closer to the latent variables
[0092] In this embodiment, the hash codes and continuous latent variables learned by the dual-path are further processed by the graph convolutional network (GCN), making the final hash codes more conform to the global structure of the data and enhancing the stability of retrieval. After being processed by the GCN, the hash codes and continuous latent variables are input into the decoder for feature reconstruction, which ensures the representational ability of the hash codes and latent variables, avoids model overfitting, and improves the retrieval performance of the model.
[0093] In step 4, the retrieval result is obtained based on the finally obtained hash codes of the data to be retrieved, which specifically includes:
[0094] Step 41: Calculate the Hamming distance between the hash code to be retrieved and the database hash codes, and sort them.
[0095] Specifically, perform an exclusive OR calculation on the hash code to be retrieved and the stored hash codes to obtain the Hamming distance of the hash codes.
[0096] Step 42: Return the corresponding original data according to the hash code with the smallest Hamming distance to achieve similarity retrieval.
[0097] Sort the hash codes according to the Hamming distance, obtain several hash codes with the smallest distance, and return several pieces of original data according to the position information of the original data corresponding to the hash codes.
[0098] In the above solution of this embodiment, the unsupervised deep hashing method based on mutual information maximization realizes efficient image similarity retrieval under the condition of no labels, and improves the accuracy and robustness of hashing retrieval. A new triple information maximization optimization strategy is proposed to enhance the deep extraction of image features, making the generated hash codes more representative and adaptive, and improving the discriminability and semantic consistency of the hash codes. The dual-path variational autoencoder (VAE) is used to jointly model the hash codes and continuous latent variables, effectively improving the representational ability and discrimination ability of the hash codes, and ensuring that the hash codes are both compact and can express rich semantic information. A new loss function is designed, combining the Kullback-Leibler (KL) divergence and the Jensen-Shannon (JS) divergence. By optimizing the gradient stability, it ensures more stable model training and improves the convergence speed and distribution rationality of the hash codes. By optimizing the network structure, improving the loss function, enhancing mutual information learning, etc., the expression ability, retrieval performance and model stability of the hash codes are improved under the unsupervised condition, significantly enhancing the accuracy and efficiency of image retrieval.
[0099] To illustrate the effect of this embodiment, a simulation experiment was conducted.
[0100] Obtain images from three datasets and resize the size of the images to 224×224×3. The training batch size is 400, and the length of the continuous latent variable L is set to 512. Adam optimizer is used for optimization with a learning rate of 1e-6. The hyperparameters are set as follows: α = 0.5, β = 0.1, γ = 0.7.
[0101] The experimental results are as follows:
[0102] Table 1 shows the comparison of the mean average precision (MAP) of the method of this embodiment (ours) with other existing technologies. In this embodiment, MAP@1000 is used on the CIFAR-10 dataset, MAP@5000 is used on the NUS-WIDE dataset, and MAP@5000 is used on the MS COCO dataset.
[0103] Table 1 Comparison results;
[0104]
[0105]
[0106] CIFAR-10, NUS-WIDE, MSCOCO: Standard datasets for evaluating hashing methods, all of which are publicly available datasets containing different types of images; The CIFAR-10 dataset consists of 60,000 color images in 10 categories, with 6,000 images in each category. The NUS-WIDE dataset consists of 269,648 color images in 81 categories, with varying resolution sizes, and it is a multi-label dataset. The MS COCO dataset is a large-scale multi-label dataset containing 122,218 color images in 80 categories selected.
[0107] In Table 1, the explanations of the existing technologies are as follows:
[0108] DeepBit (Learning compact binary descriptors with unsupervised deep neural networks): A method for unsupervised learning of compact binary descriptors for efficient visual object matching. This method does not rely on manually labeled data but generates binary codes through self-learning. During training, by optimizing three criteria: minimizing quantization loss, ensuring uniform distribution of the codes, and removing the correlation between bits, efficient and discriminative binary descriptors are generated.
[0109] SGH (Stochastic generative hashing): A generative model that combines VAE and generative adversarial networks. It enhances privacy protection by generating synthetic data to resist membership inference attacks. This method encodes discrete variables through a conditional generator, uses sampling training to ensure that the distribution of synthetic data is consistent with that of real data, and ensures training stability by improving the loss function and introducing the optimal transport distance and gradient penalty methods.
[0110] BGAN (Binary generative adversarial networks for image retrieval): An unsupervised image embedding method that directly generates binary codes through generative adversarial networks. It innovatively introduces a sign activation strategy to avoid relaxation operations and designs a composite loss function including adversarial loss, content loss, and neighborhood structure loss to improve the accuracy of image retrieval.
[0111] BinGAN (Bingan: Learning compact binary descriptors with a regularized gan): A method for generating compact binary image descriptors through generative adversarial networks. This method uses the penultimate layer of the generative adversarial network discriminator for dimensional compression and binarizes the output to obtain a low-dimensional binary representation. These binary descriptors are both compact and highly discriminative, suitable for image matching and retrieval tasks. Two loss functions are introduced, one to maximize the entropy of the binary representation to reduce the correlation between dimensions, and the other to ensure that the relationship between the low-dimensional representation and the high-dimensional features is maintained.
[0112] GreedyHash (Greedy hash: Towards fast optimization for accurate hash coding in cnn): A deep hashing method based on a greedy strategy, aiming to solve the discretization problem in deep hashing. By gradually updating the network, this method can approach the optimal discrete solution in each iteration, thus avoiding the complex optimization problems in traditional methods. This method designs a hash coding layer to ensure the maintenance of discrete constraints in forward propagation and avoid gradient vanishing in backward propagation.
[0113] HashGAN (Unsupervised deep generative adversarial hashing network): A deep unsupervised hashing method based on generative adversarial networks. Through the collaboration of a generator, a discriminator, and an encoder, combined with adversarial loss and a novel loss function, it avoids the limitations of traditional unsupervised hashing methods. This method can achieve excellent image retrieval and clustering performance without supervised pre-training. By minimizing entropy, ensuring the consistency and independence of hash bits, and combining triplet ranking loss, the quality of hash codes is significantly improved.
[0114] DVB (Unsupervised binary representation learning with deep variational networks): This method overcomes the problem that traditional projection methods cannot effectively reveal the structure of the sample space by introducing a conditional auto-encoding variational Bayesian network and using latent variables to reveal the feature space structure of the training data. Combining probabilistic inference with a hashing objective, it generates compact binary hash codes, thus improving the hashing performance in big data applications such as image retrieval.
[0115] TBH (Auto-encoding twin-bottleneck hashing): An unsupervised hashing method based on autoencoders. By introducing a code-driven graph and a twin-bottleneck structure, it optimizes graph update and hashing function learning to overcome the limitations of graph construction and hashing function learning in traditional unsupervised hashing methods. By exchanging key information through two bottlenecks, one for transmitting high-level data structures and the other for transmitting low-level detail information, the network is optimized and the image retrieval performance is improved.
[0116] DUSH (Deep unsupervised self-evolutionary hashing for image retrieval): This method solves the problems of pseudo-label calculation complexity and low accuracy in traditional unsupervised hashing methods by iteratively selecting pseudo-pairs from simple to difficult in the low-dimensional Hamming space through a curriculum learning strategy. Moreover, this method does not rely on manual labels and can effectively reduce the burden of annotating large-scale datasets.
[0117] CIBHash (Unsupervised hashing with contrastive information bottleneck): An unsupervised hashing method based on contrastive learning. By modifying the objective function and introducing a probabilistic binary representation layer, it solves the problem that traditional hashing methods overemphasize background information and neglect important semantic information. This method combines contrastive learning with the information bottleneck framework, optimizes the generation of binary hash codes, and improves the discrimination ability of the hashing task.
[0118] MeCoQ (Contrastive quantization with code memory for unsupervised image retrieval): An innovative method for unsupervised deep quantization, aiming to address the limitations of traditional reconstruction strategies. It learns binary descriptors through contrastive learning to better capture visual semantics and avoid the limitations of traditional reconstruction strategies. It introduces codeword diversity regularization to prevent model degradation and proposes a new quantization code memory module to reduce feature drift and improve the contrastive learning effect.
[0119] The results in Table 1 show that the method of this embodiment is superior to existing image retrieval methods on all the datasets used. The present invention effectively utilizes the semantic relationship between images and hash codes, thereby optimizing the generation of representative and discriminative hash codes. Compared with the MeCoQ method, on the CIFAR-10 and NUS-WIDE datasets with different code lengths, the average MAP is increased by 1.0% and 0.7% respectively.
[0120] Embodiment 2
[0121] Based on Embodiment 1, a dual-path unsupervised hashing retrieval system based on mutual information maximization is provided in this embodiment, including:
[0122] An image feature extraction module, configured to extract image features for the acquired images;
[0123] A dual-encoding module, configured to learn the image features using a dual-encoding network to obtain approximate hash codes and continuous latent variables;
[0124] The dual-encoding network is connected with a triple mutual information optimization module using a discriminator, and in the training process, a triple mutual information optimization method is used to optimize the approximate hash codes, continuous latent variables, and image features, and the output of the dual-encoding network is optimized through discriminator adversarial learning;
[0125] A graph convolution module, configured to process the obtained approximate hash codes and continuous latent variables using a graph convolutional network to obtain final latent variables;
[0126] The reconstruction retrieval module is configured to input latent variables into the decoder for reconstruction, and finally obtain the hash code of the data to be retrieved, and obtain the retrieval result based on the hash code.
[0127] It should be noted here that each module in this embodiment corresponds to each step in Embodiment 1 one by one, and its specific implementation process is the same, so it will not be repeated here.
[0128] Embodiment 3
[0129] This embodiment provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the mutual information maximization-based dual-path unsupervised hashing retrieval method of Embodiment 1 are completed.
[0130] Embodiment 4
[0131] This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps in the mutual information maximization-based dual-path unsupervised hashing retrieval method of Embodiment 1 are completed.
[0132] The above are only the preferred embodiments of the present disclosure and are not used to limit the present disclosure. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
[0133] Although the specific implementation manners of the present disclosure are described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that based on the technical solutions of the present disclosure, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present disclosure.
Claims
1. A dual-path unsupervised hashing retrieval method based on maximizing mutual information, characterized in that, It includes the following steps: Extract image features for the acquired image; Use a dual-encoding network to learn the image features to obtain an approximate hash code and a continuous latent variable; The dual-encoding network is connected with a triple mutual information optimization module using a discriminator. During the training process, the triple mutual information optimization method is used to optimize the approximate hash code, continuous latent variable, and image features, and the output of the dual-encoding network is optimized through discriminator adversarial learning; Use a graph convolutional network to process the obtained approximate hash code and continuous latent variable to obtain the final latent variable; Input the latent variable into the decoder for reconstruction, and finally obtain the hash code of the data to be retrieved. Based on the hash code, obtain the retrieval result.
2. The unsupervised hashing retrieval method based on mutual information maximization as claimed in claim 1, wherein: The dual-encoding network includes two parallel variational autoencoders, constituting a deep learning architecture, which are respectively used to learn the hash code and continuous latent variable.
3. The dual-path unsupervised hashing retrieval method based on mutual information maximization according to claim 1, characterized in that: The loss function in the training process of the variational autoencoder includes a KL divergence term and a reconstruction error term.
4. The method for dual-path unsupervised hashing retrieval based on mutual information maximization according to claim 1, wherein: The triple mutual information optimization module includes a first optimization branch, a second optimization branch, and a third optimization branch: The first optimization branch: used to optimize the matching of the hash code and image features, including a first fully connected layer and a first discriminator connected in sequence. The input of the first fully connected layer is respectively connected to the convolutional neural network and the output of the dual-encoding network; calculate the JS divergence term between the hash code and image features to obtain the mutual information between the hash code and image features, as the first mutual information, and maximize the correlation between the hash code and image features; The second optimization branch: includes a second fully connected layer and a second discriminator connected in sequence. The input of the first fully connected layer is connected to the output of the dual-encoding network, and is used to maximize the JS divergence of different hash codes by constructing positive and negative sample pairs of the hash code, so that similar hash codes are close and different hash codes are far away, improving the discriminability of the hash code, as the optimization of the second mutual information; The third optimization branch: includes a third fully connected layer and a second discriminator connected in sequence. The input of the third fully connected layer is connected to the output of the dual-encoding network, and is used to maximize the JS divergence of different continuous latent variables by constructing positive and negative sample pairs of the continuous latent variable, so that similar continuous latent variables are close and different continuous latent variables are far away, improving the discriminability of the continuous latent variable, as the optimization of the third mutual information.
5. The unsupervised hashing retrieval method with dual paths based on mutual information maximization according to claim 1, characterized in that: The graph convolutional network optimizes the node representation using the message passing mechanism, making the hash codes and latent variables of similar data points closer.
6. The unsupervised hashing retrieval method with dual paths based on mutual information maximization according to claim 1, characterized in that, Based on the finally obtained hash code of the data to be retrieved to obtain the retrieval result, specifically including: Calculate the Hamming distance between the hash code to be retrieved and the database hash code, and sort them; Return the corresponding original data according to the hash code with the smallest Hamming distance to achieve similarity retrieval.
7. The unsupervised hashing retrieval method with dual paths based on maximizing mutual information according to claim 6, characterized in that: Perform an exclusive OR calculation on the hash code to be retrieved and the stored hash code to obtain the Hamming distance of the hash code.
8. The dual-path unsupervised hashing retrieval system based on maximizing mutual information is characterized in that, It includes: An image feature extraction module configured to extract image features for the acquired image; A dual-encoding module configured to use a dual-encoding network to learn the image features to obtain an approximate hash code and a continuous latent variable; The dual-encoding network connection is provided with a triple mutual information optimization module using a discriminator, which optimizes the approximate hash code, continuous latent variable, and image features by using the triple mutual information optimization method during the training process, and optimizes the output of the dual-encoding network through discriminator adversarial learning; A graph convolution module, configured to process the obtained approximate hash code and continuous latent variable by using a graph convolutional network to obtain a final latent variable; A reconstruction retrieval module, configured to input the latent variable into a decoder for reconstruction, finally obtain the hash code of the data to be retrieved, and obtain a retrieval result based on the hash code.
9. An electronic device, characterized in that, It includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the dual-path unsupervised hashing retrieval method based on mutual information maximization according to any one of claims 1-7 are completed.
10. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the steps in the dual-path unsupervised hashing retrieval method based on mutual information maximization according to any one of claims 1-7 are completed.
Citation Information
Cited By
Hash coding optimization method based on self-consistency diffusion module
CN120579580A
A hash coding optimization method based on self-consistent diffusion module
CN120579580B