A hash coding optimization method based on self-consistent diffusion module

By optimizing the hash encoding through a self-consistent diffusion module, the problems of high computational complexity and poor performance of new samples in the deep hashing method are solved, achieving robust and efficient hash encoding suitable for image retrieval and large-scale multimedia retrieval.

CN120579580BActive Publication Date: 2025-10-24SICHUAN ZHONGTIAN YINGYAN INFORMATION TECH CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511087410.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-24
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Existing deep hashing methods suffer from high computational complexity and memory consumption, slow retrieval speed in real-time applications and large-scale datasets, and performance degradation on new sample data.

Method used

A self-consistent diffusion module is used to optimize hash encoding. The feature embedding matrix is ​​extracted by a pre-trained feature extractor and feature embedder. The hash function is then projected onto the target hash space and a smooth latent representation vector is obtained through a diffusion layer. The feature encoding capability of the hash function is optimized by minimizing the Euclidean distance.

Benefits of technology

It improves the robustness and feature diversity of hash functions, reduces computational complexity, keeps the model size constant, avoids the risk of data leakage, and enhances performance on new samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579580B_ABST
    Figure CN120579580B_ABST
Patent Text Reader

Abstract

The application discloses a hash coding optimization method based on a self-consistent diffusion module and belongs to the technical field of deep hash. The parameter-independent diffusion mechanism is used to smooth the potential representation by using a graph relation matrix, information interaction and feature diversity between samples are enhanced, and a self-encoder structure is combined to take a hash function as an encoder and a mirror decoder to reconstruct features, and the hash coding capability is refined by minimizing the reconstruction error. A two-stage strategy is adopted for training to separate feature extraction and hash optimization. The implementation effect is significant in multiple standard dataset verification: the hash code quality and retrieval precision are improved, the memory consumption and computational complexity are reduced, the generalization performance for new samples is enhanced, and the method is especially suitable for large-scale real-time retrieval scenes with limited supervision information without additional parameter burden.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of deep hashing, and in particular to a hash coding optimization method based on a self-consistency diffusion module. BACKGROUND

[0002] With the explosive growth of data, the information overload problem faced by information retrieval is becoming increasingly serious, especially in the processing of multimedia data such as images and videos. Developing an efficient and accurate retrieval method has become a key problem to be solved at present. The deep hashing method using deep neural network to learn data features and generate compact binary hash codes can optimize feature extraction and hash code generation at the same time, and automatically extract high-level semantic features of samples. Compared with traditional hash methods that rely on hand-designed features, deep hashing methods can better preserve the similarity of data, and provide an efficient and accurate retrieval method for large-scale multimedia retrieval tasks such as image retrieval, multi-modal data processing and large-scale database retrieval.

[0003] The core goal of deep hash coding is to learn a hash function through a deep neural network to map high-dimensional data to low-dimensional binary codes while preserving the semantic similarity of data as much as possible. This requires the feature extraction network to extract sufficient feature information and ensure that these feature information is uniformly distributed in the semantic space. Therefore, existing deep hashing methods mainly focus on designing powerful feature extraction networks to reduce the quantization error when mapping sample features to binary hash codes. However, a powerful feature extractor usually means higher computational complexity and memory consumption, which in turn leads to a decrease in retrieval speed, restricting the use of deep hashing methods in real-time applications or large-scale data sets. In addition, in the scene where the supervision information is limited, optimizing only the feature extractor to extract high-quality sample semantic features can perform well on the original samples, but it is difficult to capture effective feature information when encountering new samples, resulting in a decrease in performance on new sample data SUMMARY

[0004] In view of the above deficiencies in the prior art, the present application provides a hash coding optimization method based on a self-consistency diffusion module.

[0005] In order to achieve the above-mentioned application purposes, the technical scheme adopted by the present application is as follows:

[0006] A hash coding optimization method based on a self-consistency diffusion module, comprising the following steps:

[0007] S1, inputting an RGB picture into a pre-trained feature extractor and a feature embedder in sequence for feature extraction and feature embedding to obtain a feature embedding matrix;

[0008] S2, projecting the feature embedding matrix into a target hash space using a hash function to obtain a latent representation vector of the sample, and inputting the latent representation vector of the sample into a diffusion layer to obtain a smoothed latent representation vector;

[0009] S3, inputting the smoothed latent representation vector into a decoding module for decoding to obtain a reconstructed sample feature embedding matrix, and calculating the Euclidean distance between the reconstructed feature embedding matrix and the original feature embedding matrix;

[0010] S4, taking the minimization of the Euclidean distance as the optimization target of reconstruction to optimize the feature coding ability of the hash function.

[0011] Further, the pre-trained feature extractor and feature embedder in S1 are one of ResNet, DenseNet, DarkNet, and Swim Transformer, which satisfies that the generated feature embedding matrix is input to the hash function in step S2.

[0012] Further, the latent representation vector of the sample in S2 is represented as:

[0013]

[0014] In the formula, is an activation function, and b represent the weight matrix and bias vector of the hash function, respectively, is the latent representation vector of the sample, is the original feature embedding matrix.

[0015] Further, the smoothed latent representation vector in S2 is represented as:

[0016]

[0017] In the formula, is a batch of latent representation vectors, is a batch of smoothed latent representation vectors, is a Laplacian matrix composed of a degree matrix and a weight matrix, is a step size.

[0018] Further, the specific calculation method of the Euclidean distance between the reconstructed feature embedding matrix and the original feature embedding matrix in S3 is:

[0019]

[0020] In the formula, is the reconstructed sample feature embedding matrix obtained by the decoding module reconstructing the smoothed latent representation vector , denotes the original feature embedding matrix, denotes the Euclidean distance, is the Euclidean distance between the reconstructed feature embedding matrix and the original feature embedding matrix.

[0021] The present application has the following beneficial effects:

[0022] 1、The present application promotes collisions between samples through information diffusion, thereby enhancing information exchange between samples and their neighbors. The smoothed latent representation optimized through diffusion has stronger feature diversity and anti-interference ability, which improves the robustness of the hash function to changes in input data.

[0023] 2、The present application adopts a parameter-independent diffusion mechanism, and the diffusion layer does not introduce any additional parameters or labels, which can ensure that the size of the model remains unchanged and avoid the risk of data leakage.

[0024] 3、Minimizing the Euclidean distance as the optimization objective of reconstruction aims to minimize the reconstruction error and force the hash function to learn and capture the fine-grained features of the input feature embedding. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 is the implementation flowchart of the present application.

[0026] Figure 2 is the training flowchart of the deep hash coding optimization algorithm of the present application.

[0027] Figure 3 is the verification flowchart of the deep hash coding optimization algorithm of the present application. DETAILED DESCRIPTION

[0028] The specific embodiments of the present application are described below to facilitate understanding of the present application by those skilled in the art, but it should be clear that the present application is not limited to the scope of the specific embodiments, and for those skilled in the art, any changes within the spirit and scope of the present application as defined and determined by the appended claims are obvious, and all applications utilizing the concept of the present application are within the scope of protection.

[0029] As shown in Figure 1 The present application embodiment discloses a hash coding optimization method based on a self-consistent diffusion module, comprising the following steps:

[0030] Step S1, input the RGB picture into the pre-trained feature extractor and feature embedder in sequence for feature extraction and feature embedding, and obtain the feature embedding matrix;

[0031] In this embodiment, the pre-trained feature extractor and feature embedder in step S1 can be any algorithm module capable of extracting features and performing feature embedding from an RGB picture, as long as the generated feature embedding matrix can be used as the input of the hash function in step S2.

[0032] In step S2, the feature embedding matrix is projected into a target hash space by using a hash function to obtain a latent representation vector of the sample; and the latent representation vector of the sample is input into a diffusion layer to obtain a smoothed latent representation vector.

[0033] In this embodiment, the hash function is essentially a fully connected layer neural network, and the calculation of the latent representation vector can be represented by the following formula:

[0034]

[0035] wherein, represents an activation function, and b represent a weight matrix and a bias vector of the hash function (i.e., the fully connected layer), respectively, represents the latent representation vector of the sample.

[0036] The diffusion mechanism of the designed diffusion layer is as follows:

[0037] Through information diffusion, the collision between samples is promoted, thereby enhancing the information exchange between the sample and its neighbors. The smoothed latent representation after diffusion optimization has stronger feature diversity and anti-interference ability, which improves the robustness of the hash function to changes in input data.

[0038] The calculation of the smoothed latent representation vector can be represented by the following formula:

[0039]

[0040] wherein, Z represents a batch of latent representation vectors, and the shape is bs represents the batch size (batchsize), i.e., bs samples are included in a single batch, and n represents the length of the hash code to be generated, i.e., the length of the latent representation vector of a single sample. represents a batch of smoothed latent representation vectors, and the shape is also , represents a step size for controlling the smoothing strength. represents a Laplacian matrix composed of a degree matrix D and a weight matrix W, which is used to describe the connection strength between samples.

[0041] wherein, the weight matrix W is a symmetric matrix with a size of , which represents the "similarity" between the samples in the batch based on the Euclidean distance between the samples. In this embodiment, any element in the weight matrix W is calculated by the following formula: represents the similarity between sample i and sample j (i ), which is calculated by a Gaussian kernel function, i.e.:

[0042]

[0043] wherein, represents the latent representation vector of sample i and the latent representation vector of sample j between the Euclidean distance, represents the kernel width, which is used to control the rate of similarity decay.

[0044] The degree matrix D is a diagonal matrix whose diagonal elements represent the sum of the connection weights of sample i to all other samples in the batch, i.e.:

[0045] The present application adopts a parameter-independent diffusion mechanism, and the diffusion layer does not introduce any additional parameters or labels, which can ensure that the size of the model remains unchanged and avoid the risk of data leakage.

[0046] Step S3, input the smoothed latent representation vector into the decoding module for decoding to obtain a reconstructed sample feature embedding matrix; calculate the Euclidean distance between the reconstructed sample feature embedding matrix and the original feature embedding matrix.

[0047] In the present embodiment, the decoding module in step S3 is designed based on the mirror of the hash function, and the module and the hash function form an encoding-decoding structure, i.e. a self-encoder structure based on the self-consistency assumption. The self-consistency refers to a high-quality hash code that should maintain the complete information entropy of the source data and be able to recover the data before compression.

[0048] In the present embodiment, the calculation formula in step S3 is:

[0049]

[0050] wherein,

[0051] represents the reconstructed sample feature embedding matrix obtained by the decoding module reconstructing the smoothed latent representation vector represents the original feature embedding matrix, represents the Euclidean distance. Step S4, taking the minimization of the Euclidean distance as the optimization target of reconstruction, to refine the feature coding ability of the hash function.

[0052] Step S4, taking the minimization of the Euclidean distance as the optimization target of reconstruction, to refine the feature coding ability of the hash function.

[0053] ​In this embodiment, the Euclidean distance is minimized in step S4 as the optimization objective of reconstruction, aiming to minimize the reconstruction error and force the hash function to learn and capture the fine-grained features of the input feature embedding.

[0054] The embodiment of the present application adopts a two-stage training strategy, the first stage trains the feature extraction network, taking picture classification as the training task and taking the classification loss as the optimization objective. The specific loss function and optimizer design depend on the feature extraction network adopted in actual use. The second stage trains the hash function , refining its feature encoding capability. The embodiment conducts experiments on the three data sets of miniImageNet, tieredImageNet and CUB-200, each of which is divided into two non-overlapping sub-data sets D 1 and D 2 , corresponding to the first stage and the second stage respectively. The division of D 1 and D 2 is (64, 20), (351, 160) and (100, 50) respectively, where the former in the parentheses represents the number of categories contained in the first stage data set D 1 , and the latter represents the number of categories contained in the second stage data set D 2 .

[0055] Figure 2 The algorithm training process shown in the figure corresponds to the training of the second stage, i.e. optimizing the hash function. In this stage, only the Euclidean distance between the original feature embedding matrix and the reconstructed feature embedding matrix is taken as the training loss function, without introducing other losses. The implementation details will be described in detail.

[0056] In the second stage training, the feature extractor and the feature embedder are frozen, and the data set D 2 is further divided into a training set D T and a validation set D V . Then training is performed on the training set D T , and validation is performed on the validation set D V . The learning rate of the embodiment of the present application is set to 0.1, and the training round is 100.

[0057] When training on the training set D T , first input the RGB image that needs to be subjected to feature extraction, whose size is 3xHxW. Where 3 is the number of channels, H is the image height, and W is the image width. The extracted image feature has a size of CxH1xW1, where C is the number of channels, H1 is the feature height, and W1 is the feature width. Then the extracted feature is input into the feature embedder, and a feature embedding matrix with a size of H2xW2 is output.

[0058] The saved feature embedding matrix , and input it into a hash function, output a latent representation vector with length n . That is:

[0059]

[0060] Here, the length n of the latent representation vector is consistent with the length of the final output hash code, and the value of n can be 16, 32, 64, etc. depending on actual requirements.

[0061] The diffusion mechanism and smoothing formula of the diffusion layer are as described in step S2. In the embodiment of the present application, the Laplacian matrix L is a square matrix with a size of bs x bs, and by performing matrix multiplication with the latent representation vector matrix Z formed by all samples in the current batch, the feature difference between each sample in the batch and its adjacent sample can be calculated. Subtract the difference from Z, and the latent representation will adjust to the "average direction" of the adjacent sample to achieve the smoothing effect.

[0062] Subsequently, the smoothed latent representation vector is decoded to generate a reconstructed feature embedding matrix with a size of H2 x W2. The decoder used in the embodiment of the present application is designed based on the mirror of the hash function , and the purpose is to make as similar as possible to the original feature matrix , that is, the Euclidean distance between them is as small as possible. By calculating the Euclidean distance between and , and using it as a training loss function to force the model to minimize the reconstruction error, the hash function can learn and capture the fine-grained features of the input feature embedding, and generate higher-quality hash codes.

[0063] After the training converges on the training set D T , the training effect will be verified, and the verification process is as shown in Figure 3 . The difference from the training process is that after the input RGB image is smoothed by the diffusion layer to obtain a smoothed latent representation vector, the representation vector does not need to be decoded and reconstructed, but is binarized using the sign function sign(x) to obtain the final hash code. The sign function sign(x) is defined as follows:

[0064]

[0065] In the verification of the embodiment of the present application, the training set D T and the verification set D V are respectively divided into Figure 3The flow shown inputs a model and outputs two corresponding hash code sets C T and C V Each hash code in the hash code set corresponds to a sample. For each hash code in C V , the Hamming distance between it and each hash code in C T is calculated, sorted in ascending order according to the distance, and the average precision AP of a single verification sample is calculated according to the sorting result. Then, the mAP index is calculated by averaging the AP of all verification samples, and is expressed in a 95% confidence interval.

[0066] In this embodiment, ResNet18 and a global average pooling layer are taken as examples of the feature extractor and the feature embedder, respectively. In the first stage of training, SGD is taken as the training optimizer, and the cross-entropy loss is taken as the picture classification loss.

[0067] In the second stage of training, the weights of the ResNet18 trained in the first stage are frozen. The size of the final feature extracted by the ResNet18 network from a single input picture (sample m, actually, a batch is input in the actual training and verification, and a single sample is taken as an example here. The size of the batch depends on the video memory. In this embodiment of the present application, the batch size bs is 16) is , where 512 is the number of channels, is the height and width of the feature. The feature is input into the global average pooling layer for feature embedding, and the size of the embedded feature is , which is reshaped into , that is, the final feature embedding matrix is obtained.

[0068] Taking the generation of a hash code with a length of 32 bits as an example, the input size of the corresponding hash function (a fully connected layer) is 512, and the output size is 32. The feature embedding matrix is input into the hash function to output a latent representation vector with a length of 32 bits (a vector in the latent representation vector matrix Z), which is then smoothed by the diffusion layer to obtain a smoothed latent representation vector with a length of 32 bits (a vector in the smoothed latent representation vector matrix ).

[0069] In the training stage (the second stage of training), the input size of the corresponding decoder (a fully connected layer) is 32, and the output size is 512. The input of the decoder outputs a reconstruction vector with a length of 512 bits, which is reshaped into , that is, the reconstructed feature embedding matrix .

[0070] In the verification stage, the hash code of the sample m is calculated by inputting the sample m into the sign(x) function for binary processing, and outputting a 32-bit binary vector. In the verification stage, the hash code of the sample m is calculated by inputting the sample m into the sign(x) function for binary processing, and outputting a 32-bit binary vector.

[0071] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0072] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0073] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0074] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0075] Those skilled in the art will appreciate that the embodiments described herein are presented for purposes of illustration and that the inventive principles are not limited to these particular embodiments. Other variations and modifications can be made to the embodiments without departing from the spirit and scope of the inventive principles.

Claims

1. A method for hash coding optimization based on self-consistent diffusion module, characterized in that, The method comprises the following steps: S1, inputting an RGB picture into a pre-trained feature extractor and a feature embedder in sequence to perform feature extraction and feature embedding, and obtaining a feature embedding matrix; S2, projecting the feature embedding matrix into a target hash space by using a hash function to obtain a latent representation vector of the sample, and inputting the latent representation vector of the sample into a diffusion layer to obtain a smoothed latent representation vector, wherein the diffusion manner of the diffusion layer is: Through information diffusion, the collision between samples is promoted, thereby enhancing the information exchange between the sample and its neighbors, the smoothed latent representation vector after diffusion optimization has stronger feature diversity and anti-interference ability, and the robustness of the hash function to input data changes is improved; Specifically, the latent representation vector of the sample is represented as: wherein, denotes an activation function, and b denote a weight matrix and a bias vector of the hash function, respectively, denotes a latent representation vector of the sample; The smoothed latent representation vector is represented as: wherein is a batch of latent representation vectors, is a batch of smoothed latent representation vectors, is a step size, denotes a Laplacian matrix composed of a degree matrix D and a weight matrix W, and: The weight matrix W is a symmetric matrix, bs represents the batch size, the similarity between samples in the batch is represented based on the Euclidean distance between samples, and any element in the weight matrix W represents the similarity between sample i and sample j, which is calculated by a Gaussian kernel function, that is: wherein, represents the latent representation vector of sample i and the latent representation vector of sample j the Euclidean distance between them, represents the kernel width, used to control the rate of decay of similarity; The degree matrix D is a diagonal matrix with the diagonal elements representing the sum of the connection weights of sample i to all other samples in the batch, i.e.: Dii= åj6B(i) Wij ; S3, inputting the smoothed latent representation vector into a decoding module for decoding to obtain a reconstructed sample feature embedding matrix, and calculating the Euclidean distance between the original feature embedding matrix and the reconstructed feature embedding matrix; S4, taking the minimization of the Euclidean distance as the optimization target of reconstruction to optimize the feature coding ability of the hash function.

2. The hash coding optimization method based on self-consistent diffusion module according to claim 1, wherein, The pre-trained feature extractor and feature embedder in S1 are one of ResNet, DenseNet, DarkNet and SwimTransformer, and the generated feature embedding matrix meets the requirement of being input into the hash function in step S2.

3. The method of claim 1, wherein the method is based on a self-consistent diffusion module. The specific calculation method of the Euclidean distance between the reconstructed feature embedding matrix and the original feature embedding matrix in S3 is: wherein denotes reconstructing the smoothed latent representation vector denotes reconstructing the smoothed latent representation vector denotes reconstructing the smoothed latent representation vector denotes reconstructing the smoothed latent representation vector denotes reconstructing the smoothed latent representation vector

Citation Information

Patent Citations

  • Double-path unsupervised Hash retrieval method and system based on mutual information maximization

    CN120256659A

  • Stem cell quality evaluation system and method

    CN120296390A