Unsupervised image hash retrieval method based on similarity distillation
By adopting similarity distillation technology in unsupervised deep hash retrieval, the global similarity information is transmitted from the feature space to the hash space, which solves the problem of difficulty in mining global similarity structure information in the existing technology, and significantly improves the accuracy and efficiency of image retrieval.
Patent Information
- Application Number
- CN202411995477.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-23
AI Technical Summary
The existing unsupervised deep hash retrieval technology is difficult to effectively mine global similarity structure information between samples, resulting in insufficient retrieval performance and high time complexity of paired similarity calculations, which limits scalability.
Using an unsupervised image hash retrieval method based on similarity distillation, the similarity distribution building block and knowledge distillation module are used to pass the global similarity information between samples from the feature space to the hash space, and a hash code that can better reflect the similarity relationship between samples is generated.
It significantly improves the accuracy and efficiency of image retrieval, can understand the structural characteristics of the data more comprehensively, and overcomes the problems of insufficient similarity relationship extraction and inefficient learning in traditional methods.
Smart Images

Figure CN120030180A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image retrieval, and in particular relates to an unsupervised image hash retrieval method based on similarity distillation. Background Art
[0002] With the rapid growth of digital image data, the importance of large-scale image retrieval technology has become increasingly prominent. Traditional image retrieval methods usually rely on high-dimensional feature extraction and similarity calculation. However, these methods face challenges in storage and computing efficiency on large-scale data sets. Hash-based retrieval methods have become an important means to solve these problems due to their significant storage and computing advantages. Hash technology can effectively reduce storage requirements and accelerate the retrieval process by mapping high-dimensional data to low-dimensional binary codes.
[0003] Deep hashing methods have significant advantages over traditional hashing methods. They can automatically extract high-level features and capture complex nonlinear relationships through deep learning models, while traditional methods often rely on manually designed features. Deep hashing methods usually implement end-to-end training, have good flexibility and adaptability, and can easily cope with the needs of large-scale data sets. Supervised deep hashing methods can achieve good results in large-scale image retrieval with the help of precise annotation information, but this method requires expensive data annotation costs. In recent years, the emergence of unsupervised deep hashing methods has provided a new solution for image retrieval. This type of method does not rely on manually annotated training data, but automatically generates effective binary representations by learning the intrinsic features and structure of the image. This makes unsupervised deep hashing methods show strong flexibility and adaptability when processing unlabeled data.
[0004] Existing unsupervised deep hashing retrieval techniques usually rely on pairwise similarity constraints to calculate the similarity between samples. This method generates a similarity matrix to guide the model to optimize the hash code during training, so that the hash code distance between similar samples is as close as possible, while the distance between dissimilar samples is as far as possible.
[0005] Although this method can improve the accuracy of retrieval to a certain extent, it mainly focuses on local similarity and cannot effectively mine the global similarity structural information between samples. This method may not be able to effectively capture the complex structure of the data in practical applications, resulting in insufficient retrieval performance. Pairwise similarity methods may also have biases in sample selection, making the model susceptible to noise samples, thereby reducing the robustness and generalization ability of the model. In addition, the time complexity of pairwise similarity calculation is high, resulting in low learning efficiency, especially when dealing with large-scale data sets, the computational cost increases significantly. Since sample pairs need to be compared one by one, this not only increases training time, but also limits scalability. Summary of the invention
[0006] The present invention provides an unsupervised image hash retrieval method based on similarity distillation, which is used to solve the above problems or at least partially solve the above problems.
[0007] In a first aspect, an unsupervised image hash retrieval method based on similarity distillation is disclosed, the method comprising:
[0008] Step S1: Obtain a group of images for which hash codes are to be generated, and extract high-dimensional feature vectors of each image;
[0009] Step S2: inputting all the high-dimensional feature vectors into a similarity distribution construction module; the similarity distribution construction module generates, for each high-dimensional feature vector, the similarity between the high-dimensional feature vector and each other high-dimensional feature vector; and based on the similarity between the high-dimensional feature vector and each other high-dimensional feature vector, generates a similarity distribution of a feature space corresponding to the high-dimensional feature vector;
[0010] Step S3: inputting the similarity distribution of the feature space corresponding to each high-dimensional feature vector into the trained knowledge distillation module to generate the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space;
[0011] Step S4: inputting the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space into the optimized hash function to generate a corresponding hash code for each image;
[0012] Step S5: Obtain the hash code to be queried, and search the hash codes corresponding to the group of images according to the hash code to be queried.
[0013] Preferably, in step S2, the similarity distribution construction module generates, for each high-dimensional feature vector, the similarity between the high-dimensional feature vector and each other high-dimensional feature vector, wherein:
[0014]
[0015] Where i is the image number, N is the total number of images in the group, j and k are the numbers of other images except the i-th image, and v i , v j , v k are the high-dimensional feature vectors of the i-th, j-th, and k-th images respectively; p i (j) is the similarity between the high-dimensional feature vector corresponding to the image with image number i and the high-dimensional feature vector corresponding to the image with image number j.
[0016] Preferably, the step S3: inputs the similarity distribution of the feature space corresponding to each high-dimensional feature vector into the trained knowledge distillation module to generate the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space, wherein:
[0017] The knowledge distillation module includes a teacher submodule and a student submodule. The teacher submodule is used to receive the similarity distribution of the feature space corresponding to each high-dimensional feature vector and the hash feature corresponding to each high-dimensional feature vector, and calculate the similarity distribution of the hash feature corresponding to each high-dimensional feature vector in the hash space:
[0018]
[0019] Among them, q i (j) is the similarity distribution in the hash space between the hash feature corresponding to the image numbered i and the hash feature corresponding to the high-dimensional feature vector corresponding to the image numbered j, h i 、h j 、h k are the hash features corresponding to the high-dimensional feature vectors of the i-th, j-th, and k-th images respectively; h i 、h j 、h k All belong to (-1, +1) K1 , K1 is the length of the hash code to be generated;
[0020] The similarity distribution of the feature space corresponding to each high-dimensional feature vector and the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space are input into the student submodule, and the student submodule stores the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space, and then minimizes the distance between the similarity distribution of the feature space corresponding to the high-dimensional feature vector and the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space through the cross entropy function, wherein the cross entropy function is: in, is the objective function to be optimized, is the cross entropy function, p i is the similarity distribution of the feature space corresponding to the high-dimensional feature vector corresponding to the i-th image, p i =[p i (1),p i (2),…,p i (N)],q i is the similarity distribution of the hash features corresponding to the high-dimensional feature vector corresponding to the i-th image in the hash space, q i =[q i (1),q i (2),…,q i (N)].
[0021] Preferably, the hash function receives the similarity distribution of hash features corresponding to each high-dimensional feature vector in the hash space, and obtains a hash code of a specified length after dimensionality reduction.
[0022] Preferably, extracting a high-dimensional feature vector of an image includes: extracting local features and global features of the image using a convolutional neural network, and generating a high-dimensional feature vector of the image based on the local features and the global features.
[0023] In a second aspect, an unsupervised image hash retrieval device based on similarity distillation is disclosed, the device comprising:
[0024] Feature extraction module: configured to obtain a group of images for which hash codes are to be generated, and extract high-dimensional feature vectors of each image;
[0025] A first similarity distribution module: configured to input all the high-dimensional feature vectors into a similarity distribution construction module; the similarity distribution construction module generates, for each high-dimensional feature vector, the similarity between the high-dimensional feature vector and each other high-dimensional feature vector; and generates a similarity distribution of a feature space corresponding to the high-dimensional feature vector based on the similarity between the high-dimensional feature vector and each other high-dimensional feature vector;
[0026] The second similarity distribution module is configured to input the similarity distribution of the feature space corresponding to each high-dimensional feature vector into the trained knowledge distillation module to generate the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space;
[0027] Hash code generation module: configured to input the similarity distribution of hash features corresponding to each high-dimensional feature vector in the hash space into the optimized hash function to generate a corresponding hash code for each image;
[0028] The query module is configured to obtain a hash code to be queried, and query the hash codes corresponding to the group of images according to the hash code to be queried.
[0029] In a third aspect, an electronic device is disclosed, the electronic device comprising:
[0030] at least one processor; and
[0031] a memory communicatively connected to the at least one processor; wherein,
[0032] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.
[0033] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is disclosed, wherein the computer instructions are used to cause the computer to execute the method as described above.
[0034] The present invention has the following technical effects:
[0035] The present invention uses an unsupervised image hash retrieval method based on similarity distillation and the knowledge distillation technology in deep learning to effectively transfer the global similarity information between samples from the feature space to the hash space. Since the technical means can construct the global similarity distribution of training samples and minimize the similarity distribution of feature space and hash space, it can fully mine and retain the similarity structure between samples. Therefore, it solves the problems of insufficient similarity relationship extraction and low learning efficiency of traditional unsupervised hash methods, and significantly improves the accuracy and efficiency of image retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a flowchart of the unsupervised image hash retrieval method based on similarity distillation;
[0037] Figure 2 The schematic diagram of the architecture of the unsupervised image hash retrieval method based on similarity distillation;
[0038] Figure 3 Schematic diagram of the structure of the unsupervised image hash retrieval device based on similarity distillation. DETAILED DESCRIPTION
[0039] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0040] like Figure 1-Figure 2 As shown, the present invention provides an unsupervised image hash retrieval method based on similarity distillation, the method comprising:
[0041] Step S1: Obtain a group of images for which hash codes are to be generated, and extract high-dimensional feature vectors of each image;
[0042] Step S2: inputting all the high-dimensional feature vectors into a similarity distribution construction module; the similarity distribution construction module generates, for each high-dimensional feature vector, the similarity between the high-dimensional feature vector and each other high-dimensional feature vector; and based on the similarity between the high-dimensional feature vector and each other high-dimensional feature vector, generates a similarity distribution of a feature space corresponding to the high-dimensional feature vector;
[0043] Step S3: inputting the similarity distribution of the feature space corresponding to each high-dimensional feature vector into the trained knowledge distillation module to generate the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space;
[0044] Step S4: inputting the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space into the optimized hash function to generate a corresponding hash code for each image;
[0045] Step S5: Obtain the hash code to be queried, and search the hash codes corresponding to the group of images according to the hash code to be queried.
[0046] The present invention first extracts features from training samples, identifies the global similarity relationship between samples, and forms a similarity distribution. This distribution not only takes into account the local similarity of samples, but also integrates global information, ensuring that the method can more comprehensively understand the structural characteristics of the data. Subsequently, through knowledge distillation technology, this global similarity structure is converted into an optimization target for hash code learning, so that the generated hash code can better reflect the true similarity relationship between samples. In this way, the limitations of existing pairwise similarity methods in sampling and similarity calculation can be overcome, and the accuracy and efficiency of hash retrieval can be improved.
[0047] The extracting of the high-dimensional feature vector of the image includes: extracting local features and global features of the image using a convolutional neural network, and generating the high-dimensional feature vector of the image based on the local features and the global features.
[0048] In step S2, the similarity distribution construction module generates, for each high-dimensional feature vector, the similarity between the high-dimensional feature vector and each other high-dimensional feature vector, wherein:
[0049]
[0050] Where i is the image number, N is the total number of images in the group, j and k are the numbers of other images except the i-th image, and v i , v j , v k are the high-dimensional feature vectors of the i-th, j-th, and k-th images respectively; p i (j) is the similarity between the high-dimensional feature vector corresponding to the image with image number i and the high-dimensional feature vector corresponding to the image with image number j.
[0051] In the present invention, the similarity distribution construction module constructs multiple similarity distributions by calculating the similarity relationship between a certain image and all other images in the same batch. This process uses cosine similarity as a similarity metric to form a similarity matrix and a similarity distribution. The similarity distribution fully mines the similarity structural information between each image in the same batch, providing an important basis for subsequent knowledge distillation.
[0052] The step S3: inputting the similarity distribution of the feature space corresponding to each high-dimensional feature vector into the trained knowledge distillation module to generate the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space, wherein:
[0053] The knowledge distillation module includes a teacher submodule and a student submodule. The teacher submodule is used to receive the similarity distribution of the feature space corresponding to each high-dimensional feature vector and the hash feature corresponding to each high-dimensional feature vector, and calculate the similarity distribution of the hash feature corresponding to each high-dimensional feature vector in the hash space:
[0054]
[0055] Among them, q i (j) is the similarity distribution in the hash space between the hash feature corresponding to the image numbered i and the hash feature corresponding to the high-dimensional feature vector corresponding to the image numbered j, h i 、h j 、h k are the hash features corresponding to the high-dimensional feature vectors of the i-th, j-th, and k-th images respectively; h i 、h j 、h k All belong to (-1, +1) K1 , K1 is the length of the hash code to be generated;
[0056] The similarity distribution of the feature space corresponding to each high-dimensional feature vector and the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space are input into the student submodule, and the student submodule stores the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space, and then minimizes the distance between the similarity distribution of the feature space corresponding to the high-dimensional feature vector and the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space through the cross entropy function, wherein the cross entropy function is: in, is the objective function to be optimized, is the cross entropy function, p i is the similarity distribution of the feature space corresponding to the high-dimensional feature vector corresponding to the i-th image, p i =[p i (1),p i (2),…,p i (N)],q i is the similarity distribution of the hash features corresponding to the high-dimensional feature vector corresponding to the i-th image in the hash space, q i =[q i (1),q i (2),…,q i (N)].
[0057] The training process of the knowledge distillation module is to use the similarity distribution corresponding to the high-dimensional feature vectors corresponding to the training samples and the hash features corresponding to each high-dimensional feature vector for training. The loss function used in the training process is
[0058] The present invention uses knowledge distillation to transfer similarity information from feature space to hash space. Specifically, the module constrains the relationship between similarity distributions in feature space and hash space by calculating cross entropy. The teacher model generates similarity distribution in feature space, while the student model minimizes cross entropy so that the distribution of its hash code is consistent with the feature distribution of the teacher model. The student model can effectively learn the global similarity information in the feature space, thereby achieving better similarity preservation in the hash space.
[0059] In the present invention, the hash function receives the similarity distribution of hash features corresponding to each high-dimensional feature vector in the hash space, obtains a hash code of a specified length after dimensionality reduction, and in this process maintains the similarity relationship of the high-dimensional feature space between the hash codes.
[0060] like Figure 2 As shown, the unsupervised image hash retrieval method based on similarity distillation of the present invention includes four parts: feature extraction module, similarity distribution construction module, knowledge distillation module and hash code generation module. The feature extraction module uses a convolutional neural network to extract the high-dimensional features of the image. The similarity distribution construction module and the knowledge distillation module are responsible for mining the similarity structure information of multiple samples and passing the information into the hash code. The hash code generation module is responsible for the quantization and generation of low-dimensional hash codes.
[0061] like Figure 3 As shown, the present invention provides an unsupervised image hash retrieval device based on similarity distillation, the device comprising:
[0062] Feature extraction module: configured to obtain a group of images for which hash codes are to be generated, and extract high-dimensional feature vectors of each image;
[0063] A first similarity distribution module: configured to input all the high-dimensional feature vectors into a similarity distribution construction module; the similarity distribution construction module generates, for each high-dimensional feature vector, the similarity between the high-dimensional feature vector and each other high-dimensional feature vector; and generates a similarity distribution of a feature space corresponding to the high-dimensional feature vector based on the similarity between the high-dimensional feature vector and each other high-dimensional feature vector;
[0064] The second similarity distribution module is configured to input the similarity distribution of the feature space corresponding to each high-dimensional feature vector into the trained knowledge distillation module to generate the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space;
[0065] Hash code generation module: configured to input the similarity distribution of hash features corresponding to each high-dimensional feature vector in the hash space into the optimized hash function to generate a corresponding hash code for each image;
[0066] The query module is configured to obtain a hash code to be queried, and query the hash codes corresponding to the group of images according to the hash code to be queried.
[0067] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the above embodiments, it should be understood by those skilled in the art that the technical solutions described in the above embodiments can still be modified, or some or all of the technical features therein can be replaced by equivalents, and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An unsupervised image hash retrieval method based on similarity distillation, characterized in that: The method comprises the following steps: Step S1: Obtain a group of images for which hash codes are to be generated, and extract high-dimensional feature vectors of each image; Step S2: inputting all the high-dimensional feature vectors into a similarity distribution construction module; the similarity distribution construction module generates, for each high-dimensional feature vector, the similarity between the high-dimensional feature vector and each other high-dimensional feature vector; and based on the similarity between the high-dimensional feature vector and each other high-dimensional feature vector, generates a similarity distribution of a feature space corresponding to the high-dimensional feature vector; Step S3: inputting the similarity distribution of the feature space corresponding to each high-dimensional feature vector into the trained knowledge distillation module to generate the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space; Step S4: inputting the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space into the optimized hash function to generate a corresponding hash code for each image; Step S5: Obtain the hash code to be queried, and search the hash codes corresponding to the group of images according to the hash code to be queried.
2. The method according to claim 1, characterized in that In step S2, the similarity distribution construction module generates, for each high-dimensional feature vector, the similarity between the high-dimensional feature vector and each other high-dimensional feature vector, wherein: Where i is the image number, N is the total number of images in the group, j and k are the numbers of other images except the i-th image, and v i , v j , v k are the high-dimensional feature vectors of the i-th, j-th, and k-th images respectively; p i (j) is the similarity between the high-dimensional feature vector corresponding to the image with image number i and the high-dimensional feature vector corresponding to the image with image number j.
3. The method according to claim 2, characterized in that The step S3: inputting the similarity distribution of the feature space corresponding to each high-dimensional feature vector into the trained knowledge distillation module to generate the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space, wherein: The knowledge distillation module includes a teacher submodule and a student submodule. The teacher submodule is used to receive the similarity distribution of the feature space corresponding to each high-dimensional feature vector and the hash feature corresponding to each high-dimensional feature vector, and calculate the similarity distribution of the hash feature corresponding to each high-dimensional feature vector in the hash space: Among them, q i (j) is the similarity distribution in the hash space between the hash feature corresponding to the image numbered i and the hash feature corresponding to the high-dimensional feature vector corresponding to the image numbered j, h i 、h j 、h k are the hash features corresponding to the high-dimensional feature vectors of the i-th, j-th, and k-th images respectively; h i 、h j 、h k All belong to (-1, +1) K1 , K1 is the length of the hash code to be generated; The similarity distribution of the feature space corresponding to each high-dimensional feature vector and the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space are input into the student submodule, and the student submodule stores the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space, and then minimizes the distance between the similarity distribution of the feature space corresponding to the high-dimensional feature vector and the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space through the cross entropy function, wherein the cross entropy function is: in, is the objective function to be optimized, is the cross entropy function, p i is the similarity distribution of the feature space corresponding to the high-dimensional feature vector corresponding to the i-th image, p i =[p i (1),p i (2),…,p i (N)],q i is the similarity distribution of the hash features corresponding to the high-dimensional feature vector corresponding to the i-th image in the hash space, q i =[q i (1),q i (2),…,q i (N)].
4. The method according to claim 3, characterized in that The hash function receives the similarity distribution of hash features corresponding to each high-dimensional feature vector in the hash space, and obtains a hash code of a specified length after dimensionality reduction.
5. The method according to claim 3, characterized in that Extracting a high-dimensional feature vector of an image includes: extracting local features and global features of the image using a convolutional neural network, and generating a high-dimensional feature vector of the image based on the local features and the global features.
6. An unsupervised image hash retrieval device based on similarity distillation, characterized in that: The device comprises: Feature extraction module: configured to obtain a group of images for which hash codes are to be generated, and extract high-dimensional feature vectors of each image; A first similarity distribution module: configured to input all the high-dimensional feature vectors into a similarity distribution construction module; the similarity distribution construction module generates, for each high-dimensional feature vector, the similarity between the high-dimensional feature vector and each other high-dimensional feature vector; and generates a similarity distribution of a feature space corresponding to the high-dimensional feature vector based on the similarity between the high-dimensional feature vector and each other high-dimensional feature vector; The second similarity distribution module is configured to input the similarity distribution of the feature space corresponding to each high-dimensional feature vector into the trained knowledge distillation module to generate the similarity distribution of the hash features corresponding to each high-dimensional feature vector in the hash space; Hash code generation module: configured to input the similarity distribution of hash features corresponding to each high-dimensional feature vector in the hash space into the optimized hash function to generate a corresponding hash code for each image; The query module is configured to obtain a hash code to be queried, and query the hash codes corresponding to the group of images according to the hash code to be queried.
7. A computer-readable storage medium, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the method as claimed in any one of claims 1 to 5.
8. An electronic device, characterized in that: The electronic device comprises: A processor, which is used to execute multiple instructions; A memory for storing a plurality of instructions; The plurality of instructions are used to be stored in the memory and loaded and executed by the processor according to any one of claims 1 to 5.
Citation Information
Patent Citations
Application of visual language knowledge distillation in cross-modal hash retrieval
CN116594994A