Deep hash image retrieval method and system based on Hamming sphere interval
Through the deep hash image retrieval method based on Heming ball interval, the problems of slow retrieval speed, high memory consumption and high training cost in image retrieval are solved, and efficient image retrieval and storage are achieved, which is suitable for multi-category data sets.
Patent Information
- Application Number
- CN202510188474.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art has problems in image retrieval, high memory consumption, high training cost, and hash code quality affected by the number of categories in image retrieval.
The depth hash image retrieval method based on Heming ball interval is adopted. By calculating the value range of Heming ball interval, the appropriate classification network model is selected, and the value of the hash center is dynamically adjusted during the training process to generate an efficient hash code.
An efficient training process is realized, memory consumption and query time is reduced, hash code search effect is improved, and can adapt to more categories of instance search data sets.
Smart Images

Figure CN120032145A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image retrieval, and in particular relates to a deep hash image retrieval method and system based on Hamming sphere interval. Background Art
[0002] With the explosive growth of digital image content, users' demand for image retrieval is increasing. Current image retrieval technology mainly adopts a two-stage approach, that is, first recalling relevant images from massive data sets and then re-ranking relevant candidate images, and finally returning the retrieval results. The dense feature vectors output by the neural network model have achieved remarkable results in the first stage of sorting. However, similarity calculations for large-scale high-dimensional vectors are very time-consuming operations, and storing dense features of images will also consume a lot of memory. Therefore, it is necessary to use image hashing algorithms to map images to a set of binary codes to complete the first stage of retrieval tasks, so as to greatly improve the query speed and reduce memory consumption.
[0003] At present, deep hashing methods are usually used in large-scale image retrieval tasks. Deep hashing is a hashing technology based on deep learning. It generates hash codes with good consistency and discrimination by training a neural network model to learn low-dimensional embedded representations of data. The deep hashing algorithm can map a large amount of image data to a low-dimensional space. By compressing the original high-dimensional data into a binary representation, it is conducive to efficiently calculating similarity using XOR operations. In vector recall, reducing the storage space of vector indexes also greatly improves the retrieval speed.
[0004] The current research on deep hashing mainly focuses on designing objective functions that retain the similarity of the original data. The current common loss functions can be divided into single-point method, pair method, triple method, list method, etc. Among them, the pair method, triple method and list method are all based on sorting losses. These methods rely on a large number of negative samples in training and require the design of appropriate positive and negative sample sampling strategies, which greatly increases the training cost. The single-point rule represented by the hash center can avoid the above problems. Its goal is to assign mutually separated hash codes as the center point for each category and learn compact binary codes by minimizing the intra-class variance. The mainstream method based on the hash center is also limited by the number of categories, which is not conducive to practical application.
[0005] The defects of the existing technology are:
[0006] 1. Directly using the dense feature vector output by the neural network to calculate the similarity of images has the problems of slow retrieval speed and excessive memory consumption.
[0007] 2. The current image hashing method relies on a large number of negative samples in training, and requires the design of appropriate positive and negative sample sampling strategies, which increases the training cost and is not conducive to practical use.
[0008] 3. The quality of the hash code is greatly affected by the number of categories and cannot be well extended to the instance retrieval data set with more categories. Summary of the Invention
[0009] To solve the above technical problems, the present invention provides a deep hash image retrieval method based on Hamming ball distance, including the following steps:
[0010] Step S1: Obtain image data and establish a data set;
[0011] Step S2: Calculate the value range of the Hamming ball distance according to the number of categories in the data set ;
[0012] Step S3: Select a suitable , train an improved classification network model; use the bisection method to update the value, repeat training the model until all values of are enumerated; select the optimal model as the trained improved classification network model;
[0013] Step S4: Input the picture to be queried into the trained improved classification network model to obtain a hash code, and use the Hamming distance obtained by performing an exclusive OR operation on this hash code and the hash codes of the pictures in the retrieval library to judge the similarity between images.
[0014] Advantageous Effects:
[0015] 1. The present invention provides a deep hash image retrieval method based on Hamming ball distance, realizing an efficient training process. The entire training process does not require complex positive and negative sample pair mining, nor does it require generating a set of codebooks in advance as candidates for hash centers. The present invention unifies the two stages of hash encoder and hash center generation. Different from the previous center-based hash methods, it does not require finding a set of codebooks as proxies in advance, but learns the center in real time as the encoder is updated.
[0016] 2. The present invention solves the problem that the performance of the center-based hash method deteriorates when facing a large number of categories, and greatly improves the retrieval effect of the hash code. The previous methods cannot achieve good results when the number of categories is too large, and even cannot find enough codebooks to be used as centers. The present invention indirectly finds the center of the Hamming ball by constraining the distance between Hamming balls, thus solving this difficult problem.
[0017] 3. The hash coding module proposed by the present invention can work independently and can be placed after any feature extractor. By reducing the dimension of the dense features output by models such as neural networks through the hash coding module, the query and storage efficiency of image retrieval is greatly improved. Brief Description of the Drawings
[0018] Figure 1 A schematic diagram of a flow chart of a deep hash image retrieval method based on Hamming sphere interval of the present invention;
[0019] Figure 2 This is a schematic diagram of the effect of using Hamming sphere classification;
[0020] Figure 3 It is a structural diagram of the improved classification network model;
[0021] Figure 4 This is a structural block diagram of a deep hash image retrieval system based on Hamming sphere interval of the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0023] Embodiment 1
[0024] like Figure 1 As shown, a deep hash image retrieval method based on Hamming sphere interval provided by an embodiment of the present invention includes the following steps:
[0025] Step S1: Acquire image data and establish a data set;
[0026] Step S2: Calculate the Hamming sphere interval based on the number of categories in the data set The value range of
[0027] Step S3: Select the appropriate , train the improved classification network model; use the binary update The value of , repeat the training model until the enumeration is exhausted All values of ; select the best model as the trained improved classification network model;
[0028] Step S4: Input the query image into the trained improved classification network model to obtain a hash code, and use the Hamming distance obtained by performing an XOR operation on the hash code and the hash code of the image in the retrieval library to determine the similarity between the images.
[0029] In one embodiment, the above step S1: acquiring image data and establishing a data set specifically includes:
[0030] In the embodiment of the present invention, the training data set used is the Google landmark public data set. 1,186,673 landmark images are used, including 56,441 landmark categories. The test data set and the data set for constructing the retrieval library are the Paris building data set and the Oxford building data set, respectively. Each test data set contains 70 images to be queried, and the retrieval library contains 1,000,000 images.
[0031] In one embodiment, the above step S2: Calculate the Hamming sphere interval according to the number of categories in the data set The value range of includes:
[0032] Step S21: Determine the number of categories in the data set and the size of the encoding space, where the size of the encoding space is the number of bits after hashing The number of states that can be represented is different hash codes; assign a Hamming sphere to each category;
[0033] Step S22: When the code length is In the space of The volume of the Hamming sphere The definition is as follows:
[0034] ;
[0035] Step S23: Definition The minimum Hamming distance between the centers of two Hamming balls is , enumerate all the As The value range of
[0036] .
[0037] The present invention determines the number of categories and the size of the encoding space After that, a Hamming sphere is assigned to each category. During model training, it is hoped that the model will map the data of the corresponding category to the Hamming sphere assigned to it. Figure 2 The intuitive mapping process is presented, the goal of which is to maximize the separability between classes and maintain compactness within classes.
[0038] In order to improve the separability between classes, we need to find a suitable interval .and The value of can be determined by the minimum distance between the sphere centers Sure: , the radius of the Hamming sphere can be compressed arbitrarily, so the minimum distance between the sphere centers is Decided The present invention determines the minimum distance between sphere centers by using the Gilbert-Varshamov bound and the Hamming upper bound in coding theory. The possible range of values for the upper bound.
[0039] In one embodiment, the above step S3: selecting a suitable , train the improved classification network model; use the binary update The value of , repeat the training model until the enumeration is exhausted All values of ; Select the best model as the trained improved classification network model, including:
[0040] Step S31: Use binary search to Determine the current value in the range The value of
[0041] The value of the Hamming interval of the present invention has a greater impact on the entire hash coding module. If the value is too small, it will not be possible to effectively generate hash codes that can be separated between classes, and the retrieval accuracy of the model will decrease. If the value is too large, then theoretically there is no Hamming sphere that satisfies this condition, so the model cannot converge. The value of .
[0042] Step S32: input the data set into the improved classification network model for training, wherein the improved classification network model outputs the dense vector of the feature layer of the traditional classification network model Then the hash coding module was added, which is used to convert Mapping to hash code , where the hash coding module includes: n hash coding layers, a fully connected layer and a batch normalization layer; the hash coding layer includes: a fully connected layer, a batch normalization layer and an activation function:
[0043] Among them, the hash coding module will Mapping to hash code The specific steps are as follows:
[0044] Step S321: The operation of the hash coding module is expressed as:
[0045] ;
[0046] in, It is the hash encoding module;
[0047] Step S322: Update the value of the current hash center, where the hash center is defined as the point with the smallest Hamming distance to all feature points in the category:
[0048] ;
[0049] in, For the The hash code of the image, For the Category labels corresponding to the images; is the Hamming distance of the hash code; For the The value of the hash center of each category; For the A collection of elements of a category;
[0050] For a certain category, the value of its corresponding center is the value of the element that appears most times on each bit. In order to ensure efficiency, the present invention updates the value of the hash center batch by batch. The value of the hash center of each category It can be updated by the following formula:
[0051] ;
[0052] ;
[0053] in, is the symbolic function, The contribution of each bit after mapping the data of the current category to the hash code;
[0054] Step S323: Constructing the total loss based on the center loss and the contrast loss for:
[0055] ;
[0056] in, is a hyperparameter used to balance the weight of the central loss term; is the number of pictures in the current batch; is the similarity matrix. If two pictures are similar ,otherwise ; Among them, the first term in the above formula is the center loss, and the second term is the contrast loss; after calculating the values of the center loss and contrast loss, the gradient descent method is used to update the model parameters;
[0057] Step S324: Repeat steps S322 to S323 until the improved classification network model converges, completing this round of training.
[0058] The improved classification network model of the present invention is composed of an image feature extractor based on a classification network architecture and a post-positioned hash coding module. The architecture of the model is as follows: Figure 3 As shown, the activation function used by the hash coding layer is The present invention adds a batch normalization layer to the last layer of the hash coding module, which can force the network output value to become a Gaussian distribution with a mean of 0, so as to minimize the collision probability of the hash code and improve the retrieval performance.
[0059] Step S33: Update using binary search Repeat step S32 until the enumeration is complete. All values of ; select the optimal model as the trained improved classification network model.
[0060] In one embodiment, the above step S4: input the query image into the trained improved classification network model to obtain a hash code, and perform an XOR operation on the hash code and the hash code of the image in the retrieval library to obtain the Hamming distance to determine the similarity between the images.
[0061] The present invention proposes a deep hash image retrieval method based on the Hamming sphere interval, which can quickly recall relevant items from massive data. For a given query, relevant items can be quickly recalled from massive images for subsequent refined sorting, rearrangement and other processes. The design of the present invention can map high-dimensional image data into a set of binary codes, and can ensure that this set of codes still has the original similarity relationship. The method of the present invention can be applied to training data sets of any number of categories. The method of the present invention eliminates the sensitivity to the number of data categories in the previous center-based hashing scheme. The proposed training strategy starts from the interval of the Hamming sphere in the constrained coding space, thereby bypassing the difficult problem of directly finding the center, so that the center-based hashing algorithm can be extended to any multi-category data sets. The method of the present invention has better retrieval capabilities, can jointly optimize the values of the hash encoder and the hash center, and can combine the original two stages into one stage by dynamically adjusting the value of the hash center during the network training process, so that the hash code can be generated more robustly.
[0062] Embodiment 2
[0063] like Figure 4 As shown, an embodiment of the present invention provides a deep hash image retrieval system based on Hamming sphere interval, comprising the following modules:
[0064] A data set building module 51 is used to acquire image data and build a data set;
[0065] The module 52 for calculating the range of Hamming sphere intervals is used to calculate the Hamming sphere intervals according to the number of categories in the data set. The value range of
[0066] Training the improved classification network model module 53 to select the appropriate , train the improved classification network model; use the binary update The value of , repeat the training model until the enumeration is exhausted All values of ; select the best model as the trained improved classification network model;
[0067] The image retrieval module 54 is used to input the query image into the trained improved classification network model to obtain a hash code, and use the Hamming distance obtained by performing an XOR operation on the hash code and the hash code of the image in the retrieval library to determine the similarity between the images.
[0068] In one embodiment, the above training improved classification network model module adds a hash coding module to the traditional classification network model module to convert the dense vector output by the feature layer of the traditional classification network model into Mapping to hash code .
[0069] A deep hash image retrieval device based on Hamming sphere interval includes one or more electronic devices, wherein the one or more electronic devices are used to implement a deep hash image retrieval method, system and device based on Hamming sphere interval.
[0070] An electronic device includes: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement a deep hash image retrieval method, system and device based on the Hamming sphere interval.
[0071] A computer-readable storage medium stores executable instructions, which, when executed by a processor, enable the processor to implement a deep hash image retrieval method, system and device based on Hamming sphere interval.
[0072] A non-transitory computer-readable storage medium stores a computer program, which, when executed by a processor, implements a deep hash image retrieval method, system and device based on Hamming sphere interval.
[0073] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features applied herein.
Claims
1. A deep hash image retrieval method based on Hamming sphere interval, characterized in that: include: Step S1: Acquire image data and establish a data set; Step S2: Calculate the Hamming sphere interval according to the number of categories in the data set The value range of Step S3: Select the appropriate , train the improved classification network model; use the binary update The value of , repeat the training model until the enumeration is exhausted All values of ; select the best model as the trained improved classification network model; Step S4: Input the query image into the trained improved classification network model to obtain a hash code, and use the Hamming distance obtained by performing an XOR operation on the hash code and the hash code of the image in the retrieval library to determine the similarity between the images.
2. The deep hash image retrieval method based on Hamming sphere interval according to claim 1 is characterized in that: Step S2: Calculating the Hamming sphere interval according to the number of categories in the data set The value range of includes: Step S21: Determine the number of categories in the data set and the size of the encoding space, where the size of the encoding space is the number of bits after hashing The number of states that can be represented is different hash codes; assign a Hamming sphere to each category; Step S22: When the code length is In the space of The volume of the Hamming sphere The definition is as follows: ; Step S23: Definition The minimum Hamming distance between the centers of two Hamming balls is , enumerate all the As The value range of 。 3. The deep hash image retrieval method based on Hamming sphere interval according to claim 2 is characterized in that: Step S3: Select a suitable , train the improved classification network model; use the binary update The value of , repeat the training model until the enumeration is exhausted All values of ; The best model is selected as the trained improved classification network model, including: Step S31: Use binary search to Determine the current value in the range The value of Step S32: input the data set into the improved classification network model for training, wherein the improved classification network model outputs the dense vector of the feature layer of the traditional classification network model A hash coding module was then added to convert Mapping to hash code , wherein the hash coding module includes: n hash coding layers, a fully connected layer and a batch normalization layer; the hash coding layer includes: a fully connected layer, a batch normalization layer and an activation function: Step S33: Update using binary search Repeat step S32 until the enumeration is complete. All values of ; select the optimal model as the trained improved classification network model.
4. The deep hash image retrieval method based on Hamming sphere interval according to claim 3 is characterized in that: In step S32, the hash coding module Mapping to hash code , the specific steps are: Step S321: The operation of the hash coding module is expressed as: ; in, It is the hash encoding module; Step S322: Update the value of the current hash center, where the hash center is defined as the point with the smallest Hamming distance to all feature points in the category: ; in, For the The hash code of the image, For the Category labels corresponding to the images; is the Hamming distance of the hash code; For the The value of the hash center of each category; For the A collection of elements of a category; in, It can be updated by the following formula: ; ; in, is the symbolic function, The contribution of each bit after mapping the data of the current category to the hash code; Step S323: Constructing the total loss based on the center loss and the contrast loss for: ; in, is a hyperparameter used to balance the weight of the central loss term; is the number of pictures in the current batch; is the similarity matrix. If two pictures are similar ,otherwise ; Step S324: Repeat steps S322 to S323 until the improved classification network model converges, completing this round of training.
5. A deep hash image retrieval system based on Hamming sphere interval, characterized in that: Includes the following modules: Build a dataset module to obtain image data and establish a dataset; The module for calculating the range of Hamming sphere intervals is used to calculate the Hamming sphere intervals according to the number of categories in the data set. The value range of Train the improved classification network model module to select the appropriate , train the improved classification network model; use the binary update The value of , repeat the training model until the enumeration is exhausted All values of ; select the best model as the trained improved classification network model; The image retrieval module is used to input the image to be queried into the trained improved classification network model to obtain a hash code, and to determine the similarity between images by the Hamming distance obtained by performing an XOR operation on the hash code and the hash code of the image in the retrieval library.
6. A deep hash image retrieval system based on Hamming sphere interval, characterized in that: The training improved classification network model module includes a hash coding module for converting the dense vector output by the feature layer of the traditional classification network model into Mapping to hash code .
7. A deep hash image retrieval device based on Hamming sphere interval, characterized in that: The method comprises one or more electronic devices, wherein the one or more electronic devices are used to implement the method according to any one of claims 1 to 4.
8. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 4.
9. A computer-readable storage medium, characterized in that: Executable instructions are stored thereon, and when the instructions are executed by a processor, the processor implements the method according to any one of claims 1 to 4.
10. A non-transitory computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.