Image retrieval method based on locality sensitive hashing algorithm
By introducing soft winner-take-all mechanism and Herb-like learning rules into locally sensitive hash algorithms, SoftHash generates sparse and binary hash representations, solving the problem of the hard winner-take-all mechanism limiting semantic information capture in the existing technology, and improving the accuracy and performance of image retrieval.
Patent Information
- Application Number
- CN202311841185.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-12-28
AI Technical Summary
The existing locally sensitive hashing algorithms have performance limitations in image retrieval, especially BioHash adopts a hard winner-take-all mechanism, which limits its ability to capture semantic information of the input data, resulting in reduced retrieval accuracy.
A method of image retrieval based on locally sensitive hashing algorithm is proposed, called SoftHash. Combining Hebbe-like learning rules and soft-winner-take-all mechanism, it learns synaptic weights and neuronal deviations of mapping functions, thereby generating sparse, binary hash representations.
SoftHash allows weakly related data to have the opportunity to learn, generate more representative high-dimensional hash codes, improves the accuracy and performance of image retrieval, and enables fast and accurate retrieval on large data sets.
Smart Images

Figure CN120234436A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image retrieval, and particularly to an image retrieval method based on a locality-sensitive hashing algorithm. Background Art
[0002] Instance-level image retrieval is a visual search task, the goal of which is to retrieve all images containing the same object instance as a query image in a possibly very large image database. Image retrieval technology has been widely applied in many fields, such as image search engines, image copyright authentication, medical image analysis, etc. Recently, the work on image similarity retrieval has focused on large-scale high-dimensional problems. To reduce the computational complexity, a common method is to first reduce the dimensionality of the data and then effectively apply the nearest neighbor search or space partitioning method to the simplified data. To effectively reduce the dimensionality of a large-capacity data set, the locality-sensitive hashing method (LSH) has been widely used and has achieved quite successful results.
[0003] Locality-sensitive hashing has become a widely studied technique in computer science and is often used to solve image retrieval tasks. It maps the input features into binary codes, which helps to reduce the retrieval computation time. The goal of LSH is to closely cluster similar images together and separate dissimilar images far apart in the hash space. Classical LSH methods are usually used to transform high-dimensional features into low-dimensional binary codes to achieve fast and efficient nearest neighbor search. Recently, inspired by the olfactory circuit of Drosophila, some researchers have proposed a bio-inspired LSH called FlyHash, which performs sparse feature representation by transforming high-dimensional features into higher-dimensional binary codes. FlyHash is a data-independent hashing algorithm, in which the projection space is randomly constructed and cannot adapt to data streams. The learnable synaptic weight version of FlyHash is BioHash, which can update the synaptic weights of the hash network according to the input data and shows higher performance than classical FlyHash. However, BioHash adopts a hard winner-takes-all (WTA) mechanism for weight update, only allowing one winner neuron to be updated in each learning step, which greatly limits its ability to capture the semantic information of the input data.
[0004] Winner-take-all (WTA) is an important competitive mechanism in recurrent neural networks and a computational principle applied to neural network computational models. Neurons compete with each other for activation through this computational principle. WTA networks are commonly used in computational models of the brain, especially for distributed decision-making or action selection in the cortex. It can be used to develop feature selectivity through competition in simple recurrent networks. The classical form of WTA is hard WTA, in which only the neuron with the highest activation can be updated, while other neurons remain unchanged during the learning process. This kind of WTA is so crude that the input signal may not be fully utilized by the neurons, thus affecting the semantic representation of the trained network. Compared with hard WTA, soft WTA allows multiple neurons to be updated according to the input signal, which is more practical and effective in extracting semantic features from the input data.
[0005] To achieve fast retrieval in massive image data, Li et al. proposed Optimal Sparse Lifting (SOLHash), which learns the weights of the mapping function based on the training data to further improve the performance of hash representation. SOLHash consists of two key parts: Optimal sparse lifting is the sparse binary vector representation of the input image in the high-dimensional space, which can roughly maintain the pairwise similarity between images. The sparse lifting operator is a sparse binary matrix that best maps the input image to the optimal sparse lifting. To optimize the objective function, SOLHash uses Frank-Wolfe to update the parameters, where each learning update involves solving a constrained linear program involving all the training data, which is biologically unrealistic. From the perspective of computer science, the scalability of SOLHash is highly limited; not only does each update step call a constrained linear program, but the program also involves pairwise similarity matrices, which become prohibitively large for moderately sized data sets.
[0006] To improve the performance of image retrieval, Ryali et al. proposed a new locality-sensitive hashing algorithm, BioHash, inspired by FlyHash. Compared with previous work, this algorithm generates sparse high-dimensional hash codes in a data-driven manner and learns synapses in a neurobiologically plausible way. BioHash is a biologically plausible data-driven locality-sensitive hashing algorithm that combines the Hebbian rule with the hard WTA mechanism to learn synaptic weights from the data. The learning of BioHash can be formalized as minimizing the following energy function:
[0007]
[0008] where, and is the Lebesgue norm. The Rank operation sorts the inner products from largest to smallest, and
[0009]
[0010] From a computer science perspective, the implementation of BioHash is simple, scalable to large datasets, has good similarity search performance and is superior to deep hashing methods trained by backpropagation. However, the energy function of the locality-sensitive hashing algorithm BioHash is similar to the K-means clustering algorithm, where the distance from the input point to the centroid is measured by the normalized inner product rather than the Euclidean distance. Therefore, each input image can only belong to one neuron, as Figure 1 shown.
[0011] To achieve fast retrieval with binary image representation, SOLHash generates sparse binary output vectors for input image data by preserving the approximate pairwise similarity between the input image data. This is achieved by performing constrained linear programming that includes all training data at each learning step. However, in biological neural networks, neuron responses are adjusted through physically local synaptic change processes. BioHash uses a local learning rule (Hebbian) to learn the synaptic weights of the mapping function, and experiments show that BioHash has better retrieval performance than common LSH methods (such as classical LSH, data-driven hashing, deep hashing). However, when the hash length reaches a certain level, there are limitations in the BioHash representation. This may be attributed to the hard WTA mechanism adopted by BioHash, which only allows one neuron to be activated for any input image. This hard WTA mechanism greatly limits the learning and expression ability of the neural network to extract semantic information and reduces the accuracy of image retrieval. Summary of the Invention
[0012] In view of this, the present invention provides an image retrieval method based on a locality-sensitive hashing algorithm, including the following steps:
[0013] Step 1, obtain multiple image data, preprocess the multiple image data to obtain an image data set A with dense representation; represent an image data in the image data set A with a d-dimensional vector x (x ∈ A), x = [x1, x2,..., x i ,...x d ;
[0014] Step 2, initialize the SoftHash parameters; update the SoftHash parameters based on the image data set A;
[0015] wherein, the SoftHash parameters include: a synaptic weight matrix W and a neuron bias vector b, wherein, m > d;
[0016] Step 3: Given the image \(s\) with dense representation, use the updated synaptic weight matrix \(W\) and neuron bias vector \(b\) to obtain the binary representation \(v\) of the image \(s\) with hash length \(k\).
[0017] Step 4: Based on the obtained binary representation \(v\) of the image \(s\), perform nearest neighbor fast retrieval using Hamming distance to find the image with the same object instance as the query image.
[0018] Furthermore, the preprocessing in Step 1 includes:
[0019] Step 11: Use Gaussian convolution to perform blur denoising on multiple image data.
[0020] Step 12: Perform normalization on the multiple image data after blur denoising.
[0021] Furthermore, the update of SoftHash parameters in Step 2 includes:
[0022] Step 21: Given the input \(d -\)dimensional vector \(x(x\in A)\), map it to the \(m -\)dimensional space through the hash function, and the output is \(y\), where \(m>d\);
[0023] Synaptic weight The update formula is:
[0024]
[0025] where \(\tau\) is a hyperparameter and is a constant; \(\eta\) represents the learning rate, and its value ranges from 0 to 0.1; \(y\) j is used to represent the softmax output of the \(j -\)th neuron; \(x\) i is the \(i -\)th element of the \(d -\)dimensional vector \(x\); \(u\) j is the postsynaptic variable of the \(j -\)th output neuron;
[0026] The bias \(b\) of the \(j -\)th neuron j is updated as:
[0027] \(b\) j \(=b\) j-1 +\(\Delta b\) j
[0028] where \(b\) j-1 is the bias of the \((j - 1)-\)th neuron,
[0029] Step 22: Based on the synaptic weight and neuron bias \(b\) j , calculate the postsynaptic variable of the \(j -\)th neuron:
[0030]
[0031] Obtain its Bayesian posterior:
[0032] y j = Softmax(u j + b j );
[0033] Step 23, repeat Steps 21 and 22 to update the synaptic weight matrix W and the neuron bias vector b until a predetermined number of iterations is reached.
[0034] Furthermore, the binary representation v in Step 3 is specifically:
[0035]
[0036] where top k represents the top k of all outputs y j sorted from largest to smallest.
[0037] The image retrieval method based on the locality-sensitive hashing algorithm disclosed in the present invention uses the locality-sensitive hashing algorithm SoftHash to map the densely represented image to the binary space to obtain a new representation, thereby achieving fast retrieval through the Hamming distance; the SoftHash proposed in the present invention combines the Hebbian-like learning rule and the soft winner-takes-all mechanism to learn the synaptic weights and neuron biases of the mapping function so as to generate a sparse and binary hash representation; SoftHash allows weakly correlated data to have the opportunity to learn and can effectively generate more representative high-dimensional hash codes, improving the retrieval performance while achieving fast retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a schematic diagram of the clustering process of the BioHash hashing algorithm;
[0039] Figure 2 is a schematic diagram of the clustering process of the SoftHash hashing algorithm;
[0040] Figure 3 is a schematic diagram of the local Hebbian learning rule;
[0041] Figure 4 is a schematic diagram of the PN-KC-APL network architecture of the Drosophila mushroom body;
[0042] Figure 5 is a flowchart of the image retrieval method based on the locality-sensitive hashing algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The present invention discloses an image retrieval method based on the locality-sensitive hashing algorithm. Through the SoftHash algorithm, the image with dense representation is mapped into the binary space to obtain a new representation, so as to achieve fast retrieval through the Hamming distance. The energy function equation of SoftHash can be approximated as a mixture probability model, where each output neuron is regarded as an individual probability model. In this framework, the clustering process can be interpreted as the probability that each input image belongs to different neurons, as Figure 2 shown. Therefore, each input image can be associated with one or more neurons, so that the data can be represented more flexibly and meticulously. This probability modeling method in SoftHash can understand the semantic information in the input image data more richly and capture the complex patterns and associations in the data, thereby improving the expression ability and discriminability.
[0044] The SoftHash proposed by the present invention is a data-dependent hashing algorithm, which uses a Hebbian-like rule (as Figure 3 shown) to update the synaptic weights and neuron biases of the mapping function. This rule combines Bayes' theorem with a soft winner-takes-all (soft-WTA) mechanism with softmax smoothing to better learn the semantic information in the input; the present invention first generates a binary representation vector of the image through SoftHash, and the process is as Figure 4 shown by the PN-KC-APL neural network. The input picture generates its posterior probability into the KC neurons through the mapping function. The KC-APL recurrent network retains a small part of the activated KC assignment state 1, while the remaining inactive KCs are assigned state 0. Secondly, the nearest neighbor search is performed on these representation vectors using the Hamming distance to retrieve the images with the same object instance.
[0045] To achieve fast image retrieval, the position of the input image can be specified by observing the few nearest reference points selected from the hash space, generating a sparse and useful local representation. In this way, the nearest neighbor of the input image can be quickly found through the Hamming distance. The present invention models all output neurons as probability models to support the entire data distribution. The values on the output layer represent the probabilities that each input image belongs to them, which are calculated based on the input image data, synaptic weights, and neuron biases. The plasticity rule of the synaptic weights and the iterative update of the neuron biases in the probability model make their density match the density of the input image data. In other words, where the data density is high, more output neurons are needed to provide high resolution; where the data density is low, fewer neurons are needed. In this way, the SoftHash algorithm has sufficient ability to activate similar output neurons for similar inputs and closely cluster them in a new dimension-expanded hash space. Therefore, the SoftHash algorithm proposed by the present invention can learn the locality of the input representation.
[0046] Mathematically speaking, the goal of the image retrieval method based on the locality-sensitive hashing algorithm proposed by the present invention is to achieve fast and more accurate retrieval of images while representing the semantic information of images with binary values (0 and 1). The present invention learns the hash mapping function through a Hebbian-like learning rule and the idea of winner-takes-all (WTA); specifically, the synaptic weights of the mapping function are updated based on the local learning rule, that is, the change in the weights only depends on the activities of the presynaptic and postsynaptic neurons, which is biologically reasonable. On the other hand, different from previous studies that adopted the hard WTA rule, the present invention introduces a soft WTA rule, in which non-winning neurons are not completely inhibited during the learning process. This gives weakly related data the opportunity to be learned to generate more representative hash codes, thereby improving the accuracy of image retrieval.
[0047] Figure 5 Shows the flow of the image retrieval method based on the locality-sensitive hashing algorithm of the present invention. As Figure 5 shown, the method includes the following steps:
[0048] Step S1: Obtain multiple image data, preprocess the multiple image data, perform blur denoising on the multiple images using Gaussian convolution and then normalize them to obtain a dense-representation image dataset A, and represent an image data in the image dataset A with a d-dimensional vector x (x ∈ A), denoted as x = [x1, x2,..., x i ,...x d .
[0049] Step S2: Initialize the parameters of SoftHash; based on the image dataset A, update the parameters of SoftHash: the synaptic weight matrix W and the neuron bias vector b.
[0050] Step S2-1: Given an input d-dimensional vector x (x ∈ A), map it to an m-dimensional space (m > d) through a hash function, and the output is y, where y is an m-dimensional vector, expressed as The update formula of the synaptic weight based on the class Hebbian rule and the soft winner-takes-all mechanism is ( Figure 3 ):
[0051]
[0052] where τ is a hyperparameter, a constant; η represents the learning rate, taking values from 0 to 0.1; y j is used to represent the softmax output of the j-th neuron; x i is the i-th element of the d-dimensional vector x; u j is the postsynaptic variable of the j-th output neuron.
[0053] The update method of the j-th neuron bias b j is as follows:
[0054] b j = b j-1 + Δb j
[0055] where b j-1 is the bias of the (j - 1)-th neuron,
[0056] Step S2-2: Based on the synaptic weight and the neuron bias b j , calculate the postsynaptic variable of the j-th neuron, and thus further obtain its Bayesian posterior y j = Softmax(u j + b j ) of the calculation result.
[0057] Step S2-3: Based on the dense-representation image dataset A, repeat Steps S2-1 and S2-2 to update the parameters until the specified number of iterations is reached.
[0058] Step S3: Given a dense-representation image s, use the synaptic weight matrix W and the neuron bias vector b updated in Step S2 to generate a binary representation v of length k for the image s, and the method is as follows:
[0059]
[0060] Among them, top k represents the top k outputs y j sorted in descending order.
[0061] Step S4: Based on the binary representation of the image s, use the Hamming distance for nearest neighbor fast retrieval to find the image with the same object instance as the query image.
[0062] Next, two datasets, Fashion-MNIST and CIFAR10, are used to verify the performance of the image retrieval method based on SoftHash.
[0063] Table 1 Comparison results of mAP@1000 of different hashing algorithms on Fashion-MNIST and CIFAR10 datasets
[0064]
[0065] The mAP@1000 results of different hashing algorithms on the two benchmarks are shown in Table 1, where the hash length varies from 2 to 128. It can be clearly seen from Table 1 that among all hashing algorithms, the image retrieval method based on SoftHash-2 shows the best retrieval performance at large hash lengths, especially when k = 128. As the hash length k increases, the results obtained by SoftHash-2 are significantly better than those of SoftHash-1. Comparing the results of FlyHash and ConvHash, it can be seen that at different hash lengths, the superiority of FlyHash over the classical algorithm ConvHash is very obvious, demonstrating the good performance of high-dimensional sparse binary representation. BioHash-1 and BioHash-2 show the best retrieval performance at small hash lengths, but provide a small improvement from k = 32 to k = 64 and an even smaller improvement from k = 64 to k = 128, which is consistent with the conclusions in their work. In addition, it can be observed that both SoftHash-2 and BioHash-2 with synaptic weights initialized according to the standard normal distribution can obtain better results.
[0066] The experimental results show that the image retrieval method based on the locality-sensitive hashing algorithm (SoftHash) proposed in the present invention significantly outperforms the retrieval methods based on other hashing algorithms. Compared with the classical and other bio-inspired locality-sensitive hashing algorithms, SoftHash has better semantic representation ability in its binary representation.
[0067] In summary, the image retrieval method based on the locality-sensitive hashing algorithm (SoftHash) proposed by the present invention has better retrieval accuracy compared with the existing image retrieval methods based on the hashing algorithm. The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements for some technical features; without departing from the spirit of the technical solutions of the present invention, they should all be covered within the scope of the technical solutions claimed by the present invention.
Claims
1. An image retrieval method based on the locality-sensitive hashing algorithm, characterized in that The method includes the following steps: Step 1, obtain multiple image data, preprocess the multiple image data to obtain an image data set A with dense representation; represent an image data in the image data set A with a d-dimensional vector x (x ∈ A), x = [x1, x2,..., x i ,...x d ; Step 2, initialize the SoftHash parameters; Based on the image dataset A, update the SoftHash parameters; Among them, the SoftHash parameters include: the synaptic weight matrix W and the neuron bias vector b, b = [b1, b2,... b j ,... b m , where m > d; Step 3, given the image s with dense representation, use the updated synaptic weight matrix W and neuron bias vector b to obtain the binary representation v of the image s with a hash length of k; Step 4, based on the obtained binary representation v of the image s, perform nearest neighbor fast retrieval using the Hamming distance to find the image with the same object instance as the image to be retrieved.
2. The image retrieval method according to claim 1, wherein The preprocessing of Step 1 includes: Step 11, perform blurring and denoising processing on multiple image data using Gaussian convolution; Step 12, perform normalization processing on the multiple image data after blurring and denoising.
3. The image retrieval method according to claim 1, wherein The update of the SoftHash parameters in Step 2 includes: Step 21, given an input d-dimensional vector x (x ∈ A), it is mapped to an m-dimensional space through a hash function, and the output is y, y = [y1, y2,... y j ,... y m ; Synaptic weight The update formula is as follows: Among them, τ is a hyperparameter and is a constant; η represents the learning rate, with a value range of 0 - 0.1; y j is used to represent the softmax output of the j-th neuron; x i is the i-th element of the d-dimensional vector x; u j is the postsynaptic variable of the j-th output neuron; The bias b of the j-th neuron j is updated as follows: b j = b j-1 + Δb j where b j-1 is the deviation of the (j - 1)-th neuron, Step 22, based on the synaptic weights and the neuron bias b j , calculate the postsynaptic variable of the j-th neuron: Obtain its Bayesian posterior; y j = Softmax(u j + b j ); Step 23, repeat Steps 21 and 22 to update the synaptic weight matrix W and neuron bias vector b until a predetermined number of iterations is reached.
4. The image retrieval method according to claim 1, characterized in that The binary representation v in Step 3 is specifically: Among them, top k represents the top k outputs y j sorted in descending order.
Citation Information
Patent Citations
Large-scale image library retrieval method based on local similarity hash algorithm
CN104199922A
High efficiency clustering method based on locality-sensitive hashing and non-parametric Bayes method
CN106228035A
HTM sequence data analysis system and method based on locality sensitive hashing
CN114169518A
Image retrieval method based on combination of deep convolutional neural network and locality sensitive hash algorithm
CN116861022A
Image similarity search via hashes with expanded dimensionality and sparsification
US20190171665A1