A Large-scale Image Retrieval Method Based on Deep Multiple Negative Example Supervised Hashing

By using sample multiplexing strategy and multi-negative case loss function in the deep-supervised hashing method, multi-negative case tuples are constructed and convolutional neural network parameters are optimized, which solves the problems of large data usage and low training efficiency in the existing technology, and efficient image retrieval is achieved.

CN116069966BActive Publication Date: 2025-06-20XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211601313.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2025-06-20
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

When the existing deep-supervised hashing method constructs semantic similarity between images, the data usage is large and the training efficiency is low, making it difficult to effectively maintain the semantic relationship between multiple images.

Method used

The sample multiplexing strategy is used to build multi-negative example tuples, and the parameters of the convolutional neural network are optimized to improve the accuracy and efficiency of image retrieval through the new multi-negative example loss function and adaptive interval parameters.

Benefits of technology

It significantly improves image retrieval accuracy, reduces data usage, enhances optimization stability, and speeds up training convergence speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116069966B_ABST
    Figure CN116069966B_ABST
Patent Text Reader

Abstract

A large-scale image retrieval method based on deep multi-negative example supervised hashing. In each iteration process, a small number of images are randomly selected from the training set and fed into a convolutional neural network. The output of the last layer is used as the input of the hashing layer, and its hash value is calculated. The selected images are randomly divided into two equal parts. One part is used as query images. Using the supervision information, a similarity matrix is constructed for the above two parts of images, and on this basis, multi-negative example tuples are constructed for the query images. By calculating the multi-negative example loss function and performing stochastic gradient optimization on it, the network parameters are updated using the backpropagation algorithm. After completing the specified number of iterations, the trained network parameters and the binary hash code of each image are output. The present invention uses small batch data to train the network parameters and can make full use of the structural information contained in the training images, significantly improving the convergence speed and accuracy of the large-scale image retrieval method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image retrieval, and particularly to a large-scale image retrieval method based on deep multi-negative example supervised hashing. Background Art

[0002] The rapid development of digital technology, the Internet and the popularization of intelligent devices have brought a huge amount of visual data such as images and videos. For example, hundreds of millions of images are uploaded every day on major social, shopping and search platforms; a large number of monitoring devices generate a huge amount of monitoring videos all the time. It is particularly important to effectively manage and query these rich and intuitive visual information. Therefore, as an important technology for realizing the management and query of large-scale image databases, image retrieval has become a research hotspot. Image retrieval has a wide range of applications in the fields of commerce, medicine, security, etc., such as the search for product images on shopping websites; the query of similar medical diagnostic images; face recognition for access control, pedestrian recognition for monitoring, etc.

[0003] According to the different utilization methods of supervision information, the existing deep supervision hashing methods can be roughly divided into three categories: the pair-based method, the triplet-based method, and the category classification method. The pair-based method is the most common, mainly including methods such as convolutional neural network hashing (CNNH), deep supervised hashing (DSH), deep supervised discrete hashing (DSDH), and deep discrete supervised hashing (DDSH). This type of method selects image pairs and determines the similarity between images based on supervision information, so as to ensure that the hash codes can maintain the semantic similarity between each pair of images. The triplet-based method mainly includes network in network hashing (NINH), deep supervised ranking hashing (DSRH), deep triplet supervised hashing (DTSH), etc. This type of method expands the image pair into a triplet on the basis of the pair, while ensuring the similarity of the images within the triplet. Its advantage is that the supervision information is utilized more fully, and the accuracy is improved compared with the pair-based method. The methods based on category classification mainly include supervised semantic-preserving deep hashing (SSDH) and greedy hash (GH). This type of method cannot control the distance between hash codes of different categories. When the hash codes between different categories are small, it is easy to cause retrieval errors.

[0004] It can be seen that in the classification-based method, the training images do not interact with other images, and only need to ensure correct category classification; in the pair-based method, the images interact pairwise and maintain semantic similarity; in the triplet-based method, only three images interact and the hash codes maintain semantic relationships. The purpose of deep hashing is to maintain the semantic similarity between all images, and the above methods can be regarded as selecting two or three elements each time and maintaining the semantic relationship between them, and then approaching the overall semantic relationship we need through multiple iterations. Therefore, in each step of training, extracting more images to maintain similarity and using more supervision information can better approximate the overall semantic relationship. Thus, the retrieval accuracy can be improved, the training efficiency can be accelerated, and the convergence speed can be increased. Summary of the Invention

[0005] To overcome the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to propose a large-scale image retrieval method based on deep multi-negative example supervised hashing. By adopting a sample reuse strategy, multi-negative example tuples are efficiently constructed, reducing the amount of data used while ensuring the number of tuples. By adopting a new multi-negative example loss function, the semantic similarity between images in the multi-negative example tuples is maintained. By adaptively introducing a margin parameter into the cost function, the distance between dissimilar images is further increased, making the optimization more stable. The present invention adopts a batch optimization strategy to learn the parameters of the convolutional neural network.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A large-scale image retrieval method based on deep multi-negative example supervised hashing, marking X = {x1,..., x L} as the training image set, which contains L images, and marking the labels corresponding to X = {x1,..., x L} as Y = {y1,..., y L}; for any two images x i ∈X, x j ∈X, i≠j, if they are single-label images, then when x i and x j have the same label, then x i and x j are considered similar; if they are multi-label images, then when x i and x j have at least the same label, then x i and x j are considered similar; marking as the similarity matrix between training images, for any two images x i ∈X, x j ∈X, i≠j, if x i and x j are similar, then s ij = 1; otherwise, s ij = 0; including the following steps:

[0008] Step S1: Adopt a batch optimization strategy to learn the parameters of the convolutional neural network. Specifically, in each iteration, randomly select 2n images {x1,..., x n ,..., x 2n} from the training set X, and randomly divide them into two equal parts: g1 = {x1,..., x n} and

[0009] Step S2: Construct a similarity matrix S of g1 and g2 based on the supervision information. Take the images in g1 as query images, and construct multi-negative example tuples for the images in g1 based on the similarity matrix S. Without loss of generality, remove any query x in g1 i from the subscript, and mark it as x. For each query image x, mark as its corresponding multi-negative example tuple, where x + is the positive example (similar image) of x, is a negative example (dissimilar image), and there are m negative examples in total;

[0010] Step S3: Feed the images g1 and g2 into the convolutional neural network to obtain the output features u1, u2,..., u n and mark b i = h(u i ), where the function h(·) is defined as follows:

[0011]

[0012] Step S4: Calculate the partial derivatives of the multi-negative example loss function with respect to u i , respectively. Adopt the stochastic gradient optimization strategy and use the backpropagation algorithm to update the network parameters;

[0013] Step S5: Repeat Steps 1 to 4 to train the network. After repeating a fixed number of times, end the process, and output the trained network parameters and the binary hash code of each image.

[0014] The said Step S2 includes:

[0015] Step S21: Take the elements in g1 as queries. The elements in g2 may be positive examples, negative examples, or irrelevant elements. Mark s ij as the element in the i-th row and j-th column of the similarity matrix S. Based on the supervision information, construct the similarity matrix S as follows: If the i-th image in g1 is similar to the j-th image in g2, then s ij = 1; otherwise, s ij = 0;

[0016] Step S22: Construct multi-negative example tuples for the images in g1 based on the similarity matrix S. Specifically, if s ij = 1, then is the positive example of the query x i ; if s ij = 0, then is the negative example of the query x iNegative examples. Since there can only be one positive example in a multi-negative example tuple, if there are multiple positive examples, one needs to be selected as the positive example while ignoring the others. Specifically, if the image is a single-label image, a random image with the same label is selected as the positive example image. If the image is a multi-label image, the image with the largest number of identical labels is preferentially selected as the positive example image. If there are multiple images with the largest number of identical labels, one is randomly selected from them as the positive example image;

[0017] Step S23: Construct a positive example matrix P based on S. The matrix indicates which image is retained as the positive example. If p ij = 1, then is used as the query x i ∈g1's positive example. Obviously, only one element in each row of P is 1, and the rest are 0. According to the similarity matrix S and the positive example matrix P, the relationship between the images in g1 and g2 can be expressed as:

[0018]

[0019] Step S24: Mark the sets of the positive example, negative example, and irrelevant element subscripts of the query image x i as Λ i = {j|s ij = 1, p ij = 1}, Ω i = {k|s ik = 0, p ik = 0} and Υ i = {l|s il = 1, p il = 0}, where Λ i contains exactly one element. Then the multi-negative example tuple corresponding to the query image x i ∈g1 is

[0020] The said step S3 includes:

[0021] Step S31: Build the corresponding neural network. The convolutional neural network part is the pre-trained CNN-F. CNN-F is a convolutional neural network with fewer parameters, including 5 convolutional layers and 3 fully connected layers;

[0022] Step S32: Input g1 and g2 into the built neural network CNN-F. The output feature vector u = [u1, u2..., u n of the last fully connected layer and are used as the input of the following single hash layer. The hash layer calculates b i = h(u i ), Obtain vectors b and b r 。

[0023] The step S4 includes:

[0024] Step S41: Calculate the network loss according to the multi-negative example loss function provided by the present invention, that is, Equation (1), for the output of the corresponding multi-negative example tuple:

[0025]

[0026] Step S42: Calculate the derivative based on the derivatives of the multi-negative loss functions provided by Equation (2) and Equation (3), and update the parameters of the neural network through backpropagation:

[0027]

[0028]

[0029] Wherein, , Γ i ={j|s ji =1, p ij =1} and Δ i ={k|s ki =0, p ij =0} are the indexes of all similar and different images of the image ; s = 1 - m0|sgn(Φ ip )|, where m0 ∈ [0, 1) represents the scaling factor; η is a hyperparameter used to balance the contributions between the likelihood and the quantization error term.

[0030] Compared with the prior art, the advantages of the present invention are:

[0031] (1) By adopting the method of sample reuse to construct multi-negative example tuples for query images, it realizes obtaining more multi-negative example tuples by processing a small number of images in one batch, making full use of the structural information of the training image dataset and significantly improving the image retrieval accuracy;

[0032] (2) A multi-negative example loss function similar to Softmax is proposed, which allows the query to interact with multiple negative examples simultaneously and maintains the semantic similarity between the images in the multi-negative example tuples; by adaptively introducing a margin parameter in the loss function, the distance between dissimilar images is further increased, making the optimization more stable;

[0033] (3) Adopt the strategy of batch optimization to learn the parameters of the convolutional neural network. According to the derivative of the network output features with respect to the loss function, use backpropagation and stochastic gradient algorithms to optimize the network parameters, accelerating the convergence speed. Brief Description of the Drawings

[0034] Figure 1 It is a schematic flow diagram of the present invention.

[0035] Figure 2 It is the framework of the present invention; among them, (a) is the input of the network; (b) is the feature learning part; (c) is the multiple negative loss function. Specific implementation manners

[0036] The present invention will be further described in detail below with reference to the accompanying drawings.

[0037] A large-scale image retrieval method based on deep multi-negative example supervised hashing includes the following steps:

[0038] Embodiment 1

[0039] Step S1: Randomly obtain 5000 training data image sets (500 images for each class), 1000 test sets (100 images for each class) and corresponding label sets from CIFAR-10. The remaining 59000 images in CIFAR-10 are used for the retrieval database. Rescale the images in the data set to images of 224*224 pixels and perform normalization processing, and then as Figure 2 (a) shows, randomly select 2n training images from the training set, and randomly divide them into g1 = {x1,..., x n} and two sets;

[0040] Step S2: As the similarity matrix in Figure 2 (c), according to the supervision information of each image, establish a similarity matrix with g1 as columns and g2 as rows; the definition of similarity between two images in step S3 is: for single-label images, if the labels of the two images are the same, the two images are similar, and the corresponding position in the similarity matrix is 1, otherwise it is 0; this step uses the structural information in the data through the similarity matrix for subsequent selection of appropriate negative examples; use the elements in g1 as queries, and the elements in g2 may be positive examples, negative examples or irrelevant elements. Specifically, if s ij =1, then is a positive example of query x i , if s ij =0 then is a negative example of query x i . Since there can only be one positive example in a multi-negative example tuple, if there are multiple positive examples, one needs to be selected as the positive example and the other positive examples are ignored. Because the data is single-label, we randomly select one positive example. The definition of a positive example in step S31 is: for any image x in g1 i, any image similar to it in g2 is called a positive example, and vice versa is called a negative example; since the training data are all single-label data, a positive example is randomly selected and a similarity matrix P is constructed based on it;

[0041] Step S3: As Figure 2 (b) shows, build the corresponding neural network. The convolutional neural network part is the pre-trained CNN-F, followed by a single-layer hash layer. The weights of this layer are randomly initialized by a Gaussian distribution with a mean of 0 and a variance of 0.01. The final output is an n-dimensional feature vector; this step is used to convert the image into an n-dimensional vector, which is convenient for converting it into an n-dimensional hash code later. Here, a total of four comparison experiments with different lengths of hash codes are conducted, and n is 12, 24, 36, and 48 respectively. The network uses the SGD optimizer, the learning rate is set to 0.001, and the weight decay is 0.0005;

[0042] Step S31: Input g1 and g2 into the built neural network to obtain the corresponding feature vectors u1, u2..., u n and

[0043] Step S32: Input g1 and g2 into the built neural network CNN-F. Take the output feature vector u = [u1, u2..., u n and as the input of the following single hash layer. The hash layer calculates b i = h(u i ), to obtain vectors b and b r .

[0044] Step S4: Calculate the network loss according to the multi-negative example loss function provided by the present invention, that is, Equation (1), for the output of the corresponding multi-negative example tuple:

[0045]

[0046] Step S41: Calculate the derivative based on the multi-negative loss function provided by Equation (2) and Equation (3), and update the parameters of the neural network through backpropagation:

[0047]

[0048]

[0049] Among them, Γ i = {j|s ji = 1, p ij = 1} and Δ i = {k|s ki = 0, pij = 0} is the index of all similar and different images of image x i r ; s = 1 - m0|sgn(Φ ip ), where m0 ∈ [0, 1) represents the scaling factor; η is a hyperparameter used to balance the contributions between the likelihood and quantization error terms. In this experiment, m0 = 0.375 is set. When the input is a positive example, this parameter will reduce the loss, while when the input is a negative example, this parameter will increase the loss, ultimately resulting in a smaller distance between the same samples and a larger distance between different samples. η is also a hyperparameter used to balance the contributions between the likelihood and quantization error terms, and is set to 0.1 in this experiment.

[0050] Step S5: Repeat steps 1 to 4 to train the network, end after 50 epochs, and output the trained network and the binary hash code of each image.

[0051] Calculate the mAP as the experimental result. mAP represents the area under the Precision-Recall curve. The higher the score of similar images in the front position, the more comprehensive it can reflect the quality of the image retrieval result. When the hash code lengths are 12, 24, 36, and 48 respectively, the corresponding mAPs are 0.807, 0.839, 0.847, and 0.851, all higher than traditional hashing algorithms.

[0052] Example 2

[0053] Step S1: NUSE-WIDE is a multi-label scenario dataset, which is a subset of the Flicker dataset, containing approximately 270,000 images in total. These images have 81 labels, and each image can correspond to multiple labels. Select the 21 labels with the highest frequencies, and each label corresponds to at least 5,000 images. Select 500 images from each class as training images, 100 images as query images, and the images other than the retrieved ones as database images. Images with at least one same label are considered similar, otherwise not. Rescale the images in the dataset to 224*224 pixel images and perform normalization processing, and then as Figure 2 (a) shows, randomly select 2n training images from the training set and randomly divide them into g1 = {x1,..., x n} and two sets;

[0054] Step S2: As Figure 2(c) The similarity matrix is established with the supervision information of each image, taking g1 as columns and g2 as rows; the definition of similarity between two images in step S3 is: for multi-label images, if at least one label is the same between two images, then the two images are similar, and the corresponding position in the similarity matrix is 1, otherwise it is 0; this step utilizes the structural information in the data through the similarity matrix for subsequent selection of appropriate negative examples; taking the elements in g1 as queries, the elements in g2 may be positive examples, negative examples or irrelevant elements. Specifically, if s ij = 1, then is a positive example for query x i , if s ij = 0 then is a negative example for query x i . Since there can only be one positive example in a multi-negative example tuple, if there are multiple positive examples, one needs to be selected as the positive example and the others are ignored. Because the data is multi-label, we select the positive example with the largest number of labels in common with the query. The definition of a positive example in step S31 is: for any image x i in g1, any image similar to it in g2 is called a positive example, and vice versa is called a negative example; since the training data are all single-label data, a positive example is randomly selected and the similarity matrix P is constructed accordingly.

[0055] Step S3: As shown in Figure 2 (b), build the corresponding neural network. The convolutional neural network part is the pre-trained CNN-F, followed by a single-layer hash layer, the weights of which are randomly initialized by a Gaussian distribution with a mean of 0 and a variance of 0.01. The final output is an n-dimensional feature vector; this step is used to convert the image into an n-dimensional vector for subsequent conversion into an n-dimensional hash code. Here, a total of four comparison experiments with different lengths of hash codes are conducted, and n is 12, 24, 36, and 48 respectively. The network uses the SGD optimizer, the learning rate is set to 0.001, and the weight decay is 0.0005.

[0056] Step S31: Input g1 and g2 into the built neural network to obtain the corresponding feature vectors u1, u2..., u n and

[0057] Step S32: Input g1 and g2 into the built neural network CNN-F, and take the output feature vector u = [u1, u2..., u n and as the input of the subsequent single hash layer. The hash layer calculates b i = h(u i ), to obtain vectors b and b r .

[0058] Step S4: Calculate the network loss based on the outputs of the corresponding multi-negative example tuples according to the multi-negative example loss function provided by the present invention, i.e., Equation (1):

[0059]

[0060] Step S41: Calculate the derivative based on the multi-negative loss function provided by Equation (2) and Equation (3), and update the parameters of the neural network through backpropagation:

[0061]

[0062]

[0063] where Γ i ={j|s ji =1, p ij =1} and Δ i ={k|s ki =0, p ij =0} are the indices of all similar and different images of the image ; s = 1 - m0|sgn(Φ ip )|, where m0 ∈ [0, 1) represents the scaling factor; η is a hyperparameter used to balance the contributions between the likelihood and the quantization error term. In this experiment, m0 = 0.375 is set. When the input is a positive example, this parameter will reduce the loss, while when the input is a negative example, this parameter will increase the loss, ultimately resulting in a smaller distance between the same samples and a larger distance between different samples. η is also a hyperparameter used to balance the contributions between the likelihood and the quantization error term, and is set to 0.1 in this experiment.

[0064] Step S5: Repeat Steps 1 to 4 to train the network, and end after 50 epochs, output the trained network and the binary hash code of each image.

[0065] Calculate the mAP as the experimental result. The mAP represents the area of the Precision-Recall curve. The higher the score of the similar image is in the front position, the better it can comprehensively reflect the quality of the image retrieval result. When the hash code lengths are 12, 24, 36, and 48 respectively, the corresponding mAPs are 0.798, 0.827, 0.834, and 0.845, all higher than the traditional hashing algorithms.

[0066] Example 3

[0067] Step S1: The SVHN is similar to the handwritten character dataset MINIST. The difference is that this dataset collects color character images from house numbers, and the characters of the same class vary greatly, so the recognition difficulty is slightly higher. SVHN contains 10 classes, corresponding to the ten characters 0-9. Each class contains 5,000 training images and 1,000 test images. 500 images are randomly selected from each class as the training set, and 100 images are used as query images. All images except the test images are used as database images. The criterion for similarity (dissimilarity) between images is that the labels are the same (different). The images in the dataset are rescaled to images of 224*224 pixels and normalized, and then as Figure 2 (a) shows, randomly select 2n training images from the training set and randomly divide them into g1 = {x1,..., x n} and two sets;

[0068] Step S2: As the similarity matrix in Figure 2 (c), according to the supervision information of each image, a similarity matrix is established with g1 as columns and g2 as rows; the definition of similarity between two images in step S3 is: for single-label images, if the labels of the two images are the same, the two images are similar, and the corresponding position in the similarity matrix is 1, otherwise it is 0; this step utilizes the structural information in the data through the similarity matrix for subsequent selection of appropriate negative examples. The elements in g1 are used as queries, and the elements in g2 may be positive examples, negative examples, or irrelevant elements. Specifically, if s ij = 1, then is a positive example of the query x i , if s ij = 0 then is a negative example of the query x i . Since there can only be one positive example in a multi-negative example tuple, if there are multiple positive examples, one needs to be selected as the positive example and the others are ignored. Because the data is single-label, we randomly select one positive example. The definition of a positive example in step S31 is: for any image x i in g1, any image in g2 that is similar to it is called a positive example, and vice versa is called a negative example; since the training data are all single-label data, a positive example is randomly selected and the similarity matrix P is constructed accordingly;

[0069] Step S3: As Figure 2As shown in (b), a corresponding neural network is built. The convolutional neural network part is the pre-trained CNN-F, followed by a single-layer hash layer. The weights of this layer are randomly initialized with a Gaussian distribution having a mean of 0 and a variance of 0.01. The final output is an n-dimensional feature vector; this step is used to convert the image into an n-dimensional vector, facilitating the subsequent conversion into an n-dimensional hash code. Here, a total of four comparative experiments with hash codes of different lengths are conducted, where n is 12, 24, 36, and 48 respectively. The network uses the SGD optimizer, with the learning rate set to 0.001 and the weight decay to 0.0005;

[0070] Step S31: Input g1 and g2 into the built neural network to obtain the corresponding feature vectors u1, u2..., u n and

[0071] Step S32: Input g1 and g2 into the built neural network CNN-F. Take the output feature vector u = [u1, u2..., u n and as the input of the subsequent single hash layer. The hash layer calculates b i = h(u i ), to obtain the vectors b and b r .

[0072] Step S4: Calculate the network loss for the output of the corresponding multi-negative example tuples according to the multi-negative example loss function provided by the present invention, i.e., Equation (1):

[0073]

[0074] Step S41: Calculate the derivative based on the multi-negative loss function provided by Equation (2) and Equation (3), and update the parameters of the neural network through backpropagation:

[0075]

[0076]

[0077] where, Γ i ={j|s ji = 1, p ij = 1} and Δ i ={k|s ki = 0, p ij = 0} are the indices of all similar and different images of the image ; s = 1 - m0|sgn(Φ ip )|, where m0 ∈ [0, 1) represents The scaling factor; η is a hyperparameter used to balance the contributions between the likelihood and quantization error terms. In this experiment, m0 = 0.375 is set. When the input is a positive example, this parameter reduces the loss, while when the input is a negative example, this parameter enlarges the loss, ultimately resulting in a smaller distance between the same samples and a larger distance between different samples. η is also a hyperparameter used to balance the contributions between the likelihood and quantization error terms, and is set to 0.1 in this experiment.

[0078] Step S5: Repeat steps 1 to 4 to train the network, and end after 50 epochs. Output the trained network and the binary hash code of each image.

[0079] Calculate the mAP as the experimental result. The mAP represents the area under the Precision-Recall curve. The higher the score of similar images at the front position, the more comprehensive it can reflect the quality of the image retrieval result. When the hash code lengths are 12, 24, 36, and 48 respectively, the corresponding mAPs are 0.810, 0.839, 0.841, and 0.857, all of which are higher than traditional hashing algorithms.

Claims

1. A large-scale image retrieval method based on deep multi-negative example supervised hashing, denoted as X = {x1,..., x L}, which is the training image set containing L images, and the labels corresponding to X = {x1,..., x L} are Y = {y1,..., y L}; for any two images x i ∈X, x j ∈X, i≠j, if they are single-label images, then when x i and x j have the same label, x i and x j are considered similar; If they are multi-label images, then when x i and x j have at least the same label, then x i and x j are considered similar; label is the similarity matrix between training images. For any two images x i ∈ X, x j ∈ X, i ≠ j, if x i and x j are similar, then s ij = 1; Otherwise, there is s ij = 0; characterized in that it comprises the following steps: Step S1: Learn the parameters of the convolutional neural network using a batch optimization strategy. Specifically, in each iteration, randomly select 2n images {x1,...,x n ,...,x 2n} from the training set X, and randomly divide them into two equal parts: g1 = {x1,...,x n} and Step S2: Construct a similarity matrix S of g1 and g2 based on the supervision information. Take the images in g1 as query images, and construct multi-negative example tuples for the images in g1 based on the similarity matrix S. Without loss of generality, remove any query x in g1 i from its subscript, and label it as x. For each query image x, label it with its corresponding multi-negative example tuple, where x + is the positive example (similar image) of x, is a negative example (dissimilar image), and there are m negative examples in total; Step S3: Feed the images g1 and g2 into a convolutional neural network to obtain the output features u1, u2..., u of the last layer n and label b i = h(u i ), where the function h(·) is defined as follows: Step S4: Calculate the partial derivatives of the multi-negative example loss function with respect to u i , respectively. Adopt the stochastic gradient optimization strategy and use the backpropagation algorithm to update the network parameters; Step S5: Repeat Steps 1 to 4 to train the network. End after repeating a fixed number of times, and output the trained network parameters and the binary hash code of each image.

2. The large-scale image retrieval method based on deep multi-negative example supervised hashing according to claim 1, characterized in that The said Step S2 includes: Step S21: Use the elements in g1 as queries. The elements in g2 may be positive examples, negative examples, or irrelevant elements, and mark s ij is the element in the i-th row and j-th column of the similarity matrix S. Based on the supervision information, construct the similarity matrix S as follows: If the i-th image in g1 is similar to the j-th image in g2, then s ij = 1; otherwise, s ij = 0; Step S22: Construct multi-negative example tuples for the images in g1 based on the similarity matrix S. Specifically, if s ij = 1, then is a positive example of the query x i ; if s ij = 0, then is a negative example of the query x i . Since there can only be one positive example in a multi-negative example tuple, if there are multiple positive examples, one needs to be selected as the positive example while ignoring the others. Specifically, if the image is a single-label image, randomly select one image with the same label as the positive example image. If the image is a multi-label image, preferentially select the image with the largest number of the same labels as the positive example image. If there are multiple images with the largest number of the same labels, randomly select one of them as the positive example image; Step S23: Based on S, construct a positive example matrix P. The matrix indicates which images are retained as positive examples. If p ij = 1, then serves as a positive example for query x i ∈ g1. Obviously, only one element in each row of P is 1, and the rest are 0. Based on the similarity matrix S and the positive example matrix P, the relationship between the images in g1 and g2 can be expressed as: Step S24: Mark the query image x i The sets of subscripts of positive examples, negative examples, and irrelevant elements of i are Λ ij ={j|s ij =1, p i =1}, Ω ik ={k|s ik =0, p i =0} and Υ il ={l|s il =1, p i =0}, where Λ i contains exactly one element. Then the multi-negative example tuple corresponding to the query image x 3. The large-scale image retrieval method based on deep multi-negative example supervised hashing according to claim 1, characterized in that The said Step S3 includes: Step S31: Build the corresponding neural network. The convolutional neural network part is the pre-trained CNN-F. CNN-F is a convolutional neural network with fewer parameters, including 5 convolutional layers and 3 fully connected layers. Step S32: Input g1 and g2 into the built neural network CNN-F, and use the output feature vector u = [u1, u2..., u n and as the input to the single hash layer immediately following. The hash layer calculates b i = h(u i ), to obtain vectors b and b r .

4. The large-scale image retrieval method based on deep multi-negative example supervised hashing according to claim 1, characterized in that The said Step S4 includes: Step S41: Calculate the network loss according to the multi-negative example loss function provided by the present invention, that is, Equation (1), for the output of the corresponding multi-negative example tuple. Step S42: Calculate the derivative based on the derivatives of the multi-negative loss function provided by Equation (2) and Equation (3), and update the parameters of the neural network through backpropagation. Among them, Γ i ={j|s ji =1, p ij =1} and Δ i ={k|s ki =0, p ij =0} are the indices of all similar and dissimilar images of the image ; s = 1 - m0|sgn(Φ ip )|, where m0 ∈ [0, 1) represents the scaling factor; η is a hyperparameter used to balance the contributions between the likelihood and the quantization error term.

Citation Information

Patent Citations

  • Method and apparatus for learning sequential binary code using features

    KR1020140077409A

  • Image retrieval method based on variable-length deep hash learning

    WO2017092183A1