A degraded image retrieval method
Patent Information
- Application Number
- CN202310262520.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-17
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-03-17
AI Technical Summary
[0004]为了解决上述问题,本发明提出了一种大规模降质图像鲁棒检索方法,该方法突破了因图像模糊或其他降质操作导致的图像检索质量低、图像数据急剧增长引起的大数据检索效率低等技术难点,有效解决了目前图像数据挖掘中日益严重的大规模降质图像鲁棒检索性能差的问题
Smart Images

Figure CN116226428B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical fields of data mining and large-scale image retrieval, and in particular to a degraded image retrieval method, which relates to a novel robust retrieval method for large-scale degraded images. Background Technology
[0002] Due to uncertainties arising during image capture, various types of degraded images are prevalent in different scenarios, such as blurred instances, blurred foregrounds, and partially occluded areas. Currently, degraded images play a crucial role in an increasing number of practical applications and emergency situations, most commonly in implementing similarity retrieval within large datasets to identify and retrieve similar objects. For example, when someone sees something interesting, such as a speeding car, they might eagerly take out their smartphone to take a picture and use the image to search for similar products on a search engine to learn about their models, specifications, etc. These typical practical applications highlight the importance of large-scale image retrieval of degraded images.
[0003] Image retrieval, the process of retrieving similar images from an image database based on a given query image, has become a mature technology applied in various practical scenarios. However, the increasing number of noisy images in image databases presents new challenges to image retrieval. To address this, many innovative works have been proposed to enhance the robustness of image retrieval, achieving good performance in post-processed noisy images (e.g., those with geometric transformations and Gaussian noise). Unlike noisy images synthesized later, degraded images suffer from semantic loss or blurring during the imaging process, resulting in reduced semantic quality. Retrieving these types of images was not previously considered, leading to low retrieval quality in degraded image retrieval. Furthermore, given the exponential growth of image data, improving the retrieval efficiency of large-scale degraded image data is another pressing challenge. Summary of the Invention
[0004] To address the aforementioned issues, this invention proposes a robust retrieval method for large-scale degraded images. This method overcomes the technical challenges of low image retrieval quality caused by image blurring or other degrading operations, and low efficiency of large-scale data retrieval due to the rapid growth of image data. It effectively solves the increasingly serious problem of poor robust retrieval performance for large-scale degraded images in current image data mining.
[0005] To achieve the objectives of this invention, the following technical solution is adopted:
[0006] A method for retrieving degraded images includes the following steps:
[0007] Step 1: Train a deep hashing network model;
[0008] Step 2: Input the image into a deep hashing network model to learn the latent hash code represented by a binary vector;
[0009] Step 3: By introducing a minimum hash algorithm, the potential hash code represented by the binary vector is converted into a minimum hash code;
[0010] Step 4: Use the hash collision principle to retrieve degraded images.
[0011] The degraded image retrieval method includes step 1: 1.1 training sample data augmentation; 1.2 training sample data grouping; 1.3 triplet training sample data selection.
[0012] The degraded image retrieval method, wherein step 1.1 includes: performing various data augmentation methods on all images in the training sample dataset, including rotation, flipping, blurring, sharpening, and denoising.
[0013] The degraded image retrieval method, wherein step 1.2 includes: copying the original training sample dataset multiple times to obtain a sample dataset group G as a candidate benchmark sample group. a The sample dataset obtained after augmenting the original training sample data in step 1.1 will be combined into a candidate positive sample group G. p ; Make a copy of the candidate benchmark sample group G a Furthermore, the sample numbers are shuffled, and the resulting sample dataset is grouped into candidate negative sample group G. n .
[0014] The aforementioned degraded image retrieval method, wherein step 1.3 includes:
[0015] (1) In the candidate benchmark sample group G a A single sample of data is randomly selected without repetition as the baseline sample.
[0016] (2) In the candidate positive sample group G p Select one that is similar to the benchmark sample Samples with the same serial number are considered positive samples.
[0017] (3) In the candidate negative sample group G n A sample is randomly selected as the negative sample. Next, the baseline sample is calculated. Positive samples and negative samples The relative distance.
[0018] Specifically, first calculate and Differential hashing yields the DIFF file. i a ,dif i p and Calculate DIF again i a and Hamming distance d1,dif i a and dif i p The Hamming distance d2 is calculated, and finally the difference ε between d1 and d2 is calculated:
[0019]
[0020] Next, given a boundary threshold δ between positive and negative sample pairs, it is determined whether ε is less than δ. If it is less than δ, the sample is selected as a high-quality negative sample; otherwise, a new sample is selected from the candidate negative sample group as the negative sample. (w≠j), recalculate the baseline sample Positive samples and negative samples The relative distance ε is used to re-evaluate whether ε is less than δ, and this process is repeated until a negative sample with ε less than δ is found.
[0021] (4) Select the baseline sample Positive samples and corresponding high-quality negative samples Form a high-quality training triplet And put it into the high-quality training triple set. In this context, l represents the number of high-quality triplets.
[0022] (5) Repeat steps (1)-(4) until the baseline sample group G is reached. a There is no sample data available.
[0023] The degraded image retrieval method further includes step 1: 1.4 depth encoding; 1.5 hash code quantization; 1.6 calculating model loss and updating network parameters.
[0024] The degraded image retrieval method, wherein step 1.6 includes: (1) calculating the quantization loss L q (2) Calculate the discriminative loss L d (3) Calculate the network model loss L; (4) Calculate the loss gradient based on L, and backpropagate the gradient in the network model to update the parameters of each network layer in the network model.
[0025] The aforementioned degraded image retrieval method, wherein step 1.6 of calculating the quantization loss includes:
[0026] (1.1) Divide all triplet training samples in the kth batch into two halves, the first half denoted as k1 and the second half denoted as k2.
[0027] (1.2) Calculate the depth feature vector of each reference sample in k1 one by one. With each baseline sample depth feature vector in k2 Cosine similarity between and the binary vector of each benchmark sample in k1 With each benchmark sample binary vector in k2 Jaccard similarity between And calculate and The mean squared loss is the benchmark sample quantization loss.
[0028]
[0029] Among them, S jac S represents the Jaccard similarity. cos This represents the cosine similarity.
[0030] (1.3) Calculate the depth feature vector of each positive sample in k1 one by one. With each positive sample depth feature vector in k2 Cosine similarity between and the binary vector of each positive sample in k1 With each positive sample binary vector in k2 Jaccard similarity between And calculate and The mean squared loss is the positive sample quantization loss.
[0031]
[0032] (1.4) Calculate the depth feature vector of each negative sample in k1 one by one. With each negative sample depth feature vector in k2 Cosine similarity between and the binary vector of each negative sample in k1 With each negative sample binary vector in k2 Similar to Jaccard And calculate and The mean squared loss is the negative sample quantization loss.
[0033]
[0034] (1.5) Calculate the quantization loss of the benchmark sample Positive Sample Quantization Loss and negative sample quantization loss The arithmetic mean of the quantization loss L. q :
[0035]
[0036] The degraded image retrieval method, wherein step 2 includes:
[0037] Input image x into the input layer and extract its high-dimensional feature representation f;
[0038] The high-dimensional feature representation f is input into the convolutional layer to learn the high-level semantic information of the image and output the high-dimensional deep feature representation g of the image.
[0039] Input a high-dimensional deep feature representation g into a fully connected layer, perform a non-linear transformation operation while preserving semantic similarity, and output a deep feature representation z;
[0040] The deep feature representation z is input into a nonlinear binary network layer, and the deep feature representation is converted into a potential hash code b in binary vector representation.
[0041] The degraded image retrieval method, wherein step 3 includes: calculating the potential hash code represented by the binary vector output in step 2 using a minimum hash function, and finally converting it into a minimum hash code h with efficient query advantages:
[0042]
[0043] in The minimum hash function is represented by the minimum hash algorithm. The calculation process is to perform several random row permutations π on the potential hash code b represented by a given binary vector. The signature vector composed of the minimum number of rows with a value of 1 after each permutation is the minimum hash value.
[0044] The degraded image retrieval method, wherein step 4 includes:
[0045] Calculate the minimum hash code for all images, including the degraded image to be queried and the images in the database;
[0046] Divide the minimum hash code (or minimum hash vector) h into s groups, each group being a vector component occupying r = len(h) / s bits;
[0047] Design a random mapping function that can map the vector components consisting of r bits in a group to hash buckets;
[0048] The same mapping function is used to compute all groups, but a separate array of hash buckets is used for each group, so even the same column vectors in different groups will not be hashed into the same hash bucket;
[0049] Check all database images that have a certain group mapping to the same hash bucket as the degraded image to be queried, and include these database images in the candidate similar image set;
[0050] Calculate the Jaccar similarity between the degraded image to be queried and each image in the candidate similar image set;
[0051] Sort the images by similarity from highest to lowest, and select the K images with the highest similarity to find the K images that are most similar to the degraded image to be queried. Attached Figure Description
[0052] Figure 1 This is a schematic diagram of the degraded image retrieval method of the present invention;
[0053] Figure 2 Flowchart for training a deep hash network model;
[0054] Figure 3 Flowchart for selecting triplet sample data. Detailed Implementation
[0055] The following is in conjunction with the appendix Figure 1-3 The specific embodiments of the present invention will be described in detail below. These embodiments are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. Obviously, the embodiments described in this invention are only a part of the embodiments of the present invention, and not all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0056] The terms "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include the specific features, structures, or characteristics described in connection with that embodiment. Therefore, the terms "comprising," "including," "having," and variations thereof in this specification mean "including but not limited to," unless otherwise specifically emphasized.
[0057] like Figure 1-3 As shown, the degraded image retrieval method of the present invention includes the following steps:
[0058] Step 1: Train a deep hashing network model;
[0059] Step 2: Input the image into a deep hashing network model to learn the latent hash code represented by a binary vector;
[0060] Step 3: By introducing the minimum hash algorithm, the potential hash code represented by the binary vector is converted into the minimum hash code.
[0061] Step 4: Use the hash collision principle to retrieve degraded images.
[0062] Step 1: Train the deep hashing network model
[0063] 1.1 Training Sample Data Augmentation
[0064] This invention implements various data augmentation methods on the original dataset during the training sample data preprocessing stage. The purpose is to expand the training samples, explore the learning potential of sample diversity, and improve the generalization ability of the deep hash network model. The deep hash network model mainly includes a deep encoder and a nonlinear binary network.
[0065] The detailed data augmentation process is as follows:
[0066] For image datasets, such as all image data from CIFAR-10, including both training and testing data, various data augmentation techniques are applied, such as rotation, flipping, blurring, sharpening, and denoising. This is illustrated using any image x from CIFAR-10 as an example:
[0067] Rotation: Rotate the image x by 45 degrees clockwise, changing the orientation of the image content and outputting...
[0068] Flip: Flips the image x along the horizontal direction, outputting...
[0069] Blur: Apply mean blur to image x, that is, under the action of a one-dimensional convolution kernel, shift by 35 bits in both the x and y directions, and output...
[0070] Sharpening: Convolves the image x with a custom kernel, that is, applies a linear filter to the image x, and outputs...
[0071] Denoising: Similar regions are found in image x, organized by image patch. The average of these regions is then calculated to remove Gaussian noise from image x. The output is...
[0072] 1.2 Grouping of Training Sample Data
[0073] This invention first copies the original dataset multiple times. Taking image x as an example, if it is copied 5 times, then the copied sample is denoted as... and The replicated dataset serves as the candidate benchmark sample group G. a Next, the data obtained after performing five data augmentations on the images in the original dataset as described in step 1.1 will be... and The dataset consisting of ) is used as the candidate positive sample group G p Finally, duplicate the candidate benchmark sample group G. a And shuffle its sample numbers, and then use the data obtained after this operation ( and They may correspond to as follows: and The dataset consisting of ) is used as candidate negative sample group G n .
[0074] 1.3 Selection of Triple Training Sample Data
[0075] The triplet training sample data selection procedure consists of five steps:
[0076] 1. In candidate benchmark sample group G a A single sample of data is randomly selected without repetition as the baseline sample.
[0077] 2. In the candidate positive sample group G p Select one that is similar to the benchmark sample Samples with the same serial number are considered positive samples.
[0078] 3. In candidate negative sample group G n A sample is randomly selected as the negative sample. Next, the baseline sample is calculated. Positive samples and negative samples The relative distance.
[0079] Specifically, first calculate and Differential hashing yields... and Recalculate and The Hamming distance d1, and The Hamming distance d2 is calculated, and finally the difference ε between d1 and d2 is calculated:
[0080]
[0081] Next, given a boundary threshold δ between positive and negative sample pairs, it is determined whether ε is less than δ. If it is less than δ, the sample is selected as a high-quality negative sample; otherwise, a new sample is selected from the candidate negative sample group as the negative sample. (w≠j), recalculate the baseline sample Positive samples and negative samples The relative distance ε is used to re-evaluate whether ε is less than δ, and this process is repeated until a negative sample with ε less than δ is found.
[0082] 4. Select the baseline sample Positive samples and corresponding high-quality negative samples Form a high-quality training triplet And put it into the high-quality training triple set. In this context, l represents the number of high-quality triplets.
[0083] 5. Repeat steps 1, 2, 3, and 4 until the baseline sample group G is reached. a There is no sample data available.
[0084] 1.4 Depth Coding
[0085] The above-selected training triplet set The training samples (where l represents the number of high-quality triples) are divided into batches τ1, τ2, ..., τ3 according to the number of triples in each batch of N. k ,...,τ l / N The input is fed into a deep encoder built on a deep neural network to learn the deep feature representation of the image. The deep encoder consists of an input layer, convolutional layers, and fully connected layers. The specific encoding process is as follows:
[0086] 1. Take all triplet training samples in the k-th batch. Where 1≤k≤l / N, these are sequentially input into the input layer for preprocessing. The specific steps of the preprocessing operation are as follows: First, perform an image reading operation on the triples, that is, open the image handle and instantiate an Image object. Next, image transformation operations are performed on the Image object. These operations, in sequence, include: changing the image channel specifications, normalization, reconstructing the data distribution, and adjusting the size of the Image object. After these preprocessing operations, a high-dimensional feature representation <f> of the triplet training samples can be extracted. i a ,f i p ,f i n > k .
[0087] 2. Represent the high-dimensional features of the triplet training samples <fi a ,f i p ,f i n > k The input is fed into a convolutional neural network (CNN) to perform convolution operations, learning high-level semantic information of the image. Here, the CNN uses a VGG16 pre-trained model as its backbone, and then removes the last softmax layer and the fully connected (FC) layer from the network structure. The formal representation of the convolution operation is as follows:
[0088]
[0089] Where ReLU(·) represents the activation function; conv3d is a three-dimensional convolution operation, which slides the convolution kernel across the input value, calculates at each position, and generates the output value; It is a three-dimensional matrix representation of a convolution kernel used for convolution calculations; It is the convolutional layer bias vector, representing the constant bias applied to each neuron in the network.
[0090] 3. Represent the high-dimensional deep features obtained after the convolution operation. The input is fed into a fully connected network layer, where non-linear dimensionality reduction is performed to obtain a deep feature representation while preserving semantic similarity. The formal representation is as follows:
[0091]
[0092] Among them W i fc This is the weight matrix of the fully connected layer, representing the weights of the connections to the input values. This is the bias vector for the fully connected layer, representing the constant bias applied to each neuron in the network.
[0093] 1.5 Hash Code Quantization
[0094] To facilitate image data computation and storage, we represent the depth features encoded by the depth encoder. The input is fed into a nonlinear binary network, and the output is a latent hash code, which is called a binary vector representation. The nonlinear binary network function is expressed as:
[0095]
[0096] Where sgn(·) represents the sign function, which can transform the deep feature vector z i The j-th component Mapped to -1 or 1, b i The j-th component.
[0097] 1.6 Calculate model loss and update network parameters
[0098] The steps are as follows:
[0099] (1) Calculate the quantization loss
[0100] (1.1) Divide all triplet training samples in the kth batch into two halves, the first half denoted as k1 and the second half denoted as k2.
[0101] (1.2) Calculate the depth feature vector of each reference sample in k1 one by one. (0<u≤N / 2) and the depth feature vector of each benchmark sample in k2 Cosine similarity between (N / 2 < v ≤ N) and the binary vector of each benchmark sample in k1 With each benchmark sample binary vector in k2 Jaccard similarity between And calculate and The mean squared loss is the benchmark sample quantization loss.
[0102]
[0103] Among them, S jac S represents the Jaccard similarity. cos This represents the cosine similarity.
[0104] (1.3) Calculate the depth feature vector of each positive sample in k1 one by one. With each positive sample depth feature vector in k2 Cosine similarity between and the binary vector of each positive sample in k1 With each positive sample binary vector in k2 Jaccard similarity between And calculate and The mean squared loss is the positive sample quantization loss.
[0105]
[0106] (1.4) Calculate the depth feature vector of each negative sample in k1 one by one. With each negative sample depth feature vector in k2 Cosine similarity between and the binary vector of each negative sample in k1 With each negative sample binary vector in k2 Similar to Jaccard And calculate and The mean squared loss is the negative sample quantization loss.
[0107]
[0108] (1.5) Calculate the quantization loss of the benchmark sample Positive Sample Quantization Loss and negative sample quantization loss The arithmetic mean of the values is the quantization loss L. q :
[0109]
[0110] (2) Calculate the binary vector of the benchmark sample in the kth batch. With positive sample binary vector Jaccard distance and the binary vector of the reference sample With negative sample binary vector Jaccard distance Next calculation and The triplet loss is also known as the discriminative loss L. d :
[0111]
[0112] d jac =1-S jac δ represents and The minimum difference between them, which is 0.5 here.
[0113] (3) Calculate the sample quantization loss L using weighted average. q And sample discriminative loss L d The final loss calculated is the network model loss L:
[0114] L = L q +αL d
[0115] α represents the loss weight hyperparameter, which is set to 0.1 here.
[0116] (4) Calculate the loss gradient based on L, and backpropagate the gradient in the network model to update the parameters of each network layer in the network model until all triplet training samples in all batches have completed the above operations 1.4, 1.5 and 1.6.
[0117] Step 2: Input the image into a deep hashing network model to learn and generate a potential hash code in binary vector representation.
[0118] The image is input into the deep hashing network model trained in step 1, which outputs a latent hash code represented by a binary vector with high semantic similarity and discriminative power. The detailed process is as follows:
[0119] 1. Input image x into the input layer and extract its high-dimensional feature representation f;
[0120] 2. The high-dimensional feature representation f is input into the convolutional layer to learn the high-level semantic information of the image and output the high-dimensional deep feature representation g of the image;
[0121] 3. Input the high-dimensional deep feature representation g into the fully connected layer. While maintaining semantic similarity, perform nonlinear transformation operations and output the deep feature representation z.
[0122] 4. The deep feature vector z is input into a nonlinear binary network layer, which converts the deep feature representation into a potential hash code b represented by a binary vector.
[0123] Step 3: Introduce the minimum hash algorithm to convert the potential hash code represented by the binary vector into the minimum hash code.
[0124] Without significantly affecting or minimizing the hash representation capability, the potential hash code represented by the binary vector output from step 2 is calculated using a minimum hash function, ultimately transforming it into a minimum hash code h with efficient query capabilities. The calculation process is represented as follows:
[0125]
[0126] in The minimum hash function is represented by the minimum hash algorithm. The calculation process is to perform several random row permutations π on the potential hash code b represented by a given binary vector. The signature vector composed of the minimum number of rows with a value of 1 after each permutation is the minimum hash value.
[0127] Step 4: Use the hash collision principle to complete (large-scale) image (robust) retrieval.
[0128] 1. Calculate the minimum hash code for all images, including the degraded image to be queried and the images in the database;
[0129] 2. Divide the minimum hash code (or minimum hash vector) h into s groups, each group being a vector component occupying r = len(h) / s bits;
[0130] 3. Design a random mapping function that can map the vector components consisting of r bits in a group to hash buckets;
[0131] 4. The same mapping function is used to calculate all groups, but a separate hash bucket array is used for each group. Therefore, even the same column vectors in different groups will not be hashed into the same hash bucket.
[0132] 5. Check all database images that have a certain group mapped to the same hash bucket as the degraded image to be queried, and include these database images in the candidate similar image set;
[0133] 6. Calculate the Jaccard similarity between the degraded image to be queried and each image in the candidate similar image set;
[0134] 7. Sort the images by similarity from highest to lowest, and select the K images with the highest similarity to find the K images that are most similar to the degraded image to be queried.
[0135] This invention can fully explore the potential information of training samples, improve the generalization ability of retrieval methods to cope with various degradation operations; it can greatly alleviate the huge human annotation cost caused by the need for a large amount of data with supervision signals to learn good semantic similarity and discriminability of degraded images; it can adaptively select high-quality triple training samples to accelerate model convergence and improve model training efficiency; under the guidance of quantization loss and discriminability loss, the image hash code output by the trained model can maintain high semantic similarity and discriminability at the same time, thereby improving the retrieval quality of degraded images; and it can realize efficient retrieval of large-scale degraded image data by utilizing the hash collision principle.
[0136] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for retrieving degraded images, characterized in that... Includes the following steps: Step 1: Train the deep hashing network model; Step 2: Input the image into a deep hashing network model to learn the latent hash code represented by a binary vector; Step 3: By introducing a minimum hash algorithm, the potential hash code represented by the binary vector is converted into a minimum hash code; Step 4: Use the hash collision principle to retrieve degraded images; Step 1 includes: 1.1 Training sample data augmentation; 1.2 Training sample data grouping; 1.3 Triple training sample data selection; Step 1.1 also includes: performing various data augmentation methods on all images in the training sample dataset, including rotation, flipping, blurring, sharpening, and denoising; Step 1.3 also includes: In the candidate benchmark sample group A single sample of data is randomly selected without repetition as the baseline sample. Candidate benchmark sample group It is a dataset obtained by copying the original dataset multiple times; In the candidate positive sample group Select one that is similar to the benchmark sample Samples with the same serial number are considered positive samples. The candidate positive sample group is a dataset consisting of images from the original dataset that have undergone five different data augmentations. In the candidate negative sample group A sample is randomly selected as the negative sample. The candidate negative sample set is a data set obtained by copying the candidate benchmark sample set and shuffling the sample numbers. Calculate the baseline sample Positive samples and negative samples The relative distance; first calculate , and The difference hash is obtained. , and , then calculate and The Hamming distance d1, and The Hamming distance d2 is calculated, and finally the difference between d1 and d2 is calculated. : ; Given a boundary threshold between positive and negative sample pairs ,judge Is it less than If less than If the negative sample is selected as a high-quality negative sample, then a new sample is selected from the candidate negative sample group; otherwise, a new sample is selected as the negative sample. , Recalculate the baseline sample Positive samples and negative samples relative distance Reassess Is it less than This process continues until the solution is found. Less than Negative samples; Select the benchmark sample Positive samples and corresponding high-quality negative samples Form a high-quality training triplet And put it into the high-quality training triple set. In this context, l represents the number of high-quality triplets; repeat the above steps until the baseline sample group is reached. No sample data is available. Step 1 also includes: calculating the model loss and updating the network parameters; wherein, calculating the model loss and updating the network parameters further includes: All triplet training samples in the kth batch are split in half, into a first half k1 and a second half k2. Calculate the depth feature vector of each reference sample in k1 one by one. With each baseline sample depth feature vector in k2 Cosine similarity between and the binary vector of each benchmark sample in k1 With each benchmark sample binary vector in k2 Jaccard similarity between and calculate and The mean squared loss is used to obtain the benchmark sample quantization loss. : ; Calculate the depth feature vector of each positive sample in k1 one by one. With each positive sample depth feature vector in k2 Cosine similarity between and the binary vector of each positive sample in k1 With each positive sample binary vector in k2 Jaccard similarity between and calculate and The mean squared loss is used to obtain the positive sample quantization loss. : ; Calculate the depth feature vector of each negative sample in k1 one by one. With each negative sample depth feature vector in k2 Cosine similarity between and the binary vector of each negative sample in k1 With each negative sample binary vector in k2 Cosine similarity between and calculate and The mean squared loss is used to obtain the negative sample quantization loss. : ; Calculate the quantization loss of the benchmark sample Positive sample quantization loss and negative sample quantization loss The arithmetic mean of the samples yields the quantization loss. : ; Calculate the binary vector of the benchmark sample in the kth batch. With positive sample binary vector Jaccard distance and the binary vector of the reference sample With negative sample binary vector Jaccard distance and calculated and Triple loss yields discriminative loss. : ; Weighted average calculation of sample quantization loss and sample discriminative loss The final network model loss L is obtained: ; The loss gradient is calculated based on L, and the gradient is backpropagated in the network model to update the parameters of each network layer in the network model. Step 4 further includes: calculating the minimum hash code for all images, including the degraded image to be queried and the database images; dividing the minimum hash code h into s groups, each group being a vector component occupying r = len(h) / s bits; designing a random mapping function to map the r-bit vector component in the group to a hash bucket, using the same mapping function for all groups, and each group using an independent hash bucket array; checking all database images that have a group mapped to the same hash bucket as the degraded image to be queried, and including these database images in the candidate similar image set; calculating the Jaccard similarity between the degraded image to be queried and the images in the candidate similar image set one by one; sorting the images by similarity from high to low, and selecting the K images with the highest similarity as the top K images most similar to the degraded image to be queried.
Citation Information
Patent Citations
Remote sensing image retrieval method and device based on category-level semantic hash
CN113190699A